DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk4 min

What a Coding Agent Taught Me About A/B Test Telemetry

A case study in using a coding agent for A/B test instrumentation and BigQuery queries—and why session modeling, data validation, and human interpretation still matter.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A coding agent helped Evgeny Khramov instrument a three-variant scanner-screen experiment, prepare its Firebase Analytics data for BigQuery, and write SQL queries. The more important lesson was about the data model: make a scan attempt the unit you can analyze, and keep the human responsible for deciding what the experiment can actually tell you.

What the agent helped build

Khramov describes an Android app used by store staff to scan price tags. The team tested three versions of the scanning screen. A coding agent helped define event attributes, implement the instrumentation, configure Firebase Analytics export to BigQuery, prepare a scanner_ab.sessions table, and write queries for analysis.

The agent could help translate questions such as “Compare A/B/C for the last three days,” “Break the results down by business unit,” or “Analyze by device model” into queries. But queries can only answer questions the experiment and its data can support. Khramov supplied the product question and retained responsibility for interpreting the result.

Firebase documents exporting Analytics data to BigQuery for SQL analysis. Export availability is not necessarily immediate: Firebase says the first export may take time, and describes daily syncs. See the Firebase BigQuery export documentation. Firebase also documents inspecting experiment and variant membership in Analytics event tables through BigQuery: Firebase A/B Testing and BigQuery.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Model the scan attempt, not just the events

The useful analytical unit in this example was a scan attempt, modeled as a session. The app sent a start event with a shared session_id, the test variant, store, device, and launch context. A finish event used the same identifier and included the result and scan details. That shared key made the beginning and outcome of an attempt joinable.

The prepared scanner_ab.sessions table represented one scan session per row. Instead of rebuilding each attempt from raw event records for every question, analysts could query a business-level attempt and segment it by relevant context. This is Khramov’s implementation choice, not a Firebase requirement.

Represent endings accurately

A user explicitly cancelling a scan is not the same as a session with no finish event. The first is an observed outcome; the second may indicate missing telemetry, an interrupted session, or a crash. Keep those cases distinguishable rather than silently classifying every absent finish as a normal cancellation. Where a session has no finish event, compare it with crash reports to investigate what happened.

Keep the prepared layer traceable

A session table saves repeated reconstruction work, but it should remain possible to trace a row back to the events that produced it. Khramov also describes a daily merge as part of his setup. Google Cloud supports recurring scheduled queries, but neither that feature nor Firebase export requires this particular table schema or merge strategy. See Google Cloud’s scheduled queries documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate what the app actually sends

A runbook can describe intended instrumentation without matching production data. Khramov found a disagreement between the runbook and observed parameter names or values. A query using the wrong value can return zero rows without an obvious SQL error, so verify the data itself before trusting a result.

  • Inspect actual event names, parameter names, values, and types in the exported data.
  • Check that each start event can be paired with the intended finish event using the shared session identifier.
  • Confirm that variant and context values are populated as expected, rather than relying only on documentation.
  • Convert string-valued fields safely before numeric analysis, and check that a field still measures what its name suggests.

Raw exports need care, too. Khramov reports that querying daily and intraday tables together with a wildcard can double-count overlapping data. Deduplicate or filter the records before aggregation, and verify that the query does not include the same events through both table types.

Align the analysis with assignment

The variants in this experiment were assigned by store. That assignment unit matters: scans from the same store are not automatically independent participants. Treating each scan as an independent observation can misrepresent how much independent evidence the experiment contains.

Before comparing variants, establish how assignment occurred and make the analysis reflect that unit. Then examine the outcome alongside relevant guardrails and segments, such as business unit, store, or device model, while checking data quality. A device or store breakdown can reveal an issue worth investigating; it does not by itself establish why the difference occurred.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the reported device result does—and does not—show

Khramov reports a 68.2% success rate on a Lenovo TB-8504X running Android 7.1.1, compared with a rate above 90% elsewhere. He says a crash was later confirmed by comparison with Crashlytics. This is a project-specific anecdote reported by the author in 2026, not an independent benchmark, representative estimate, or causal finding about the device or operating system.

The article’s variant comparison is illustrative and does not provide results. It establishes no winning variant, sample size, confidence interval, or overall effect. A query can summarize the recorded data; it cannot supply missing experimental evidence or settle whether a result is meaningful.

The human judgment that remains

Khramov puts the division of work this way: “I brought the product question, asked the questions in plain language, and remain responsible for the part that doesn’t come out of a query: how the experiment is set up and which conclusion the data actually allows.” That is his account of this project, not a universal claim about every coding agent or analytics workflow.

In this case, the agent helped with instrumentation, data preparation, and query execution. The experiment’s design, the choice of assignment unit, and the interpretation of evidence remained human responsibilities. The practical takeaway is to use automation to make sound telemetry easier to build and explore—not to treat a convenient query as a substitute for a sound experiment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.