Precision can survive a broken experiment

A small p-value and narrow interval describe calculations under a model; they do not certify that assignment, exposure, logging, exclusions, or the decision metric were valid. Begin diagnosis before reading the treatment effect. Compare configured and observed allocation, check eligibility counts and invariant pre-treatment measures, inspect releases and event volume, and reconstruct several units end to end. Online experimentation research treats sample-ratio mismatch as a serious warning because unexpected variant proportions can reveal assignment, logging, or data-loss defects. If the trust checks fail, hide the scorecard. Explaining the mismatch is the task; narrating the apparent winner is not.

Evidence: Encyclopedia of Machine Learning and Data Science; National Institute of Standards and Technology

Assignment and exposure may describe different experiments

Random assignment supports a causal comparison when units retain their assigned condition and outcome measurement follows the design. Diagnose cross-device reassignment, cookie deletion, shared accounts, link sharing, cached interfaces, staff overrides, and treatment features that fail to load. A treatment-on-the-treated analysis based on who successfully received the feature can reintroduce selection; intent-to-treat usually preserves the assignment comparison. Check how many assigned units were actually eligible, rendered, and measurable, but do not casually exclude failures caused by the treatment itself. If people influence one another, the independent-unit assumption may also fail. The next observation is an assignment-to-exposure transition audit by unit and version.

Evidence: National Institute of Standards and Technology; Encyclopedia of Machine Learning and Data Science

Instrumentation can move while behavior stays still

A variant may emit an event twice, suppress it during latency, rename a property, change consent flow, or alter when a session closes. Compare server and client counts, raw and derived events, pre-treatment invariants, and an A/A test where feasible. Trace one synthetic account through assignment, exposure, action, delayed outcome, and aggregation. Watch for different missingness by variant and for filters defined after outcomes are known. A visually plausible dashboard is weak evidence if its lineage is opaque. The separating observation is whether an independent operational record—orders, confirmations, or support state—moves consistently with the analytics event under the same assignment.

Evidence: Encyclopedia of Machine Learning and Data Science; UK Government Digital Service

Analysis flexibility can manufacture a winner

Search the audit trail for repeated peeking, a changed stop date, newly privileged outcomes, dropped days, unplanned segments, several variants, and one-sided tests selected after seeing direction. Each additional opportunity to declare success alters the false-positive risk. NIST's multiple-comparison material explains why a family of comparisons needs procedures that control joint error rather than separate unadjusted claims. Compare the scorecard with the timestamped plan. Re-run the planned analysis without post-outcome exclusions, label all other analyses exploratory, and report the full set attempted. If no plan exists, the result may still generate a hypothesis, but its confidence label should not imitate a clean confirmatory test.

Evidence: National Institute of Standards and Technology; Center for Open Science

A valid effect can answer the wrong decision

Suppose a banner reliably increases clicks but purchases, comprehension, or refund-adjusted value do not improve. The estimate can be statistically correct and operationally irrelevant. Revisit the decision sentence, primary outcome, unit, window, and minimum worthwhile effect. Check guardrails and groups likely to bear harm. A short test may capture novelty, while delayed cancellations fall outside the window. A session metric can overweight frequent visitors compared with account-level welfare. The next observation is not another segment search; it is a decision-alignment review that connects each metric to the action and consequence it claims to represent.

Evidence: UK Government Digital Service; Encyclopedia of Machine Learning and Data Science

Close the diagnosis with one of four verdicts

Classify the run as valid and decision-relevant, valid but inconclusive, exploratory only, or invalid. For a valid run, report effect size, uncertainty, quality checks, guardrails, and the attribution boundary. Inconclusive means the interval still contains materially different decisions, not that treatment and control are identical. Exploratory findings become inputs to a new prospectively defined test. Invalid runs preserve incident evidence but cannot crown a variant. Center for Open Science's connected-record approach supports linking the prior plan, deviations, data, analysis, and outcome. Record the defect owner and retest condition. A disciplined invalidation is a successful diagnosis because it prevents false confidence from becoming product policy.

Evidence: Center for Open Science; Encyclopedia of Machine Learning and Data Science; National Institute of Standards and Technology

Sources and further reading

These references informed this article. A source supports a claim; it does not imply endorsement of TenMultigure or any future product reference.

  1. Completely randomized designsNational Institute of Standards and Technology · Accessed August 10, 2026

    Defines the assignment logic against which contamination, unit instability, and post-assignment selection are diagnosed.

  2. Multiple comparisonsNational Institute of Standards and Technology · Accessed August 10, 2026

    Grounds the diagnostic warning about interpreting many outcomes, variants, or subgroups with isolated unadjusted thresholds.

  3. How the alpha phase worksUK Government Digital Service · Accessed August 10, 2026

    Supports returning to the risky assumption and decision evidence when a technically valid metric does not answer the product question.

  4. Online Controlled Experiments and A/B TestsEncyclopedia of Machine Learning and Data Science · Accessed August 10, 2026

    Contributes real-world failure checks for sample ratios, exposure, power, instrumentation, and extreme scorecard results.

  5. Lifecycle Open ScienceCenter for Open Science · Accessed August 10, 2026

    Supports diagnosing deviations against a prospective plan and retaining null, exploratory, contradictory, and invalid results in one traceable record.

Reviewed for clarity and evidence

Reviewed by TenMultigure Editorial Review. See an error or a source that has changed? Tell the editorial team.

Review method: AI-assisted desk research with editorial checks. Reviewed ; next scheduled review . Built a diagnostic sequence for allocation, exposure, instrumentation, analysis flexibility, decision alignment, and four evidence verdicts.