Compare designs on the claim you need to make
Use the same criteria for every option: decision to support, causal strength, assignment feasibility, expected effect, traffic and duration, spillover, reversibility, mechanism uncertainty, harm, and operational cost. A randomized test is not automatically ethical or informative; a qualitative study is not a small A/B test; a before–after chart is not automatically causal. Write the desired claim in advance. 'People can understand the new refund explanation' requires different evidence from 'offering this explanation changes qualified purchase by at least two percentage points.' If the decision can be made with a lower-risk method, do not collect a larger behavioral data set merely because a platform makes it easy.
Evidence: UK Government Digital Service; Center for Open Science
Randomized comparison: strongest for a bounded treatment effect
Random assignment is favored when eligible units can receive stable variants concurrently, interference is manageable, outcomes are measurable, and enough information can accumulate before the context changes. It estimates the effect of assignment under the design's assumptions, with quality checks for allocation, exposure, missingness, and instrumentation. It is a poor choice when the treatment cannot be isolated, spillover is dominant, traffic is too low for the decision-relevant effect, or serious harm cannot be contained. Avoid it when the team has not yet defined a coherent alternative; randomizing two confusing experiences produces a precise comparison of the wrong options.
Evidence: National Institute of Standards and Technology; Encyclopedia of Machine Learning and Data Science
Before–after or interrupted time study: practical but confounded
A time-based comparison may be necessary for infrastructure, policy, pricing, or population-wide changes that cannot run concurrently. Multiple stable pre-periods, an unaffected comparison series, a prespecified interruption, and consistent measurement can make it more informative than a single week-over-week number. Seasonality, campaigns, competitor moves, learning, traffic mix, and simultaneous releases remain rival explanations. Choose it when rollout is irreversible or interference prevents individual assignment and the historical series is rich enough to model. Avoid a causal label when only one pre point and one post point exist. The attribution boundary must name the alternatives the design could not rule out.
Evidence: UK Government Digital Service; Encyclopedia of Machine Learning and Data Science
Qualitative probe: best for mechanism and design uncertainty
Task observation, interviews tied to a real decision, prototype sessions, and support-case analysis can reveal misunderstanding, workarounds, trust gaps, and recovery paths before scale testing. They help decide what treatment and outcome are worth testing. They do not estimate a population effect from a small purposive sample. Choose this option when the competing mechanism is unclear, the interface is early, failure could harm people, or traffic cannot support a useful randomized comparison. Avoid asking preference questions as a proxy for behavior or turning participant counts into percentages. The deliverable is an evidence-linked mechanism hypothesis and revised design, not a declaration of conversion lift.
Evidence: UK Government Digital Service; Center for Open Science
Staged evidence often dominates a forced either-or choice
A responsible sequence may use support evidence to identify a risky transition, qualitative sessions to expose the mechanism, a randomized test to estimate a bounded effect, and a later time-series review for durability. Each stage has its own stop rule. GOV.UK's alpha guidance is useful here: test the riskiest assumption with the cheapest adequate prototype rather than build the whole solution. Do not call the sequence triangulation if every source depends on the same flawed event. Preserve contradictory results. If the qualitative mechanism disappears in a representative implementation or the randomized effect fails to materialize, update the model instead of averaging unlike evidence into a positive story.
Evidence: UK Government Digital Service; National Institute of Standards and Technology; Encyclopedia of Machine Learning and Data Science
The decision matrix ends with who should avoid each option
Avoid randomized testing when assignment cannot remain stable, likely spillover overwhelms contrast, or the required sample outlives the decision. Avoid before–after claims when concurrent changes and seasonality cannot be observed well enough to challenge the story. Avoid qualitative-only evidence when the decision depends on prevalence or a small causal effect. Avoid all three when the primary action is deceptive, unsafe, or lacks a legitimate user benefit; methodology does not cleanse the intervention. Record the selected design, rejected alternatives, assumptions, cost of error, and recheck trigger. This comparison recommends no commercial tool because design suitability precedes platform procurement.
Evidence: Center for Open Science; National Institute of Standards and Technology; Encyclopedia of Machine Learning and Data Science
Sources and further reading
These references informed this article. A source supports a claim; it does not imply endorsement of TenMultigure or any future product reference.
- Completely randomized designsNational Institute of Standards and Technology · Accessed August 10, 2026
Defines the randomized-design option and clarifies why it depends on eligible experimental units, assignment, treatment, and outcome structure.
- How the alpha phase worksUK Government Digital Service · Accessed August 10, 2026
Provides the staged-evidence principle of testing risky assumptions with proportionate prototypes and explicit success criteria before full implementation.
- Online Controlled Experiments and A/B TestsEncyclopedia of Machine Learning and Data Science · Accessed August 10, 2026
Contributes practical feasibility and trust criteria for online randomized tests, including power, assignment, sample ratios, and instrumentation.
- Lifecycle Open ScienceCenter for Open Science · Accessed August 10, 2026
Supports prospective design records, transparent deviations, connected outputs, and preserving null or contradictory evidence across staged methods.
Reviewed by TenMultigure Editorial Review. See an error or a source that has changed? Tell the editorial team.
Review method: AI-assisted desk research with editorial checks. Reviewed ; next scheduled review . Compared randomized, time-based, and qualitative evidence on one decision matrix, included disqualifying contexts, and proposed a staged alternative.