What is a product experiment, and how is it different from user research? #

A product experiment, or business experiment, tests a stated hypothesis by putting something realistic in front of real users under real conditions and measuring what they do. User research explores what people need and why; an experiment is built to confirm or refute a specific belief with evidence. In the Innovation Mode methodology it is the core instrument of opportunity validation: the way a team concludes with confidence whether an opportunity is worth building.

  • Research explores and an experiment decides. Interviews tell you what people say they want; an experiment shows what they actually do when it is offered
  • Every experiment starts from a written hypothesis with a metric and a threshold, for example 'at least 15% of visitors who see the offer ask for access'
  • It captures two kinds of data: implicit data from how people interact, such as clicks and usage events, and explicit data from what they tell you in the product or in a follow-up survey
  • The thing being tested is usually a prototype, but a focused one, built to test the hypotheses rather than to show how the finished product would look
  • The audience is chosen deliberately and reached through a defined channel, so the result can be read and repeated
  • Research and experiments work in sequence: discovery research produces the hypotheses, and experiments test them
Key Takeaway

If you already know what you believe and need to know whether it is true, you need an experiment. If you do not yet know what to believe, you need research first.

Inside Ainna Many ideas solve the wrong problem. Ainna reframes it before you build on a shaky premise. Frame my opportunity

What is the difference between a risk and an uncertainty, and why does it matter for testing? #

A risk is a negative outcome you can name and estimate: how likely it is, how much it would hurt and what it costs to prevent. An uncertainty is something you cannot estimate in advance because the information does not exist yet, such as how people will behave with a genuinely new product. Risks are handled with analysis and mitigation; uncertainties are handled with experiments. I find the popular word 'derisking' misleading for exactly this reason: teams plan for risks and miss the uncertainties that decide the outcome.

  • Risks you can assess early: the product does not solve the problem, adoption is lower than expected, it fails under real-world load, competitors respond, it cannot scale, it breaches a regulation, it damages the brand
  • Uncertainties you cannot plan away: how users will behave with a new experience, what an emerging technology does to the market, shifting social norms, changes in the ecosystem, sudden regulatory or geopolitical shifts
  • Testable uncertainties, such as user behaviour, pricing and features, call for experiments. Untestable ones, such as a regulatory change, call for agility and a prepared pivot path
  • Silent assumptions are a third category: beliefs nobody has written down, so nobody tests them. They are often the most dangerous, because a whole plan rests on them unnoticed
  • Trying to control every risk creates a new one: process and cost that make the innovation effort slow and expensive
  • The aim is risk intelligence, knowing the source, nature and likely impact of each unknown, rather than a promise that no risk remains
Key Takeaway

Before designing any test, sort your unknowns into risks, testable uncertainties, untestable uncertainties and silent assumptions. Only the second group needs an experiment, and it is usually the group that decides whether the product works.

Which assumptions should you test first? #

Test the assumption that would kill the opportunity if it were wrong and that you currently have the least evidence for. That is usually a belief about demand or behaviour, not about technology. Rank your assumptions by how much the plan depends on them and how little you actually know, then run the cheapest experiment that can settle the top one.

  • List every belief the plan depends on, including the ones that feel obvious, because the obvious ones are where silent assumptions hide
  • For each, ask two questions: if this is false, does the opportunity survive, and what evidence do we have today?
  • Demand and behaviour usually come first: will people want it, will they change what they do today, will they pay?
  • Technical feasibility comes first only when the technology is genuinely unproven; otherwise a quick feasibility check is enough
  • Test one critical assumption per experiment where you can. Experiments that test five things at once rarely produce a clear answer
  • Write down in advance the result that would make you stop. Without it, every result gets read as encouraging
Key Takeaway

Order matters more than volume. One experiment on the assumption that can kill you is worth more than ten on the ones that cannot.

What is a fake door test, and when should you use one? #

A fake door test, also called a smoke test, presents a product or feature as if it existed, usually on a landing page or as a button inside an existing product, and measures how many people try to go through the door. It is the cheapest way to measure interest before building anything. It is also the easiest to misread: a click shows curiosity, not commitment, so it should inform a decision rather than make it alone.

  • Use it when the biggest uncertainty is demand: would anyone want this enough to act?
  • Describe the offer realistically, with a clear call to action such as 'request access' or 'see pricing', and measure the share of visitors who take it
  • Be honest at the door: tell people it is not available yet and offer to notify them. Deceiving users beyond that costs trust you will need later
  • Compare against a threshold set in advance, and ideally against a second variant, so you read a difference rather than a raw number
  • Treat a strong result as permission for a more expensive test, such as a functional prototype, not as proof of product-market fit
  • Buffer started with a two-page test of this kind: a page describing the app and a page collecting the emails of people who wanted it; a pricing page added between the two then measured which plan people clicked
Key Takeaway

A fake door test answers one question well, whether anyone is interested, and every other question badly. Use it early, cheaply, and as the first step in a sequence.

What is a Wizard of Oz test? #

A Wizard of Oz test gives users what looks like a working, automated product while people deliver it by hand behind the scenes. Users interact as if the system were real, which shows whether the experience works and whether they value it, before you build the automation. It is the right method when the product's promise depends on something expensive to build, such as an AI capability or a complex back end.

  • The user sees a finished-looking front end; the machine behind it is your team doing the work manually
  • It tests the experience and the value together: do people use it the way you expected, and do they come back?
  • It suits AI products well: a person produces the output the AI would, and you learn what users need from it before training or integrating anything
  • Keep it small and time-boxed. Manual delivery does not scale, and that is fine, because scale is not what you are testing
  • Disclose responsibly: run it with participants who have agreed to try an early service, tell them afterwards that people delivered it, never let staff see personal or confidential data users believe only a machine will process, and never charge for automation that does not exist
  • Zappos began this way: its founder photographed shoes in local stores, put them online, and bought each pair from the store at full price only when an order came in
Key Takeaway

Use a Wizard of Oz test when you are confident about the solution and uncertain whether people will adopt it. It lets you learn from real behaviour at a fraction of the cost of building the real thing.

What is a concierge MVP, and how is it different from a Wizard of Oz test? #

A concierge MVP delivers the service by hand, openly: users know a person is doing the work, and the team works closely with each one to learn what they really need. A Wizard of Oz test hides the manual work and simulates the finished product. Concierge is for learning what to build; Wizard of Oz is for testing whether what you plan to build will be adopted.

  • Concierge happens in plain view, Wizard of Oz behind a curtain, and that single difference changes what each can teach you
  • Use concierge when you are still unsure what the solution should be. Working with users one by one reveals needs no survey would
  • Use Wizard of Oz when the solution is defined and the open question is adoption and usability
  • Concierge works with very few users, often a handful, because each one takes real effort. That is its strength, not a flaw
  • The two often run in sequence: a concierge MVP to learn what to build, then a Wizard of Oz test to check the automated version will be used
  • The meal-planning service Food on the Table, an example in The Lean Startup, began this way: one customer, menus planned by hand and a weekly in-person meeting, and it signed up more customers before investing in automation
Key Takeaway

If you cannot yet describe the product precisely, start with concierge. If you can, and the question is whether people will use it, move to Wizard of Oz.

How do you run an experiment with a functional prototype? #

Build an inexpensive working version that contains only the features tied to your hypotheses, release it to a small, deliberately chosen audience, and measure both how they use it and what they tell you. In my experience this produces the richest evidence of any experiment type, because people use something real under real conditions. It also costs the most, so it should come after cheaper tests have justified it.

  • Map each hypothesis to the features that will produce evidence for it, and build those features properly. Everything else can be missing or simulated
  • Leave out supporting functions far from the core idea, such as search, profiles or settings, unless a hypothesis depends on them
  • Instrument it before launch: decide which usage events will show whether each hypothesis holds
  • Release it in stages, invitation only at first, to an audience you recruited for the purpose
  • Collect explicit feedback alongside the usage data, through short in-product prompts and a follow-up survey
  • Set an end date and early-stop conditions, so a failing test is ended rather than extended
Key Takeaway

A functional prototype is a measuring instrument, not an early version of the product. Build it to answer the hypotheses, and resist every feature that does not.

Did you know? The Judge, Ainna's AI evaluator, scores your opportunity across ten dimensions and shows the reasoning behind every score, with suggestions for improving it. Get an honest assessment

What is the difference between an A/B test and a multivariate test? #

An A/B test shows two or more versions of one thing to randomly split groups of users and measures which performs better. A multivariate test changes several elements at once and measures which combination works best. Both need a formal hypothesis and enough traffic to reach statistical significance, and a multivariate test needs far more traffic, because every combination needs its own sample.

  • A/B: one variable, one clear answer. Use it to compare two onboarding flows, two headlines or two layouts
  • Multivariate: several variables, the best combination. Use it only when traffic is high and the elements interact
  • Randomise who sees what, or the result reflects who happened to arrive rather than what you changed
  • Fix the metric and the smallest difference worth detecting before you start, and do not stop the moment the result looks good
  • Both methods optimise what already exists. They tell you which version of a feature is better, not whether the product should exist
  • For a new concept, an out-of-product experiment is usually the better first step
Key Takeaway

Choose A/B for a clear answer to one question, and multivariate when traffic is abundant and elements interact. Neither replaces the earlier question of whether the idea deserves a product at all.

How do feature experiments and pricing experiments work? #

A feature experiment adds or removes a capability for a subset of users and measures the effect on the metrics that matter, such as activation, retention or task completion. A pricing experiment tests different price points or pricing models to find what customers accept and what maximises value. Both run inside a live product, usually released to a small random group first, and the winning variant then goes to everyone.

  • Feature experiments answer 'does this change behaviour?' rather than 'do people like it?', so measure usage and outcomes, not opinions
  • Remove as well as add: testing whether a feature can be taken away is one of the cheapest ways to simplify a product
  • Release through feature flags to a small share of users, so a bad variant can be switched off without a new release
  • Pricing experiments can test price levels, packaging, trial length or the model itself, such as subscription against usage-based
  • Price tests carry more risk than feature tests, because customers compare notes and consumer-protection rules apply. Test on new customers or in separate markets where you can, and check the pricing rules in each market before showing different prices to different people
  • Look past conversion. A lower price can raise sign-ups and still reduce revenue or attract the wrong customers
Key Takeaway

In-product experiments turn the product into a learning system. The discipline is the same as for any experiment: a hypothesis, a threshold and a decision agreed before the data arrives.

Very often, silent assumptions hide in plain sight, embedded in the ways of thinking and the collaboration patterns, and they are not even recognized or called out.

How do you choose the right type of experiment? #

Start from the uncertainty, not the method. If the question is demand, a fake door test can answer it cheaply. If it is whether a new experience will be adopted, you need something people can really use: a functional prototype, or a Wizard of Oz test when the automation behind it is expensive to build. If you do not yet know what to build, start with concierge. Then decide the scale, meaning how many people and for how long, and in most cases combine more than one method in sequence.

  • Do you have a live product? If yes and the change is small, test inside it. If the concept is new, or testing it would disrupt the product, test outside it
  • Uncertainty about demand: a fake door or landing page test
  • Uncertainty about adoption of a new experience: a functional or clickable prototype
  • Uncertainty about what the solution should be: concierge. About whether a defined solution works before you automate it: Wizard of Oz
  • Uncertainty about price or packaging: pricing experiments, inside the product if you have customers, on a landing page if you do not
  • Sequence from cheap to expensive, and let each result justify the next test
Key Takeaway

The most common mistake is choosing the method first because the team already knows how to run it. Name the uncertainty, and the method usually follows.

Six tiles pairing an uncertainty with the experiment that answers it: demand, adoption, what to build, adoption before automating, price or packaging, and a small change to a live product.
Name the uncertainty and the method follows; then set the scale and sequence tests from cheap to expensive, each result justifying the next.

What should a business experiment template include? #

The Innovation Mode Business Experiment Framing Template has six parts: identifiers, learning goals, form factor, audience, planning and decisions. Learning goals are the core: the learning objective, the list of hypotheses with the metrics and thresholds that would confirm or refute each one, and the success criteria for the experiment itself. Writing all six down before launch is what separates an experiment from an activity.

  • Identifiers: a unique ID, a title, a short description and one named owner
  • Learning goals: the learning objective, the product or initiative it relates to, the hypothesis list with thresholds and a baseline, and the success criteria
  • Form factor: what participants will interact with, such as a landing page, a clickable or functional prototype, a physical prototype or a proof of concept
  • Audience: named segments, the channel that reaches each one, and the target sample size
  • Planning: a launch checklist, start and end dates, and end conditions that stop the experiment early if it clearly fails
  • Decisions: what happens on each outcome and who decides, written before the data arrives and completed with the actual decision afterwards
Key Takeaway

A shared template does more than improve each experiment. Kept in one place, experiment proposals become discoverable, so product teams can pick up promising ones instead of starting from nothing.

Six stacked parts of an experiment template: identifiers, learning goals, form factor, audience, planning and decisions, each with what it records.
Write all six parts down before launch; that is what separates an experiment from an activity.

What form can an experiment take: landing page, prototype or proof of concept? #

There are five common form factors, and each collects different evidence. A landing page measures interest. A clickable prototype tests usability and perceived value. A functional prototype captures real usage and feedback. A physical prototype tests form, fit and function for hardware. A proof of concept shows that the technology can work. Choose the cheapest form that can produce the evidence your hypotheses need.

  • Landing page: presents the offer realistically and measures engagement, before any development
  • Clickable prototype: a basic model of the interface for the core flows, good for usability and value feedback
  • Functional prototype: an inexpensive working version that captures interaction data and direct feedback
  • Physical prototype: a draft of a physical product for hands-on testing, sometimes connected with electronics and data
  • Proof of concept: a small implementation of specific aspects, built to prove technical feasibility rather than demand
  • A proof of concept answers 'can we build it?', so do not read it as evidence that anyone wants it. See prototype, proof of concept and MVP
Key Takeaway

Fidelity should follow the question. Building a functional prototype to test interest wastes weeks, and testing adoption with a landing page produces an answer you cannot trust.

Start from the right assumption Know which assumption to test first Ainna frames your idea and surfaces the silent assumptions behind it, so the experiment you design targets the unknown that matters. Surface my assumptions

What makes a business experiment successful? #

An experiment is successful when it produces clear, real-world evidence that settles a decision, whether that decision is to continue or to stop. An experiment that gives you the evidence to kill an opportunity has done its job perfectly. The failed experiment is the one that ends ambiguous, because it costs another round of prototyping and time without moving the decision.

  • Judge the experiment on what it taught you, not on whether the hypothesis held. Confirmation and refutation are both results
  • Success criteria belong to the experiment as a learning exercise: did it produce evidence strong enough to decide?
  • Killing a weak opportunity early is one of the highest-value outcomes in innovation, because it frees money and people for better ones
  • Ambiguous results usually trace back to design: vague hypotheses, too small an audience, or a prototype that did not represent the concept
  • Report results against the thresholds set in advance, not against whatever looks encouraging afterwards
Key Takeaway

Measure an experiment by the quality of the decision it enables. By that standard, a clean 'no' is a success and a muddy 'maybe' is the real failure.

Why do experiments give misleading results, and how do you avoid it? #

The most common cause is a prototype that does not represent the concept well. When the experience is poor, users react to the prototype rather than the idea, and a good concept is rejected for the wrong reason. The fix is to build the parts that carry the hypotheses to a genuinely good standard and cut everything else, rather than building everything badly.

  • If users struggle with a clumsy interface, the data measures the interface, not the value of the concept
  • Invest quality where the hypotheses live and leave out the rest. A narrow, polished prototype beats a broad, rough one
  • Watch the audience as closely as the prototype: the wrong participants produce confident but irrelevant results
  • Guard against reading noise as signal: small samples and short runs produce swings that look like findings
  • Beware the team's own hopes. Thresholds set in advance and a named decision-maker keep the reading honest
Key Takeaway

Before trusting a negative result, ask whether the prototype gave the concept a fair chance. Before trusting a positive one, ask whether the audience was the one you will actually sell to.

What happens after the experiment ends? #

Compare the results with the thresholds, record what you learned, and make the decision you described in advance: build, iterate with a new experiment, pivot or stop. In the Innovation Mode methodology the validation team closes every experiment with a report and a recommendation, and the learning is stored so the next team does not have to repeat it.

  • Write the result against each hypothesis, confirmed, refuted or inconclusive, with the evidence
  • Make the decision explicit and name who made it. Drift, where nobody decides, is the most expensive outcome
  • If you continue, the next step is usually a more expensive test or an MVP, scoped by what the experiment taught you
  • If you stop, record why. A well-documented kill stops the same idea returning in six months untested
  • Store the experiment, its design and its data where other teams can find them
Key Takeaway

The experiment ends when the decision is made, not when the data stops arriving. What it taught you should outlive the experiment that produced it.

Beliefs like “users will naturally share this content” or “customers will pay a premium for this feature” should instantly raise a red flag.

Inside Ainna

Test the idea before you build it

Ainna applies the Innovation Mode methodology to your idea: it frames the problem, challenges your assumptions and scores the opportunity through The Judge, Ainna's AI evaluator, which shows its reasoning.

Ideas in →
Opportunities out.