Skip to main content

Running Growth Experiments One at a Time

Experiments · 10 min read ·

A disciplined loop for small teams: pick one hypothesis, change one thing, measure, decide. How to run growth experiments that teach you something.

Illustration: A circular loop drawn as an isometric track with four nodes labelled Hypothesis, Change, Measure, Decide, a small cyan dot circling it

A team that makes six changes in a week and sees sign-ups rise has learned very little. Which change helped? Did any hurt? Was it the changes at all, or a mention in a newsletter that nobody noticed? Growth that cannot be explained cannot be repeated.

The remedy is a habit borrowed from science: change one thing, predict what will happen, measure and decide. It is slower in the short run and faster in the long, because every result teaches something. This guide describes how to run growth experiments in that way, in the small, noisy conditions that early products face.

The idea of an experiment

The scientific method, in outline, is a cycle: observe, form a hypothesis, test it, analyse the result and revise. Its power lies in the discipline of prediction. A hypothesis, stated before you look at the data, can be wrong, and being wrong is how you learn.

Applied to a product, an experiment is simply a deliberate change with a stated prediction and a defined measure. It is not a hunch carried out with enthusiasm. The difference lies in the writing down.

One change at a time

The central rule is to change one thing. Two changes at once make a result ambiguous: if sign-ups rise, you cannot say which change did it, and if they fall, you cannot say which to reverse.

This is the idea behind a control group, a set of subjects not exposed to the change, used as a baseline for comparison. You may not have a formal control, but you can always compare with the period before, or with a group that did not receive the change. The cleaner the comparison, the clearer the lesson.

Practical consequences:

  • Do not redesign the homepage, change the pricing and send a newsletter in the same week.
  • If several ideas seem good, rank them and queue them.
  • If a change is bundled by necessity, such as a new release, record everything that changed.
  • Avoid other major events during the test, such as launches or holidays, when you can.

Start with a question

Good experiments answer questions you care about. At your stage, the question should come from your current milestone. If your milestone is to raise the share of new users who complete setup, your questions are about setup.

Examples of useful questions:

  • Why do people leave before finishing setup?
  • Which message brings people from the page into the product?
  • Does a short welcome email improve return use?
  • Which audience responds best?

Avoid questions with no decision attached. If any answer would lead to the same next step, do not spend a week answering it.

Write the hypothesis

A hypothesis states what you will change, what you expect and why. A simple template:

"If we [change], then [measure] will [direction and amount], because [reason]."

For example: "If we cut the sign-up form from five fields to two, then the share of visitors who finish sign-up will rise from 30 to at least 38 per cent over two weeks, because fewer fields reduce effort."

The reason matters. It turns the experiment into a test of an idea about users. When the result comes, you learn whether the idea holds, not only whether the number moved.

Choose the measure

Decide what you will measure, and how, before you begin.

  • Primary measure: the one number that decides success, tied to your milestone.
  • Guard measures: things that must not get worse, such as retention or support requests.
  • Baseline: the current figure, measured over a similar period.
  • Time window: how long you will run it, set in advance.

Pick measures that you can actually observe. If your analytics cannot see the step you are testing, fix that first.

Avoid vanity measures. A change that increases clicks while reducing completed use is not an improvement.

Small numbers, honest results

Early products have few users, so results are noisy. Statistical significance is a way of judging whether an observed difference is unlikely to have arisen by chance alone. With small samples, it is hard to reach, and a naive reading of the numbers can mislead.

A few practical principles:

  • Look at counts. Fifteen out of thirty and nineteen out of thirty is a difference of four people.
  • Do not stop early because the numbers look good. Stopping when the result is favourable inflates false positives.
  • Expect noise. A normal week-to-week swing can be larger than the effect you are trying to measure.
  • Prefer big effects. If a change needs a large sample to detect, it may be too small to matter at your stage.
  • Repeat important results. A second run is stronger evidence than a first.
  • Combine with conversations. Numbers say what happened; people say why.

A/B testing, comparing two versions shown at random to different users, is the standard method when you have enough traffic. With a few hundred visitors a week, it will rarely give a clear answer. In that case, run sequential tests: one version for a week, another for the next, aware of the weaknesses of that method, and supplement with user interviews.

Run it

  1. Write the experiment card. Hypothesis, change, measure, guard measures, baseline, start and end dates, owner.
  2. Make the change, and only that change.
  3. Record the start with the date and version.
  4. Leave it alone until the end date. Do not tweak midway.
  5. Watch for problems, such as errors, but do not judge the result yet.
  6. Collect the data at the end, using the same method as the baseline.
  7. Compare and decide.

Keep every card in one place. After ten experiments, the collection is a valuable record of what you have learned.

Deciding

At the end, make a decision from three options.

  • Keep the change. The result met your prediction and the guard measures held.
  • Revert it. The result was worse, or the change caused harm.
  • Inconclusive. The difference is within the noise. Decide whether to extend, redesign or move on.

Be honest about the third. Many experiments are inconclusive, especially at small scale. An inconclusive result is not a failure, but treating it as a success is a mistake.

Whatever the outcome, write down what you learned about users, not only about the number. "People do not mind a longer form if they know how many steps remain" is a lesson that travels.

Choosing what to test first

With a queue of ideas, a quick ranking helps.

  • Potential impact: how much could this move the main measure?
  • Confidence: how sure are you, from evidence, that it will work?
  • Effort: how long will it take?

Prefer high impact and low effort. Be wary of experiments that need a long build before you learn anything. If an idea is big, find a smaller version that tests the same assumption: a manual version, a single page or a hand-sent message.

Ethics and care

Experiments involve real people.

  • Do not deceive users in ways that harm them. Testing wording is fine; testing a false claim is not.
  • Respect privacy. Collect only what you need and tell people what you collect.
  • Watch for harm. If a variant confuses or disadvantages a group, stop.
  • Do not manipulate. Tactics that pressure or trick people may raise a number for a week and erode trust for a year.
  • Avoid fake urgency and fake scarcity.

A growth practice that users would be uncomfortable seeing explained is a growth practice to avoid.

Common mistakes

  • Testing without a hypothesis, which turns the exercise into a random walk.
  • Changing several things.
  • Moving the goalposts after seeing the result.
  • Ignoring seasonality, such as comparing a holiday week with a normal one.
  • Over-reading small samples.
  • Running too many at once, which collides with the one-at-a-time rule.
  • Never concluding. Experiments that run forever produce nothing.
  • Forgetting to record. A lesson not written down is likely to be relearned at cost.

A worked example

A team's milestone is to raise the share of new sign-ups who create a first project from 35 to 45 per cent. They suspect that the empty dashboard confuses people. Their hypothesis: "If we show a short example project on first login, then at least 45 per cent of new users will create their own project within a day, because they will see what to do."

They write the card: change is the example project, primary measure is the share creating a project within a day, guard measures are support questions and one-week return. Baseline: 35 per cent over the last three weeks, 140 sign-ups. They run it for two weeks, alone, with no other changes.

In the test period, 118 people sign up, and 51 create a project, which is 43 per cent. Support questions are unchanged and return is similar. The result is close to the prediction but short of it, and with 118 people the margin is uncertain. They decide to keep the change, since nothing got worse, and run the next experiment on the empty dashboard's wording, noting that the first result is suggestive, not conclusive. They also talk to six new users and learn that the example helped, but that the button to start was hard to find. They move it, and run that as the next experiment.

Questions makers ask

Is this too slow? It feels slower than changing everything, but you stop repeating mistakes.

Can I test marketing messages the same way? Yes. Change one message, one place, for a set period.

What if I have almost no traffic? Use interviews and manual tests: show two versions to five people each and watch.

How many experiments should run at once? One main one per milestone, perhaps a small second in an unrelated area.

Summary

Change one thing, state what you expect, define the measure and decide in advance how long to run. Respect small numbers: look at counts, avoid stopping early and repeat the results that matter. Decide to keep, revert or extend, and record what you learned about users, not just about the figure. Rank ideas by impact, confidence and effort, behave ethically and keep every card. Growth that you can explain is growth that you can repeat.

Questions and answers

What is a growth experiment?
A deliberate change to the product or its promotion, made to test a stated prediction and measured against a defined result.
Why change only one thing at a time?
So that you can tell which change caused the result. With several changes, the cause is hidden.
Can I run A/B tests with few users?
Often not reliably, because small samples produce noisy results. Use sequential tests and conversations instead.
How long should an experiment run?
Long enough to see a full cycle of use, often one to two weeks, and decided in advance.
What if the result is unclear?
Treat it as no evidence. Do not claim success or failure; refine the hypothesis or the measure and try again.

Sources

Support