A pilot is useful when it helps you make a specific next decision. Before it starts, name the question, the boundary, the evidence you will collect, and the date you will review it. At the end, separate three things: whether the test ran as planned, whether the intended signal appeared, and whether you can repeat the approach with the resources you have. Then choose whether to scale, adapt, test again, or stop.
A pilot is not a smaller version of a rollout. It is a bounded test that makes uncertainty easier to act on. If you start with a vague goal like “see how it goes,” a busy week or one enthusiastic response can become the story. A clear review rule gives the work a fairer read.
Start with the decision at the end
Write the decision you expect to make before choosing a metric. Are you deciding whether to offer the approach to another team, change one part of the process, run a longer test, or put the idea aside? Each answer needs different evidence.
“Did people like it?” is rarely enough by itself. An idea can feel promising while taking too much time to deliver. A process can run smoothly without improving the result it was meant to change. A useful pilot review holds outcome and practicality in view together.
Keep the question narrow enough for one test to inform. For example, a small studio testing a new project intake form might ask: “Does this form give us enough information to estimate a request without a separate clarification round?” That is easier to assess than “Will this improve how we work?” It names the decision the evidence should support.
Set the boundary before the test
Define the smallest real setting that can answer the question. Record:
- What is included: the part of the product, service, or process being tested.
- Who and where: the people, requests, or setting included in this round.
- When it starts and ends: the review date, with enough time for the signal you chose to appear.
- What stays the same: the existing approach or conditions you will compare against.
- Who decides: the person who will read the evidence and make the next call.
If several parts change at once, the pilot may tell you that the package felt different without showing which change mattered. Keep the test small enough to understand. A project baseline can help record the starting point and the comparison you intend to use.
Collect three kinds of evidence
A simple pilot can use a small evidence set. Pick a primary outcome, then collect enough context to interpret it.
1. Delivery
Did the test happen as described? Note who took part, what steps were completed, what changed during delivery, and where the process stopped. If only part of the pilot ran, record that before interpreting the outcome. A result from a different version of the test is still useful, but it answers a different question.
2. The intended signal
Choose one outcome that connects to the decision. For a workflow change, it might be whether requests arrive with the information needed to plan them. For a product test, it might be whether the intended customer can understand the offer and take the next step. Use a measure you can actually observe, and write down its source before collecting it.
Pair counts with a short explanation when the number alone could mislead. A total with no denominator, a change with no starting point, or a favorable comment from one person cannot carry more weight than it can support. If the pilot is too small to distinguish a pattern from a few individual cases, say so plainly.
3. Feasibility
Track the time, cost, attention, and support the test required. Ask whether people could complete it, whether the work fit alongside existing responsibilities, and what would have to be in place to repeat it. A process that only works with extra help from its designer may not be ready to expand unchanged.
This separation follows a useful distinction in the Australian Centre for Evaluation’s 2025 guide to evaluating pilot programs: understand implementation, look for early evidence of promise, test feasibility, then assess what wider use would require. A small team can use the same questions at a lighter scale.
Write the review rule while it is easy to be honest
For each kind of evidence, write what would count as enough, what would raise a concern, and what would leave the answer uncertain. Avoid borrowing a threshold from another business or inventing a number just to make the plan look precise. Choose a threshold that matches the decision, the stakes, and the amount of evidence the pilot can collect.
For the intake-form example, the review note might say: “We will test the form on one defined request type through the review date. We will compare the information received with the fields we need to estimate the work, record any clarification needed, and note the time spent by both sides. We will extend the test if too few requests arrive to judge it. We will revise the form if the same important information is repeatedly missing. We will consider wider use only if the information is sufficient and the extra effort is workable.”
This is a hypothetical planning example, not a result or a universal threshold. Its value is that the decision rules are visible before anyone sees the outcome. If there are material risks or dependencies, capture them in a project risk register and define what would pause the test.
Review the test before judging the idea
At the review date, begin with what actually ran. Compare the delivered test with its agreed scope. Then read the outcome and feasibility evidence. Ask whether the result reflects the idea, the way it was delivered, or a condition that changed during the test.
Keep “we did not observe the signal” separate from “the approach cannot work.” A short test, low participation, missing data, or a changed process can leave the answer open. Write that uncertainty into the decision rather than filling it with confidence.
If customer behavior is part of the question, short conversations about recent actions can add context to the numbers. The Journal’s guide to customer discovery interviews explains how to ask about what people have already done, rather than rely on a hypothetical promise.
Choose one of four next moves
- Scale carefully when the signal is promising, the test was delivered reliably, and the resources needed to repeat it are available. Start with the next comparable setting and keep monitoring.
- Adapt when the idea has promise but a clear friction point needs to change. Name the change and its reason, then treat the next round as a revised test.
- Test again when the evidence is incomplete or mixed. State exactly what you still need to learn and what the next test can show.
- Stop when the agreed result did not appear, the cost or risk is unacceptable, or there is no workable path to repeat the approach. Record what the pilot ruled out and close the open tasks.
Stopping a pilot can be a sound decision. So can refusing to scale a result that has only been shown in one narrow setting. The point is to make the next commitment proportional to what the test actually established.
Use a one-page pilot brief
Before the next test, save a short record with these fields:
- Decision this pilot will inform
- Question and boundary
- Starting point and comparison
- Primary signal, source, and review date
- Delivery and feasibility checks
- Stop condition and decision owner
- What evidence leads to scale, adapt, another test, or stop
After review, add the observed result, the limits of the evidence, the decision, and the next owner. That gives the next person something they can inspect and continue. For more practical systems for doing the work, visit The Self Made Journal.
A steady work block helps make room for the review and the next iteration. The Forged in Repetition Tee is part of the uniform we make for showing up to that work.
0 comments