Decision One: What Is Worth Building Around?

Task

Run the Stage 1 gate review.

Summary

Review the evidence and decide whether one customer problem and use case are clear enough to build around.

Decide What Customer Work Is Ready to Standardize

Task ID: S1-15

A Stage 1 gate review turns months of customer work into one explicit decision: either standardize one narrow offer or do not invest in doing so yet. This article explains how to judge demand, customer results, delivery repeatability, and financial evidence without mistaking enthusiasm, founder intuition, or a polished package for proof.

Why the gate exists

A consulting-led company can appear more repeatable than it really is.

Several customers may describe similar problems. The founder may use the same proposal deck, repeat familiar advice, and reuse parts of the delivery process. Revenue may be growing. Yet each sale can still depend on the founder’s reputation, each engagement can require a different solution, and each customer result can depend on unplanned senior effort.

That is the situation the Stage 1 gate review must resolve.

The review asks whether the company has found one sufficiently consistent combination of customer, problem, offer, and result to justify making the next investment in standardization. It is not a review of whether the company is good at consulting, whether customers like the founder, or whether the underlying technology works. It is a narrower decision:

Is there enough paid, operational, customer, and financial evidence to turn one part of the existing work into a repeatable offer?

This distinction matters because service productization is more than naming a package. Research describes it as a combination of defining and formalizing the offer, making it easier for customers to understand, and systematizing the processes through which it is sold and delivered. The objective is an offering that can be sold, delivered, and invoiced with less ambiguity—not merely a new page on the company’s website.

A gate review provides a deliberate stopping point between learning from customer work and investing in a standardized offer. Formal review systems in engineering and public projects use the same underlying principle: agree on the evidence required to enter and leave a phase, inspect that evidence, record unresolved actions, and decide whether the work is ready to proceed. NASA’s current guidance explicitly recommends tailoring review criteria to the project rather than treating a generic checklist as universally sufficient.

That is the right level of discipline here. The company needs a real decision, but not a committee-heavy bureaucracy.

What the decision must establish

By the time this review occurs, the team should already have collected evidence from its earlier sales, customer, delivery, and financial work. The gate is not the place to conduct all that research for the first time. It is where the team tests whether the separate findings tell one coherent story.

A credible go decision should establish five things.

First, a recognizable type of customer repeatedly experiences the problem. “Technology companies” or “small businesses” is usually too broad. The useful definition is specific enough to affect qualification, messaging, scope, delivery, and price. It may include the customer’s situation, operating maturity, trigger event, constraints, and people involved in buying the work.

Second, the customer considers the problem important enough to pay to address. Compliments, interview interest, newsletter sign-ups, and statements that an idea “sounds useful” are not equivalent to a purchase. Research on willingness-to-pay measurement has repeatedly found differences between hypothetical statements and decisions involving real economic consequences. A signed contract, deposit, paid pilot, renewal, or expansion therefore deserves more evidentiary weight than a survey answer about possible future buying.

Third, customers are buying substantially the same result. They do not need identical projects, but the core promise must be stable. A cybersecurity assessment for one customer, a full infrastructure rebuild for another, and an interim leadership role for a third may use some of the same expertise, but they are not one offer merely because the same founder delivers them.

Fourth, the work contains a repeatable delivery core. The team can identify what normally happens, in what sequence, with what inputs, using which tools, producing which outputs, and requiring how much time. Some professional judgment may remain essential. The issue is whether judgment is applied inside a stable method or whether every engagement has to be invented again.

Fifth, the financial pattern can support further investment. Revenue alone is not enough. The review must consider the labour, contractor expense, rework, customer support, software, transaction costs, and senior attention required to produce that revenue. Contribution margin—the amount remaining after the variable costs associated with delivering a sale—helps show how much revenue remains to cover fixed costs and profit. Break-even and sensitivity analysis then show how changes in price, cost, and volume affect the viability of the offer.

A no-go decision means that one or more of those conclusions cannot yet be supported. It does not necessarily mean the problem is unimportant or the company should abandon the market. It means the evidence does not justify depending on this offer as the company’s next repeatable foundation.

That is a valuable outcome. A 2024 replication involving four randomized trials and 759 firms found that teaching entrepreneurs a more scientific decision process increased the termination of ideas and produced a more selective pattern of strategic changes. The researchers interpret this as a combination of more efficient search and greater willingness to doubt the current idea. An earlier randomized trial involving 116 startups similarly found that hypothesis-driven testing improved the precision of continuation and pivot decisions.

A good gate does not maximize the number of projects approved. It improves the quality of the projects that consume the next round of time and money.

Evidence that deserves weight

The review should separate observations, interpretations, assumptions, and hopes. Problems arise when these categories are mixed together in a persuasive presentation.

“We sold five projects” is an observation. “Those projects demonstrate a repeatable market” is an interpretation. “We can reduce delivery time by half” is an assumption until tested. “Customers will buy the software version” is a hope unless customers have taken an economically meaningful action.

The following evidence table is a practical starting point. It is not an industry benchmark, and the amount of evidence required will vary with the offer’s price, risk, implementation difficulty, buying cycle, and cost of being wrong.

Decision areaStronger evidenceEvidence that often looks stronger than it isGate question
CustomerMultiple paid customers with similar situations, triggers, decision-makers, and constraintsA broad persona assembled from opinions or market reportsCan the team identify who should and should not buy this offer?
ProblemCustomers independently describe the same costly, urgent, or risky problemThe company uses the same problem statement in every pitchDoes the problem exist in customer language without prompting?
Buying behaviourContracts, deposits, paid pilots, renewals, expansions, or referrals tied to the problemPositive interviews, demo interest, survey intent, or verbal enthusiasmHas the customer accepted a real cost or commitment?
ResultCustomers sought and recognized a similar business resultEvery project has different success criteriaCan the offer promise one result without misleading the buyer?
DeliveryA stable sequence, known inputs, reusable assets, bounded exceptions, and measurable effortA checklist created after the fact that does not match actual workCan another qualified person follow the method successfully?
EconomicsCollected revenue, direct delivery cost, rework, support, discounts, write-offs, and senior timeContract value without delivery cost or unpaid effortDoes the work produce an acceptable contribution after its real delivery burden?
TransferabilitySome work can move from the founder to employees, software, automation, or documented decisionsThe founder can do the work quickly from experienceCan capacity increase without recreating the founder?
Market reachEvidence from more than one referral cluster, customer type, or relationship pathSeveral customers introduced by one highly connected advocateIs the observed demand broader than one network effect?

The last issue is especially important in a founder-led, referral-driven business.

Warm introductions reduce buyer uncertainty and can produce excellent early customers. They can also create a biased sample. Customers from the same network may share budgets, attitudes, industries, or trust in the founder. A product may appear broadly attractive when what has actually been validated is one relationship cluster.

Research on entrepreneurial experiments has found that a mismatch between the target market and the people included in early testing can have persistent effects on subsequent performance. The practical inference is that the gate review must assess not only how much evidence exists, but who generated it and how representative those customers are of the intended market.

This does not require cold-market proof before proceeding. It does require honesty about the boundary of what has been learned.

Use actual economics, not an ideal future model

For the gate review, calculate the economics of the work as it is performed now.

A useful internal calculation is:

Offer contribution = collected revenue − variable delivery costs

Variable delivery costs should include the costs that rise when another engagement is sold: delivery labour, contractors, engagement-specific tools, travel, payment fees, expected rework, and routine support. The company should decide deliberately how to value founder delivery time rather than treating it as free.

Then calculate:

Contribution margin ratio = offer contribution ÷ collected revenue

This is a management calculation for comparing offers, not a substitute for formal financial reporting. Its purpose is to expose whether the apparent opportunity depends on unpaid founder labour, inconsistent time recording, unusually easy customers, or costs that have been classified elsewhere.

The review should show the result as a range, not a single precise forecast. UK government cost-estimating guidance recommends documenting assumptions and exclusions, connecting cost to scope and risk, presenting ranges that reflect uncertainty, and explaining changes between decision gates. Those principles apply equally well to a small company evaluating an offer with limited historical data.

At minimum, test what happens when:

  • delivery takes longer than expected;
  • a non-founder employee performs the work;
  • the customer needs more support;
  • the company pays market rates for specialist labour;
  • discounts are required;
  • fewer customers buy than forecast;
  • onboarding or implementation becomes more complex.

An offer that works only in the most optimistic case is not ready to become the foundation for further investment.

How to run the review

The review should be short enough to force a decision and formal enough that the company cannot quietly reinterpret the result later.

The process below separates evidence preparation from the meeting itself.

flowchart LR
    A[Prepare evidence pack] --> B[Check completeness]
    B --> C[Review customer, result, delivery, and economics]
    C --> D{Enough evidence for one narrow offer?}
    D -->|Go| E[Approve offer to standardize]
    D -->|No-go| F[Pause, narrow, test, or stop]
    E --> G[Record scope, assumptions, owners, and limits]
    F --> G

The diagram shows a binary recorded decision. A no-go may lead to another test or a narrower offer, but it is still a no-go for the current proposal.

Prepare one evidence pack

The team should circulate the evidence before the meeting. It should include:

  • the proposed customer definition;
  • the problem in customer language;
  • the result customers bought;
  • the engagements included in the analysis;
  • sales history and buying triggers;
  • customer outcome evidence;
  • the actual delivery sequence;
  • delivery time and cost by engagement;
  • exceptions, rework, and support requirements;
  • the financial model and sensitivities;
  • founder-dependent activities;
  • evidence that contradicts the proposal;
  • unresolved assumptions;
  • the proposed offer to standardize;
  • the requested next investment.

The evidence pack should contain underlying records where practical: contracts, invoices, time records, interview notes, support records, project retrospectives, customer usage information, and financial calculations. A summary slide without traceable support is not enough.

Agree on the decision criteria before debating the answer

The people advocating for the offer should not be allowed to change the criteria after seeing weak evidence.

The criteria should be approved before the meeting and tailored to the risk of the next step. A modest investment to document a service can pass with less evidence than a decision to hire a product team, stop consulting work, or build a large software platform.

Formal review guidance commonly distinguishes between entrance criteria, success criteria, evidence, findings, and open actions. It also recommends identifying required participants and tracking review actions until they are resolved.

For this gate, the entrance criteria might require that customer, delivery, and financial records are complete enough to compare the relevant engagements. The success criteria should describe what must be true for one offer to proceed.

Put the right people in the room

A small company does not need a large steering committee, but it does need more than the founder selling the idea to everyone else.

The review should include:

  • one person accountable for the decision;
  • the person who owns sales evidence;
  • the person closest to delivery;
  • someone capable of challenging the financial assumptions;
  • someone who can represent customer outcomes and support burden;
  • a facilitator or reviewer who is willing to record disagreement.

One person may hold several roles. What matters is that sales optimism, delivery reality, customer value, and financial risk are all represented.

The founder should contribute evidence and judgment but should not be able to convert every missing fact into an assurance that “we can figure it out.” Where possible, have someone who did not create the proposal test the evidence and calculations.

Review one proposed offer at a time

The gate becomes weak when the team evaluates several possible customers, problems, and offers as one blended opportunity.

“Some version of this will work” is not an approval criterion.

The proposal should be narrow enough to state in one sentence:

For this type of customer, facing this problem under these conditions, we will provide this defined work to produce this result, within these scope and delivery boundaries, at this price or pricing basis.

If the team cannot write that sentence without several alternatives, the offer probably needs further narrowing before approval.

Force the meeting to produce a decision

The meeting should end with one of two outcomes:

Go: The company will invest in standardizing the proposed offer.

No-go: The company will not invest in standardizing the proposed offer at this time.

The no-go record should then state whether the team will narrow the proposal, gather specified evidence, continue selling it only as custom work, or stop pursuing it.

A “conditional go” often becomes a way to approve an idea while leaving important risks unresolved. When the missing condition is material to customer demand, delivery feasibility, or economics, record a no-go and define the test needed to return to the gate. Reserve post-approval actions for issues that do not undermine the basic decision.

Stage-gate systems can themselves become slow and bureaucratic. Research on combining structured gates with agile product development suggests that the discipline of decision points can coexist with iterative work, but review systems must remain adaptable and responsive to customer learning. The evidence for these hybrid approaches was initially limited, which is another reason not to treat any gate format as a universal formula.

How to record the decision and measure progress

The primary measure for this task is deliberately simple:

Was a go or no-go decision recorded?

That is a useful operating target, not an industry performance benchmark. Its value is governance: the company can identify what it decided, why it decided, and what investment was authorized or withheld.

A complete decision record should include:

Record fieldWhat it should contain
DecisionGo or no-go
Date and approverWho made the decision and when
Proposed offerCustomer, problem, result, scope, and pricing basis
Evidence reviewedThe records, analyses, and customer cases used
Evidence excludedCases omitted and the reason for omitting them
AssumptionsImportant claims not yet supported by direct evidence
Contradictory evidenceCustomers, projects, or numbers that do not fit the thesis
Decision rationaleWhy the evidence is sufficient or insufficient
Authorized next stepThe work, spending, and people approved
BoundariesCustomers, features, services, regions, or exceptions excluded
Open actionsOwner, due date, and required evidence
Revisit conditionWhat would trigger another review or reversal

The approved checklist is part of the evidence that the review occurred, but it is not the main result. The result is the decision and the clearly bounded offer—or the clearly stated reason not to proceed.

Supporting measures should then preserve the evidence base and reveal whether the gate judgment was correct. Useful measures include:

  • the share of relevant sales that fit the approved customer, problem, and result;
  • collected price and discounting;
  • delivery hours and elapsed delivery time;
  • variation in delivery cost between customers;
  • rework, support, and exception rates;
  • time required from the founder;
  • time until the customer receives the promised first result;
  • completion and outcome attainment;
  • repeat purchases, renewals, expansions, and qualified referrals;
  • contribution per engagement;
  • demand from outside the original referral network.

Do not require arbitrary universal thresholds. A high-priced, high-risk enterprise engagement needs different evidence than a low-cost service that can be tested and reversed quickly. The review should instead set company-specific ranges based on historical performance, target economics, customer risk, and the size of the next investment.

For example, a company may decide that reducing founder delivery work is essential even if the present margin is attractive. Another may accept founder involvement during the first version because the standardized method itself is the learning objective. The decision record should make that tradeoff explicit.

Failure modes and judgment calls

A polished package is mistaken for a repeatable offer

The company creates a name, price, landing page, proposal, and delivery checklist. None of these proves that customers share the same problem, value the same result, or can be served with consistent effort.

Service productization research treats commercial definition and delivery systemization as connected activities. Standardizing the description while leaving the underlying work unstable solves only part of the problem.

The team counts customers instead of examining similarity

Ten customers do not validate one offer if they bought ten different results. Conversely, a smaller set of closely comparable paid engagements may provide better initial evidence.

The correct question is not “How many customers have we served?” It is “How many relevant observations do we have for this exact customer-problem-result combination, and how independent are they?”

The founder’s sales ability is treated as market proof

Founder-led selling is useful because the founder can listen closely, adapt quickly, and explain unfamiliar work. It can also hide weaknesses in positioning, qualification, proof, and pricing.

The gate should identify which sales steps depend on personal reputation, relationships, improvisation, or the ability to change scope during the conversation. That does not invalidate the offer, but it limits claims about repeatability.

Customer satisfaction is confused with a measurable result

A customer may value responsiveness and expertise while failing to obtain the intended business outcome. The team must distinguish satisfaction with the relationship from evidence that the defined offer produced the promised result.

Where outcomes take a long time to appear, use leading evidence carefully. Record what has actually happened, what is expected to happen later, and what other factors could explain the change.

Average economics hide dangerous variation

An average engagement can look healthy while a few easy projects subsidize several difficult ones. Show the range and identify the causes of variation.

Particular attention should go to:

  • projects led entirely by the founder;
  • customers who arrived unusually well prepared;
  • work delivered with reusable assets that will not apply elsewhere;
  • unpaid changes;
  • emergency support;
  • unusually low contractor rates;
  • revenue recognized before the main delivery burden occurred.

The team uses the review to defend sunk costs

Time already spent does not make the next investment more attractive. The relevant question is whether current evidence justifies future effort.

The scientific-decision research is useful here because it treats termination as a legitimate learning result rather than proof that the earlier work was wasted. Better decision processes can increase the willingness to end weak ideas before further resources are committed.

The company standardizes too aggressively

Some customer work produces value precisely because it is adapted to context. The aim is not to eliminate judgment or force every customer into an identical process.

A sensible offer can have a stable core and controlled options. The gate should identify:

  • what must always be the same;
  • what can be selected from defined modules;
  • what can be adapted within limits;
  • what requires separate custom work;
  • what the company will refuse.

A successful example is treated as a formula

Basecamp provides a useful illustration. Its founders were operating a web design firm when they built a project-management tool to solve their own recurring delivery problem. Clients saw the tool and asked to use it. The company released it commercially, and its official history says the product was producing more revenue than the web design business roughly a year later, after which the company stopped providing web design services. Independent reporting also describes the product as emerging from the system the firm used to work with clients.

The lesson is not that every consultancy should turn an internal tool into software. The useful pattern is narrower:

  • a repeated operational problem was visible;
  • the solution was already being used in real work;
  • customers independently expressed demand;
  • customers paid for the separate product;
  • product revenue produced a clear basis for a later strategic shift.

Most companies will not see evidence that clean. The gate exists to determine whether the available evidence is strong enough for the next limited investment—not to demand that every offer resemble a famous software origin story.

At the end of the Stage 1 gate review, the company should be able to write one of two statements:

Go: We will standardize this defined offer for this customer because paid demand, customer results, delivery experience, and financial evidence support the decision. These assumptions and boundaries still apply.

Or:

No-go: We will not standardize this offer now because these material questions remain unresolved. We will run this specific test, narrow the proposal, continue it only as custom work, or stop.

Anything less specific leaves the company in the same position it was in before the meeting: busy, informed, but not decided.

Sources

Primary and official sources

  • NASA, “Program/Project Life Cycle,” including guidance on tailoring formal review criteria.
  • NASA Software Engineering Handbook, “Entrance and Exit Criteria.”
  • NASA Software Engineering Handbook, “Software Peer Reviews and Inspections—Checklist Criteria and Tracking.”
  • UK Government, Cost Estimating Guidance.
  • Basecamp, “Where We Came From.”
  • OpenStax, Principles of Accounting, Volume 2: Managerial Accounting, sections on contribution margin, break-even analysis, and sensitivity analysis.

Open research

  • Camuffo, Arnaldo, et al. “A Scientific Approach to Entrepreneurial Decision-Making: Large-Scale Replication and Extension.” Strategic Management Journal, 2024.
  • Camuffo, Arnaldo, Alessandro Cordova, Alfonso Gambardella, and Chiara Spina. “A Scientific Approach to Entrepreneurial Decision Making: Evidence from a Randomized Control Trial.” Management Science, 2020.
  • Cao, Ruiqing, Rembrand M. Koning, and Ramana Nanda. “Sampling Bias in Entrepreneurial Experiments.” NBER Working Paper, revised 2023.
  • Härkönen, Janne, Arto Tolonen, and Harri Haapasalo. “Service Productisation: Systematising and Defining an Offering.” Journal of Service Management, 2017.
  • Härkönen, Janne. “Exploring the Benefits of Service Productisation: Support for Business Processes.” Business Process Management Journal, 2021.
  • Schmidt, Jonas, and Tammo H. A. Bijmolt. “Accurately Measuring Willingness to Pay for Consumer Goods: A Meta-Analysis of the Hypothetical Bias.” Journal of the Academy of Marketing Science, 2020.
  • Cooper, Robert G., and Anita F. Sommer. “The Agile–Stage-Gate Hybrid Model: A Promising New Approach and a New Research Opportunity.” Journal of Product Innovation Management, 2016.

Public reporting

  • WIRED, “Work Smarter: 37signals,” on Basecamp’s origins in the company’s client-project system.