Decision Two: Does the Offer Repeat?

Task

Run the Stage 2 gate review.

Summary

Decide whether sales and delivery results support investing in the first standalone product.

Decide Whether the Offer Is Truly Repeatable

Task ID: S2-15

A Stage 2 gate review decides whether a service has become a repeatable offer rather than a collection of similar custom projects. The review tests sales evidence, delivery consistency, customer results, economics, and founder dependence. Its purpose is not to reward activity. It is to record a defensible go, conditional-go, hold, or no-go decision before the company invests in further growth or software.

A named package is not yet a repeatable offer

Three customers have signed. Revenue is real. The website gives the service a name, the proposal has a standard cover page, and the team describes the work as a package.

Yet the first customer received a six-week engagement, the second negotiated a different result, and the third required the founder to rewrite the scope, approve the price, join every sales call, and solve delivery problems. Each project was profitable when viewed on an invoice, but nobody has counted the unpaid presales work, management time, rework, or exceptions.

The company has evidence of demand. It does not yet have evidence of a repeatable offer.

A productized service is a service with a defined customer, problem, promise, scope, price, buying process, delivery method, and expected result. Research on professional-service productization describes three related forms of work: specifying and standardizing the offer, making expert knowledge visible and understandable to buyers, and systemizing how the service is produced. Research on business-process standardization similarly treats consistency as something created through a defined reference process, not merely through a common label or template.

That does not mean every customer must receive an identical experience. Service modularity allows a company to combine standard components in approved ways, preserving some customer choice without redesigning the entire engagement. The important distinction is between controlled variation and unplanned customization. Controlled variation is priced, documented, and deliverable. Unplanned customization is discovered after the sale and absorbed through extra work, delay, or reduced margin.

The Stage 2 gate review exists to make that distinction visible. It asks:

Can the same offer be sold to similar customers and delivered to an acceptable standard without rebuilding the sale or the work each time?

The answer must come from completed sales and delivery records, not confidence, pipeline forecasts, or enthusiasm about the next stage.

What the gate is deciding

The gate is a resource-allocation decision. It determines whether the company should depend on the offer as a stable foundation for growth.

By this point, the company should already have done the underlying work: identified the customer and problem, written the promise, bounded the scope, set the price, developed the sales conversation, defined onboarding, documented delivery, tested customer results, and measured the basic economics. The gate review does not redo all of those tasks. It examines whether the combined system works in practice.

The decision should answer four questions:

  1. Can the offer be sold repeatedly? Similar customers understand it, accept roughly the same promise and commercial structure, and buy without a complete proposal redesign.
  2. Can it be delivered predictably? Delivery follows a known path, variation stays within an acceptable range, and exceptions do not routinely require senior intervention.
  3. Does the customer receive the promised result? The company can show evidence of the intended outcome or a credible leading indicator of it.
  4. Does the offer make operational and financial sense? Price, direct cost, capacity, rework, and support obligations leave enough room to operate and improve the service.

A recorded “go” is not a claim that uncertainty has disappeared. It means the remaining uncertainty is small enough, understood well enough, and affordable enough for the company to proceed.

A useful gate also allows more than a forced yes or no. Public-sector service assessments in the United Kingdom use green, amber, and red outcomes: green permits progression, amber permits progression while noncritical issues are corrected and tracked, and red keeps the service in its current phase until critical gaps are resolved. That model is useful because it distinguishes a manageable weakness from a reason not to proceed.

For a commercial offer, the equivalent choices are:

flowchart LR
    A[Evidence from sales and completed delivery] --> B[Stage 2 gate review]
    B --> C{Is the offer repeatable enough?}
    C -->|Yes| D[Go with stated limits]
    C -->|Minor gaps| E[Conditional go with owners and dates]
    C -->|Insufficient evidence| F[Hold and run specific tests]
    C -->|Weak model| G[No-go, redesign, or stop]

The diagram shows a review based on observed evidence. A go decision advances the existing offer. A conditional go advances it within explicit limits. A hold requests defined evidence rather than “more work” in general. A no-go stops further dependence on the present design.

Stopping or changing an offer should not be treated automatically as failure. Four randomized controlled trials involving 759 firms found that teaching a more scientific approach to entrepreneurial decisions increased the termination of ideas and produced a more selective pattern of strategic change. A 2026 randomized study of 553 technology-oriented startups likewise found that trained ventures shut down more often and earlier; among those that survived, the treated ventures subsequently showed stronger funding, employment, and revenue outcomes. These findings do not prove that every early stop improves performance, but they support the broader principle that disciplined termination can preserve resources for better opportunities.

The evidence that should be on the table

The review needs a compact evidence pack. A large presentation is less useful than a small set of reconciled records that show what was sold, what was delivered, what happened to the customer, and what the work actually cost.

Sales repeatability

Sales repeatability is not simply the number of signed contracts. The review should determine whether those contracts represent the same commercial proposition.

For each opportunity, record:

  • customer segment and buying trigger;
  • sales source, such as referral, targeted outbound, or an existing relationship;
  • stated problem and expected result;
  • offer version sold;
  • quoted and final price;
  • discounts, concessions, and special terms;
  • promised scope and timetable;
  • sales-cycle length;
  • people required to close the sale;
  • reason won, lost, delayed, or disqualified.

The review should compare the contracts and proposals themselves. A company can report five sales of the “same” offer while carrying five materially different obligations.

Useful measures include:

Standard-offer sale rate=Sales completed without material redesignTotal sales of the offer \text{Standard-offer sale rate} = \frac{\text{Sales completed without material redesign}} {\text{Total sales of the offer}}

A material redesign is a change that alters the promised result, delivery model, core scope, risk, price logic, or resource requirement. Minor wording changes, scheduling adjustments, or preapproved modules do not necessarily count.

Other useful signals are the median discount, range of sales-cycle lengths, share of deals requiring founder participation, and percentage of proposals using the standard agreement without substantial commercial exceptions.

Separate channel evidence rather than combining it indiscriminately. A referral from a trusted former customer may close quickly even when the sales message is unclear. Targeted outbound provides a harder test because the buyer has less inherited trust. A repeatable direct-sales motion should eventually work beyond the founder’s closest relationships, but the Stage 2 gate does not require every future channel to be proven.

Delivery predictability

Predictability means that the company can estimate the time, work, cost, dependencies, and likely result of an engagement within useful limits.

Do not rely only on average delivery time. Averages can hide severe variation. Record the median, the range, and a high-percentile result such as the eightieth percentile. If the median engagement takes 40 hours but one in five takes more than 100, staffing and pricing decisions based on the median will be unsafe.

Relevant measures include:

On-time delivery rate=Engagements completed by the committed dateCompleted engagements \text{On-time delivery rate} = \frac{\text{Engagements completed by the committed date}} {\text{Completed engagements}}
Standard-scope completion rate=Engagements completed without unplanned scope expansionCompleted engagements \text{Standard-scope completion rate} = \frac{\text{Engagements completed without unplanned scope expansion}} {\text{Completed engagements}}
Rework rate=Hours spent correcting or repeating workTotal delivery hours \text{Rework rate} = \frac{\text{Hours spent correcting or repeating work}} {\text{Total delivery hours}}
Founder dependency=Founder hours in sales, delivery, and escalationTotal hours used by the offer \text{Founder dependency} = \frac{\text{Founder hours in sales, delivery, and escalation}} {\text{Total hours used by the offer}}

Also examine onboarding time, waiting time caused by the customer, handoff failures, change orders, support work after completion, and the number of critical exceptions. Customer-caused delays should be visible rather than silently removed; they may reveal that the onboarding requirements or customer qualification rules are weak.

Process research defines capability by comparing the variation of a stable process with acceptable limits. Manufacturing capability indexes should not be copied mechanically into a young service business: common versions assume a sufficiently large sample, independent observations, process stability, and often a normal distribution. The transferable lesson is narrower and more useful—an average alone cannot show whether delivery is dependable. The team must understand variation and decide what range is acceptable.

Large-scale healthcare research provides evidence that consistency can matter in complex services, although it should not be generalized directly to consulting or software implementation. A study using about 35 million California inpatient stays found that greater procedural consistency was associated with lower cost per discharge, lower readmission rates, and less variation in readmission rates. The effects differed across operating units and medical conditions, reinforcing the point that standardization should be evaluated within comparable types of work rather than imposed uniformly.

Customer result and promise integrity

A service can be easy to sell and easy to deliver while failing to solve the customer’s problem.

Each engagement should therefore have an agreed success measure. The measure may be a business result, such as reduced processing time, a completed migration, a working control, or an approved operating plan. When the final business result takes longer to appear, use a leading measure that is closely connected to it, and state that limitation.

Useful records include:

  • baseline condition before the service;
  • expected result and date;
  • delivered result;
  • customer acceptance;
  • time to first useful value;
  • unresolved issues;
  • satisfaction or recommendation evidence;
  • repeat purchase, renewal, referral, or expansion;
  • reasons the result was not achieved.

The strongest evidence comes from comparable customers who reached the agreed outcome through substantially the same delivery model. Testimonials can support the case, but they do not replace operational records. A pleased customer may have received extraordinary founder attention or free work that cannot be repeated.

The review should also examine whether the sales promise matches the delivery team’s understanding. Repeated requests for “small extras” often show that the promise, exclusions, or qualification process is unclear.

Economics and capacity

Revenue proves that someone paid. It does not prove that the offer can support the company.

Calculate direct delivery margin using the costs that rise when another engagement is sold:

Direct delivery margin=Offer revenuedirect delivery costsOffer revenue \text{Direct delivery margin} = \frac{\text{Offer revenue} - \text{direct delivery costs}} {\text{Offer revenue}}

Direct costs can include employee delivery time at a realistic loaded rate, contractors, customer-specific software, travel, data, transaction charges, implementation resources, and post-delivery support. Presales work, project management, senior review, and rework should not disappear merely because employees are salaried.

The review should also show contribution after the recurring costs needed to sell and operate the offer. The purpose is not to force every early engagement to meet a mature margin target. Early work often includes learning costs. The purpose is to identify which costs should decline with repetition, which are structural, and which have merely been hidden.

Capacity should be expressed in operational terms:

  • engagements one delivery unit can complete per month;
  • work required from each role;
  • bottleneck role or step;
  • maximum safe work in progress;
  • time required to train another employee;
  • support burden after delivery;
  • effect of one additional sale on cash and staffing.

A service that sells faster than it can be delivered may create a backlog, missed promises, and cash pressure. Growth is not evidence of repeatability when every additional sale worsens the underlying economics.

How to run the review

The gate review should be a decision meeting, not a status meeting. Participants should receive the evidence pack early enough to examine it. Missing or unreliable data should be marked explicitly rather than replaced with confident estimates.

The core group normally includes the person responsible for the offer, the sales owner, the delivery owner, a finance or operations representative, and the executive making the investment decision. In a small company, one person may hold several roles, but the questions should still be separated. The founder should not be the only person presenting the evidence and judging it.

The evidence pack should contain completed engagements and closed sales up to a stated cutoff date. It should also include lost deals, delayed projects, dissatisfied customers, and costly exceptions. Removing inconvenient cases defeats the purpose of the gate.

A practical review can be organized around the following record:

AreaEvidence to examineDecision questionLikely owner
Customer and problemBuyer profiles, triggers, disqualifications, lost-deal reasonsAre substantially similar customers paying to solve the same problem?Sales or offer owner
Promise and scopeProposals, contracts, exclusions, change requestsIs the same result being promised under controlled terms?Offer owner
SalesPrices, discounts, sales cycles, channels, founder involvementCan the offer be sold without rebuilding the commercial case?Sales owner
Onboarding and deliveryPlans, actual hours, cycle times, delays, rework, exceptionsCan the team forecast and complete the work within acceptable limits?Delivery owner
Customer resultBaselines, acceptance, outcome measures, complaints, referralsAre customers receiving the result that justified the purchase?Customer or delivery owner
EconomicsRevenue, direct costs, contribution, cash timing, support burdenDoes each additional sale improve rather than weaken the business?Finance or operations
Dependence and capacityFounder hours, escalations, bottlenecks, training timeCan another qualified person sell or deliver the offer reliably?Executive sponsor
Data qualitySource, owner, definition, refresh date, missing recordsIs the evidence reliable enough to support the decision?Review chair

The meeting itself should follow a short sequence.

First, restate the exact offer being reviewed. Include the customer, trigger, result, standard scope, price structure, delivery method, and major exclusions. If participants describe different offers, the gate cannot proceed.

Second, present the evidence without interpreting it. Show the number and type of opportunities, completed engagements, delivery range, exceptions, outcomes, and economics. Distinguish observed facts from estimates.

Third, compare the evidence with the criteria agreed before the meeting. Late changes to the criteria should be documented. Moving the threshold downward because a launch is already scheduled is not an evidence-based decision.

Fourth, examine counterevidence. Ask which sale was least standard, which project was least profitable, which result was weakest, and what would happen if the founder were unavailable. This is not an attempt to reject the offer. It is a way to identify the actual limits of the model.

Fifth, make and record the decision before discussing broad future plans. The record should identify the outcome, evidence considered, unresolved assumptions, limits, owners, dates, and the condition for another review.

An independent or cross-functional challenge improves the process. The United Kingdom’s service-assessment guidance, for example, uses a panel that reviews a demonstration and questions the team about its decisions rather than relying solely on the producing team’s self-assessment.

However, rigor should not become inflexibility. Research based on 120 new-product projects found that repeatedly applying strict, objective gate criteria could increase project inflexibility and learning failure, particularly in turbulent technology settings. A young offer therefore needs clear decision criteria, but it also needs room to revise a hypothesis when credible new information appears.

How to set thresholds without false certainty

There is no universal number of customers that proves repeatability.

A low-priced, short, low-risk service may produce useful evidence after several independent sales and completions. A complex enterprise migration may require fewer engagements but deeper operational evidence because each engagement is expensive, long, and consequential. The required confidence should rise with the cost of a wrong decision.

Thresholds should reflect:

  • price and contractual exposure;
  • sales-cycle length;
  • delivery duration and complexity;
  • degree of customer variation;
  • regulatory, security, or safety consequences;
  • capital and hiring required after the gate;
  • cost of failure to the customer;
  • quality and independence of the evidence.

A company can use internal working thresholds, but they should be described as decision rules rather than industry benchmarks. For example:

  • at least several completed engagements from the intended customer group;
  • more than one source of business, where practical;
  • no recurring critical failure in onboarding or delivery;
  • most engagements completed within an agreed time-and-effort range;
  • customer outcomes measured for a defined share of completed work;
  • known, acceptable direct economics;
  • no undisclosed dependence on unpaid founder labour;
  • a standard scope and agreement that survived real negotiations;
  • documented reasons for every major exception.

The sample must also be examined for independence. Five engagements referred by one close partner, sold by the founder, and delivered by the same expert provide less evidence of a general repeatable motion than five engagements obtained across separate relationships and delivered by more than one qualified person.

Do not mix unlike customers merely to enlarge the sample. A service may be repeatable for a narrow segment and unpredictable elsewhere. A valid go decision can therefore be limited to a customer size, industry, region, use case, integration pattern, or delivery configuration.

Use ranges and distributions wherever possible:

  • median and eightieth-percentile delivery hours;
  • shortest, median, and longest sales cycle;
  • standard price and actual price range;
  • percentage of work completed within scope;
  • percentage of engagements requiring escalation;
  • expected and worst-observed direct margin;
  • result attainment by customer segment;
  • founder time per engagement.

Small samples still require judgment. A single outlier may represent bad luck, a preventable process failure, or evidence that the target segment contains materially different customer types. The review should investigate the cause rather than either ignoring the case or allowing it to dictate the entire decision.

Data quality belongs inside the decision. Every important measure should have a definition, owner, source, cutoff date, known bias, and readiness status. A precise-looking dashboard built from inconsistent time entries or incomplete cost records is weaker than a simple table that clearly states what is missing.

What public examples teach

GitLab’s professional-services catalogue illustrates one way to separate repeatable packages from broader advisory work. Its public service portfolio names defined offerings such as Implementation QuickStart, Migration QuickStart, health checks, training, and specialized migration services. The catalogue also describes the customer situation and broad purpose of each offer. This does not prove that every engagement is operationally identical or financially successful. It shows the commercial value of making standard service units visible while keeping different problems in distinct offers rather than presenting all expertise as one undefined consulting service.

The lesson for a Stage 2 review is not to copy those packages. It is to ask whether a buyer and delivery employee can tell where the standard offer ends, which modules are approved, and when a request belongs in a separate custom engagement.

MoviePass provides an extreme warning from a different type of business. It was a consumer subscription rather than a professional service, but it demonstrates why sales growth alone is not a gate criterion. In a June 2018 filing, its parent reported an average monthly cash deficit of about $25 million from September 2017 through May 2018. The May 2018 deficit was about $40 million, and the company expected at least $45 million for June, attributing the increase partly to rapid subscriber growth.

For the first nine months of 2018, the company reported approximately $205 million in revenue and a gross loss of about $219.4 million. It also disclosed that failure to pay processing partners had caused a service interruption in July 2018. The Federal Trade Commission later finalized a settlement over allegations that operators used deceptive tactics to prevent subscribers from using the service as advertised.

The direct lesson is simple: repeated buying is not proof of a sound offer when the company cannot afford or reliably fulfil what it has sold. A Stage 2 review must join demand evidence to delivery capacity, promise integrity, and unit economics.

Record a decision that can be acted on

The final output is not the meeting. It is a written gate decision.

A useful decision record begins with the exact offer version and evidence period. It then records one of four outcomes.

Go means the company has enough evidence that the defined offer sells repeatedly, can be delivered within acceptable limits, produces a credible customer result, and has workable economics. The decision should still state its boundaries—for example, approved customer segments, maximum concurrent engagements, authorized modules, standard price range, and situations that require a new review.

Conditional go means the main model is supported and the unresolved weaknesses are not critical. Each condition needs an owner, due date, measure, and operating limit. “Improve onboarding” is not a condition. “Reduce the eightieth-percentile onboarding time below ten business days before increasing simultaneous engagements beyond six” is actionable.

Hold means the evidence is insufficient or unreliable. The record should state precisely what must be learned. Examples include completing three engagements without founder delivery, testing the standard price with outbound prospects, measuring the customer result after 90 days, or tracking delivery time correctly for the next cohort.

No-go or redesign means the present offer should not be used as the basis for further growth. The company may narrow the customer, change the promise, remove costly scope, raise the price, create separate modules, alter delivery, or stop the offer. The decision should preserve what was learned rather than simply declaring that the work failed.

At minimum, the signed record should include:

  • offer name and version;
  • customer and use-case boundaries;
  • review date and evidence cutoff;
  • criteria used;
  • sales, delivery, customer-result, and financial evidence;
  • data limitations;
  • major exceptions and their causes;
  • assumptions that remain untested;
  • decision and rationale;
  • constraints on proceeding;
  • actions, owners, and due dates;
  • date or event that triggers reassessment.

The working target is therefore not a universal sales count, margin, or delivery-time benchmark. It is a recorded gate decision supported by evidence. The appropriate threshold depends on the company’s market, price, customer risk, delivery complexity, sample quality, and the investment that follows.

The review is complete when leadership can say more than “customers seem interested.” It should be possible to show who buys, what they buy, how consistently the company delivers it, what result customers receive, what each engagement consumes, where the model still varies, and why the remaining risk is acceptable—or is not.

Only then can the company responsibly depend on the offer, hire around it, increase sales activity, automate parts of delivery, or use its repeated work as evidence for a future standalone product.

Sources

Primary and official sources

  • UK Government Digital Service, “What Happens at a Service Assessment,” updated May 2024.
  • GitLab, “Professional Services,” current service catalogue.
  • National Institute of Standards and Technology, “What Is Process Capability?”
  • Helios and Matheson Analytics, Form 8-K, June 2018.
  • Helios and Matheson Analytics, Form 10-Q for the quarter ended September 30, 2018.
  • US Federal Trade Commission, “FTC Finalizes Settlement With Operators of MoviePass,” October 2021.

Open research

  • Kanika Goel, Wasana Bandara, and Guy Gable, “Conceptualizing Business Process Standardization: A Review and Synthesis,” Schmalenbach Journal of Business Research, 2023.
  • Elina Jaakkola, “Unraveling the Practices of Productization in Professional Service Firms,” Scandinavian Journal of Management, 2011.
  • Ahm Shamsuzzoha, Hannele Blomqvist, and Josu Takala, “Service Productisation Through Standardisation and Modularisation: An Exploratory Case Study,” International Journal of Sustainable Engineering, 2023.
  • Rajesh Sethi and Zafar Iqbal, “Stage-Gate Controls, Learning Failure, and Adverse Effect on Novel New Products,” Journal of Marketing, 2008.
  • Arnaldo Camuffo et al., “A Scientific Approach to Entrepreneurial Decision-Making: Large-Scale Replication and Extension,” Strategic Management Journal, 2024.
  • Esther Bailey et al., “Learning to Quit? A Multi-Year, Multi-Site Field Experiment With Innovation-Driven Entrepreneurs,” NBER Working Paper, January 2026.
  • Anand Bhatia and Jayashankar M. Swaminathan, “Measuring Consistency in Service Delivery: Examining the Effect of Process Standardization on Hospital Performance,” Journal of Service Research, 2026.