Sell the Pilot and Learn from the Buyer
Task
Run pilot sales and collect conversion evidence.
Summary
Use pilot sales to learn about the buyer, budget, objections, implementation burden, and decision process.
Run Pilot Sales to Learn Whether the Product Can Stand on Its Own
Task ID: S3-11
A pilot sale should test more than whether a customer likes the product. It should show whether a defined buyer will commit money, complete onboarding, reach a useful result, and continue without the company rebuilding the product or surrounding it with custom service. This article explains how to run that test and interpret the evidence honestly.
A pilot is a sales test, not a friendly beta
A company has extracted a useful piece of software from work it previously delivered as a service. Several customers have asked to try it. The founder can demonstrate it well, the early users are enthusiastic, and the product team receives a long list of feature requests.
None of that yet proves that the company has a standalone product business.
The unresolved question is whether a customer will buy the product on defined commercial terms and receive value without the company quietly recreating the old service around it. A pilot sales program is the controlled test of that question.
The operating principle is simple:
Run the pilot as a real sale with a limited learning period, not as an open-ended favor to a friendly customer.
That distinction matters because praise, survey answers, product use, pilot participation, and payment are different forms of evidence. Research on willingness to pay consistently finds that hypothetical statements overstate what people will pay when their own money is at stake. A meta-analysis covering 77 studies found that hypothetical willingness-to-pay estimates exceeded real willingness to pay by about 21% on average, although the size of the gap varied by method and product.
Payment is not perfect evidence either. A customer may pay because of a founder relationship, a bundled consulting agreement, a large discount, or a promise of custom development. But a purchase made under clearly defined conditions is stronger evidence than “we would probably buy this.”
This is why pilot sales belong after the repeated customer problem and first useful product have become reasonably clear. Earlier customer discovery helps a team understand problems, buyers, alternatives, and language. For example, the U.S. National Science Foundation’s I-Corps program emphasizes direct interviews with potential customers and industry stakeholders to evaluate possible product-market fit and inform the business model.
A sales pilot asks a harder question: Will the buyer act?
The answer must include not only whether the product was purchased, but also what work was required to produce that purchase. A product that sells only when the founder performs extensive analysis, imports data manually, redesigns workflows, and promises new features may still be a valuable offer. It is simply not yet evidence of a repeatable standalone product.
A rigorous pilot helps prevent teams from accepting false positives. Randomized research involving startups has found that teaching founders to state hypotheses, specify tests, and evaluate evidence systematically improves the quality of entrepreneurial decisions. The benefit is not merely learning faster; it is becoming less likely to continue with an idea because of selective or ambiguous feedback.
Decide what the pilot must prove
A weak pilot starts with a customer and a product demonstration. A strong pilot starts with written claims that the team is trying to test.
For a standalone product sold through demonstrations and design-partner relationships, the central claims usually concern the buyer, sale, product, onboarding, result, and support burden.
Before recruiting participants, the operating team should write down answers to these questions:
| Question | Claim being tested | Useful evidence |
|---|---|---|
| Who buys? | A specific customer type experiences the problem strongly enough to act. | Qualified opportunities by segment, role, trigger, and use case |
| Why now? | The buyer has an active reason to change rather than general interest. | Trigger event, deadline, cost of delay, executive priority |
| What is being purchased? | The product has a stable scope that can be described without inventing a new offer for each account. | Standard proposal, package, price, exclusions, implementation terms |
| Can the product be sold? | A repeatable demonstration and sales conversation can produce a commercial commitment. | Demo-to-proposal and proposal-to-pilot results, questions, objections |
| Can the customer start? | The buyer can provide the data, access, people, and decisions required for onboarding. | Time to start, onboarding completion, customer effort, internal dependencies |
| Does the customer receive value? | The product produces a result that the buyer recognizes as useful. | Agreed success measure, baseline, observed result, buyer confirmation |
| Can the company deliver it economically? | Support and implementation do not recreate an unpriced consulting project. | Employee hours, founder hours, custom work, support cases, infrastructure cost |
| Will the customer continue? | The customer will enter a normal paid relationship after the test. | Signed product agreement, paid invoice, renewal or continuation decision |
These claims should be narrow enough to fail.
For example, “buyers like the dashboard” is not useful because almost any positive comment can be treated as confirmation. “At least four of eight qualified operations leaders who complete the pilot will sign a twelve-month product agreement at no more than a 15% pilot discount within 30 days of the final review” is testable.
That does not make four of eight a universal benchmark. It makes the team’s expectation explicit.
A pilot target should state:
- the customer segment;
- the event that qualifies an account for the pilot;
- the product scope and minimum acceptable price;
- the numerator and denominator used for conversion;
- the deadline for making the purchase decision;
- the amount of implementation and support the company is willing to provide;
- the result the customer is expected to achieve;
- the conditions that count as a pass, a partial result, or a failure.
The design-partner label should not relax these requirements. A design partner may receive early access, additional influence, or a temporary discount in return for structured participation. It should not receive an unlimited right to redirect the product.
Early testers can also mislead a company when they do not resemble the intended market. Research on digital-product launches found that a mismatch between a product’s target customers and the people who participated in early beta testing had persistent effects on subsequent growth. The practical lesson is that pilot recruitment is part of the test design: convenient customers are not necessarily representative customers.
Recruit across the intended customer profile rather than filling the cohort solely with:
- long-standing consulting clients;
- personal contacts of the founder;
- unusually sophisticated users;
- companies willing to tolerate unfinished products;
- customers whose needs require exceptions unavailable to the wider market.
Existing relationships can be included, but they should be identified in the data. A customer who trusts the founder after five years of service is not equivalent to a buyer encountering the product for the first time.
Run a controlled commercial process
A pilot should follow the same core path for every qualified account. The conversations will vary, but the stages, definitions, and records should not.
flowchart LR
A[Qualified target account] --> B[Commercial discovery]
B --> C[Standard product demo]
C --> D{Real problem and buying intent?}
D -->|No| E[Disqualify and record reason]
D -->|Yes| F[Pilot proposal]
F --> G[Agree scope, price and success criteria]
G --> H[Run time-limited pilot]
H --> I[Review result and delivery effort]
I --> J{Purchase decision}
J -->|Buy| K[Standalone product contract]
J -->|No| L[Loss interview and reason]
J -->|Not now| M[No-decision reason and follow-up date]
The diagram shows the essential discipline: qualification comes before the pilot, commercial terms are discussed before delivery begins, and every completed pilot reaches a recorded decision.
The first stage is commercial discovery. The seller should establish that the account has the problem, the problem matters now, the proposed user can participate, and someone has authority to approve a purchase.
The discovery conversation should reconstruct the customer’s current situation rather than lead the buyer toward the product. Useful questions include:
- What is happening now that has made this problem important?
- How is the work completed today?
- What does the current approach cost in time, money, delay, risk, or missed opportunity?
- Who experiences the problem, and who owns the result?
- What other options are being considered, including doing nothing?
- What would have to be true for the organization to purchase?
- Who must approve the product, budget, security review, implementation, and ongoing use?
- When does the customer need a result?
These questions also expose false opportunities. A prospect may enjoy the demonstration but have no deadline, budget owner, implementation capacity, or reason to replace the current method.
The demonstration should then connect the product to the customer’s actual work. Research on business-to-business selling has found a gap between abstract seller descriptions of value and the concrete solutions buyers expect. In a pilot sale, the demo should therefore show the buyer’s use case, required inputs, operating workflow, expected output, and path to first value—not merely a tour of features.
The pilot proposal should be standardized enough to compare results across accounts. It should specify:
| Pilot term | What to define |
|---|---|
| Problem and use case | The particular job the customer is trying to complete |
| Users | Named roles and expected participation |
| Scope | Included product capabilities, data sources, integrations, and locations |
| Exclusions | Custom development, additional departments, unsupported data, consulting work |
| Duration | Start date, end date, and decision date |
| Customer obligations | Data, access, training attendance, internal owner, review participation |
| Company obligations | Setup, training, support response, product access, reporting |
| Success criteria | Observable result, baseline, target, and evidence source |
| Commercial terms | Pilot fee, credit toward purchase, post-pilot price, contract term |
| Decision process | Who decides, what documentation is required, and when the decision will occur |
Where possible, charge for the pilot. The fee need not equal the eventual subscription price, but it should require a genuine budget decision. A paid pilot tests seriousness, procurement, price sensitivity, and the customer’s ability to execute.
There are legitimate reasons to run an unpaid pilot. A new category may require learning before a price can be defended. A strategically important account may provide access to unusual data or a critical use case. A regulated buyer may require a proof of concept before procurement can approve a purchase.
The exception should be deliberate. Record the full cost, why the company is funding the test, what learning is expected, and the maximum investment allowed.
A revealing public example comes from Palantir Technologies. In its annual report for the year ended December 31, 2025, the company said it sometimes provides short-term pilot deployments at no or low cost and warned that these engagements may not convert into longer-term revenue. It also described a sales model that can require months of work and significant resources.
The lesson is not that unpaid pilots are always wrong. Palantir deliberately pursues large, complex opportunities and can support a high-touch model. The lesson is that a pilot has a cost whether or not the prospect pays. A smaller company testing a standalone product should not copy an enterprise deployment model accidentally.
Run early pilots sequentially when the team still expects major learning. Review the first few before recruiting a much larger cohort. Research on entrepreneurial experimentation suggests that sequential improvement based on customer feedback is particularly useful when demand and customer preferences remain poorly understood. Parallel testing becomes more useful when the product, customer, and decision criteria are stable enough to compare results across a larger sample.
Keep evidence that can survive review
A pilot pipeline is not a list of companies that expressed interest. It is a record of every account that entered the defined process, including the accounts that declined, stalled, or were disqualified.
At minimum, each opportunity record should contain:
| Evidence category | Fields to record |
|---|---|
| Account context | Segment, size, industry, existing relationship, source |
| Buyer | Primary user, champion, decision-maker, budget owner, technical approver |
| Problem | Current method, trigger, urgency, consequence, desired result |
| Qualification | Fit decision, reasons, date, person making the decision |
| Sales activity | Discovery date, demo date, proposal date, follow-ups |
| Offer | Package, list price, discount, pilot fee, contract term |
| Pilot delivery | Start, completion, user participation, data readiness, support time |
| Product evidence | Key actions completed, usage pattern, first-value milestone |
| Outcome | Converted, declined, no decision, extended, disqualified |
| Decision timing | Planned and actual decision dates |
| Feedback | Objections, alternatives considered, purchase criteria, reasons |
| Economics | Founder hours, employee hours, outside cost, custom work, gross contribution |
Maintain both structured fields and narrative notes.
Structured fields make comparisons possible. Narrative notes preserve the buyer’s language, internal politics, unexpected objections, and details that do not fit a drop-down menu.
The record should distinguish observations from interpretations:
| Type | Example |
|---|---|
| Observation | “The operations director used the product three times and exported two reports.” |
| Buyer statement | “The report removed about four hours of manual work.” |
| Commercial fact | “Procurement rejected the proposed annual prepayment.” |
| Interpretation | “The buyer may prefer monthly billing because the product is still unproven.” |
| Hypothesis | “Offering a quarterly commitment may improve conversion in this segment.” |
| Decision | “Test quarterly terms with the next five otherwise identical prospects.” |
This separation prevents a team from turning one buyer’s comment into a market conclusion.
Collect objections during the sale, but interview buyers again after the outcome. The post-decision conversation should reconstruct how the decision was made, not ask the buyer to grade the product.
Useful questions include:
- What originally caused you to examine a different way of doing this work?
- How did you evaluate the available options?
- What requirements became more important during the pilot?
- Where did the product help?
- Where did it create extra work?
- What concerned you but did not come up during the sales process?
- How did the price compare with the value and with alternatives?
- What would have needed to change for the decision to be different?
- Which parts of onboarding or support would be difficult to repeat?
- Why did the organization buy, decline, delay, or continue evaluating?
Interview wins, losses, and no-decisions. Wins identify the triggers, proof, and sales actions worth repeating. Losses identify gaps. No-decisions reveal whether the problem lacked urgency or the buying process was incomplete.
Do not rely solely on the salesperson’s loss code. The salesperson’s explanation is useful, but it is an interpretation from one participant in the sale. Whenever possible, someone who was not personally responsible for the deal should conduct the buyer interview. The purpose is learning, not evaluating the salesperson.
The team should review the pipeline weekly while pilots are active. The review should focus on evidence and decisions:
- Which opportunities entered or left each stage?
- Which success criteria have been observed?
- What work was performed outside the standard product and process?
- Which objections are repeating?
- Which customer characteristics appear in both wins and losses?
- What should change now?
- What must remain unchanged until the cohort is complete?
Changing everything after every conversation makes the test impossible to interpret. Ignoring repeated problems is equally damaging. Record each change to the offer, product, price, demo, qualification rules, and onboarding process, together with the date. This allows the team to compare cohorts rather than mixing several versions into one conversion number.
Read conversion rate in context
The primary measure is the pilot conversion rate:
That definition is suitable when the purpose is to test whether a completed pilot leads to a standalone sale. It should not be the only rate in the funnel.
Track at least these stage measures separately:
| Measure | What it helps diagnose |
|---|---|
| Qualified-account-to-demo rate | Ability to gain access to the intended buyer |
| Demo-to-pilot rate | Strength of problem, message, product presentation, and offer |
| Pilot-start rate | Procurement, data, technical, and customer-readiness friction |
| Pilot completion rate | Onboarding quality and customer participation |
| Pilot-to-paid conversion rate | Commercial proof after product experience |
| Time to first value | Speed at which the product produces a useful result |
| Time from completion to decision | Buying-process clarity and unresolved risk |
| Support hours per pilot | Delivery burden |
| Custom work per pilot | Whether the product is remaining standalone |
| Post-purchase use | Whether conversion is followed by adoption |
The denominator must be fixed before results are reviewed.
A team can make performance look better by excluding stalled pilots, counting only enthusiastic design partners, removing accounts that received the first version, or treating an extension as a conversion. This is denominator shopping.
Keep separate rates for:
- all qualified opportunities;
- pilots contractually started;
- pilots operationally started;
- pilots completed;
- completed pilots eligible for purchase;
- full-price conversions;
- discounted or exceptional conversions.
Suppose eight of twelve completed pilots buy. The observed conversion rate is 67%. That is encouraging, but twelve observations do not establish that the underlying rate is close to 67%.
Using a Wilson confidence interval, the approximate 95% interval for eight conversions from twelve trials is 39% to 86%. The range is wide because the sample is small. The U.S. National Institute of Standards and Technology recommends methods such as Wilson intervals for small binomial samples rather than relying on the familiar normal approximation.
The confidence interval should not become a reason to delay every decision until a large statistical study is possible. Early-stage companies rarely have hundreds of identical opportunities. It should prevent false precision.
Report:
Eight of twelve completed pilots converted to qualifying paid contracts, an observed rate of 67%. The approximate 95% Wilson interval is 39%–86%. Ten of the twelve accounts came from the same industry, and five had prior service relationships with the company.
That statement is more useful than “our pilot conversion rate is 67%.”
The appropriate target depends on the test design. It will differ with:
- customer fit and qualification standards;
- whether the pilot is free or paid;
- product price;
- contract length;
- number of approvers;
- implementation complexity;
- security and procurement requirements;
- strength of the customer’s trigger;
- existence of an incumbent;
- founder involvement;
- size of the discount;
- decision window;
- degree of custom work;
- maturity of the product and category.
Set a working target by combining commercial requirements with current evidence.
For example, if the company can support twelve pilots during the test period and needs six paid customers to justify the next investment, the working target is at least six qualifying conversions. That target is a decision threshold based on the company’s economics and capacity. It is not an external industry standard.
Conversion should also be interpreted alongside cost and repeatability. Six conversions may be inadequate if each requires 150 hours of founder work. Three may be valuable if each uses the same scope, pays the intended price, reaches value quickly, and opens a large repeatable segment.
A useful review table is:
| Result pattern | Likely interpretation | Next decision |
|---|---|---|
| High conversion, low custom work, acceptable support | Strong early evidence of a repeatable standalone sale | Expand carefully and test with less founder involvement |
| High conversion, heavy custom work | Buyers value the outcome, but the offer may still be a service-product hybrid | Standardize or price the service component before scaling |
| Low conversion, strong product use | Users may value the product while the buyer, price, contract, or business case is wrong | Revisit buyer and commercial design |
| Low completion | Onboarding, data readiness, customer selection, or product usability is blocking the test | Repair the start-to-value process before recruiting more pilots |
| High pilot acceptance, low paid conversion | Free access may be attracting interest without purchase intent | Tighten qualification, charge, or discuss post-pilot terms earlier |
| Conversion concentrated among existing clients | Relationship strength may be masking weak cold-market demand | Test with independent accounts |
| Mixed results by segment | The product may fit a narrower customer group | Separate the cohorts and focus the next test |
| Repeated no-decisions | The problem may lack urgency or the buying process may be incomplete | Strengthen trigger-based qualification or stop targeting that group |
Know what decision the evidence supports
Pilot sales are complete when the company can make a clearer decision, not when every participant has been made happy.
The final evidence package should contain:
- a stage-by-stage pilot pipeline;
- the exact qualification and conversion definitions;
- account-level sales and delivery records;
- original buyer notes and post-decision interviews;
- a coded list of objections and reasons for wins, losses, and no-decisions;
- product-use and first-value evidence;
- price, discount, and contract outcomes;
- implementation, support, custom-work, and founder time;
- conversion rates by relevant cohort;
- uncertainty and limitations;
- changes made during the test;
- a recommendation to proceed, narrow the market, change the offer, repeat the test, or stop.
The work should not be accepted as complete merely because a target was established. The target must be connected to a defined cohort, commercial terms, decision period, cost model, and next decision.
Before the company depends on the result, the following should be true:
The customers were representative enough. The cohort included plausible target buyers, not only friendly users who were unusually patient or committed.
The product was identifiable. Buyers understood what was included, what was excluded, and what they would receive after the pilot.
The purchase was real. Conversion meant a qualifying paid product agreement, not a verbal expression of interest, an indefinite pilot extension, or a product bundled invisibly into consulting fees.
The result was useful. The customer reached an observable first-value milestone rather than merely logging in.
The delivery burden was visible. Founder effort, support, data work, configuration, and custom development were measured rather than absorbed as free labor.
Non-conversions were understood. The team did not label every loss “timing” or “price” without reconstructing the buyer’s decision.
Uncertainty was acknowledged. Small samples, relationship effects, segment concentration, changing product versions, and discounts were reported openly.
The next test follows from the evidence. The team knows what it will keep stable, what it will change, and what result would reverse the current conclusion.
The strongest outcome is not necessarily a high percentage. It is a truthful explanation of who buys, under what conditions, at what price, after what experience, with how much company effort.
When that explanation is clear, leadership can decide whether the product is ready for a wider direct-sales motion, whether it needs another controlled cohort, whether a service component must be priced explicitly, or whether the standalone product thesis is not yet supported.
Sources
Primary and official sources
- U.S. National Science Foundation, “Information for Accepted National Teams” and “About Teams,” describing direct customer discovery and evaluation of commercial potential.
- U.S. Securities and Exchange Commission, Palantir Technologies annual report for the year ended December 31, 2025, describing pilot cost, conversion risk, and enterprise sales complexity.
- U.S. National Institute of Standards and Technology, guidance on confidence intervals for binomial proportions and small samples.
- Stanford Lean LaunchPad, course description of evidence-based entrepreneurship and direct testing with customers.
Open research
- Camuffo, Gambardella, Cordova, Spina, and related replication authors, research on applying a scientific approach to entrepreneurial decisions.
- Cao, Koning, and Nanda, “Sampling Bias in Entrepreneurial Experiments,” National Bureau of Economic Research.
- Schmidt and Bijmolt, “Accurately Measuring Willingness to Pay for Consumer Goods: A Meta-Analysis of the Hypothetical Bias.”
- Davis, Muzyrya, and Yin, “Experimentation Strategies and Entrepreneurial Innovation,” Stanford Institute for Economic Policy Research.
- Terho and co-authors, research on digital transformation and value-based business-to-business selling.
