Build the Core Value Path
Task
Build a demo, prototype, or MVP that shows the core value path.
Summary
Create a demo, prototype, or minimum viable product that makes the central outcome visible.
Build the Smallest Product That Proves the Customer Result
Task ID: S3-05
A useful demo, prototype, or minimum viable product does more than show features. It lets a suitable customer experience the shortest path from a real problem to a result they can judge. This article explains how to define that path, choose the right level of product fidelity, test it credibly, and decide whether it is ready for design partners.
A convincing demo must survive customer use
A product team can spend weeks preparing a smooth demonstration and still learn very little. The presenter knows which buttons to press, the sample data has been selected to produce an attractive result, and an engineer may be quietly fixing problems behind the scenes. The meeting goes well. The customer says the product looks interesting. Yet nobody has shown that the customer can use it, obtain a worthwhile result, or fit it into real work.
That is the problem this task solves.
The operating principle is simple:
Build the smallest credible version of the product that allows a suitable customer to travel from a recognizable problem to an observable result.
“Smallest” limits cost and prevents premature expansion. “Credible” means the experience is realistic enough for the customer to make a meaningful judgment. “Observable result” means the test ends with evidence, not merely praise.
Research on minimum viable products supports treating the MVP as a set of deliberate tradeoffs rather than as a universally defined feature threshold. The product can vary in functionality, technical completeness, presentation, distribution, and the type of evidence it is intended to produce. Reducing one dimension may save time, but it can also weaken the test if that dimension is essential to the customer’s judgment.
That distinction matters at this stage of company development. The company is not merely testing whether people understand an idea. It is trying to determine whether one useful part of previously delivered work can become a standalone product. The test therefore has to address more than visual appeal:
- Can the intended customer recognize the problem being solved?
- Can the customer complete the essential workflow?
- Does the product produce a result the customer considers useful?
- Can the company provide that result without rebuilding the product for each customer?
- Is the remaining human work deliberate and limited, or is custom service work hiding behind the interface?
A demo can answer some of these questions. A prototype can answer others. A working minimum viable product may be necessary when the uncertainty concerns real data, integrations, reliability, repeated use, or operational delivery. The correct artifact is the least expensive one that still exposes the important risk.
Define the core value path before choosing what to build
A feature list is not a value path.
A feature list might include authentication, dashboards, alerts, exports, permissions, integrations, and artificial intelligence. A value path describes what happens to the customer:
- A situation creates a reason to act.
- The customer supplies information or connects a source.
- The product performs the essential work.
- The customer receives a result.
- The customer can verify, use, or act on that result.
This sequence should begin where the customer’s problem becomes active, not where the software interface begins. It should end when the customer has received evidence of value, not when the software finishes processing.
Government Digital Service guidance makes the same distinction in service design: teams are expected to understand the user’s full context, test prototypes with likely users, and consider the complete journey, including support and offline steps. Testing only the screens owned by the product team can miss the work required before and after the interface.
A useful core value path can be written as a single concrete sentence:
When [customer] encounters [situation], they can give the product [required input], receive [specific output] within [relevant time], and use it to [customer action or result].
For example:
When an operations manager receives a large exception report, they can upload it, identify the records requiring attention, assign the next action, and export a reviewed list before the daily operations meeting.
That statement is more useful than “an artificial-intelligence-powered operations platform.” It identifies the customer, trigger, input, processing step, result, time constraint, and next action. Each can be tested.
flowchart LR
A[Customer encounters a problem] --> B[Provides the minimum input]
B --> C[Product performs the essential work]
C --> D[Customer receives a result]
D --> E{Can the customer verify and use it?}
E -->|Yes| F[Core value demonstrated]
E -->|No| G[Revise the product or the value claim]
The diagram shows the minimum end-to-end test. A product is not ready merely because the central algorithm or interface works. The customer must be able to reach the result and decide whether it is useful.
Before building, the team should identify the assumptions contained in this path. At minimum, examine four types:
| Assumption | Practical question | Evidence the prototype should produce |
|---|---|---|
| Customer problem | Does this situation occur often enough and matter enough? | Customers recognize recent examples and are willing to spend time addressing them |
| Workflow | Can the customer provide the required input and complete the steps? | Observed task completion with representative inputs |
| Product result | Is the output accurate, timely, understandable, and useful? | The customer can assess the output and use it in a real decision or task |
| Delivery model | Can the company produce the result repeatedly without extensive custom work? | Known setup time, support effort, manual work, and technical constraints |
The team should then name the most dangerous assumption: the one that would make further development wasteful if it proved false. An experiment should be designed around that assumption rather than around the feature the team is most excited to show.
This is more than a tidy planning exercise. A study of two software startups found that MVPs and business hypotheses were often only loosely connected: some important hypotheses were not tested by an MVP, while some product work had no identifiable hypothesis behind it. The practical lesson is that every substantial part of the prototype should have a reason to exist: it either enables the value path, makes the test credible, or measures an important uncertainty.
A larger set of randomized trials involving 759 firms also found that teaching entrepreneurs a more scientific approach to decisions increased disciplined idea termination and produced more focused strategic changes rather than either no change or repeated uncontrolled pivoting. The point of an MVP is therefore not to defend the current idea. It is to make continuation, revision, or termination more evidence-based.
Match the prototype to the uncertainty
“Build an MVP” is incomplete advice because different artifacts answer different questions.
A product team should choose fidelity according to the decision it needs to make. A rough artifact is often sufficient to test whether customers understand a workflow. It may be useless for testing processing speed, output quality, technical feasibility, or trust.
The following comparison provides a practical starting point:
| Artifact | Best suited to testing | What it cannot establish by itself |
|---|---|---|
| Storyboard, sample report, or demonstration video | Whether the problem and promised result are understandable and interesting | Whether customers can operate the product, whether the technology works, or whether value persists in real use |
| Clickable interface prototype | Navigation, terminology, sequence, comprehension, and basic usability | Real processing, integration behavior, output quality, reliability, or delivery cost |
| Concierge prototype | Whether the customer values the outcome when humans perform much of the work manually | Whether the process can be automated economically or supported at scale |
| “Wizard of Oz” prototype | Whether the visible experience works while selected functions are performed manually behind the interface | Whether hidden functions are technically feasible, fast, reliable, or affordable |
| Functional vertical slice | Whether one narrow workflow works through the interface, business logic, data, and output | Whether the wider product, edge cases, or operating model are ready |
| Limited working MVP | Whether customers can use the product in a realistic environment, obtain value, and return to it | Long-term retention, broad market demand, mature reliability, or full unit economics |
The appropriate level of polish also depends on what the customer is being asked to judge. Research involving entrepreneurs from 12 startups found that even early products need enough approachability and professionalism to communicate the intended value. An unnecessarily confusing or careless experience can cause customers to reject the artifact rather than the underlying idea.
That does not mean every prototype should look finished. In controlled usability research, low- and high-fidelity prototypes have sometimes uncovered similar interface problems, suggesting that expensive visual refinement is not always necessary for learning about task flow. Those findings should not be stretched too far: fidelity becomes important when realism itself affects trust, performance, technical behavior, or the result being evaluated.
A sensible rule is:
Increase fidelity only where a lower-fidelity version would allow the team or customer to reach the wrong conclusion.
A clickable prototype may be enough to test whether users understand an approval workflow. It will not prove that an analysis engine can process a customer’s data within ten minutes. A manually generated recommendation may reveal whether the recommendation is valuable. It will not prove that the company can produce thousands of recommendations profitably.
For products derived from service work, this boundary deserves special attention. Manual work is acceptable in an early test when it is used knowingly to isolate demand or value. It becomes misleading when the team presents a labour-intensive service as though the software produced the result independently.
The prototype record should therefore identify:
- what is fully functional;
- what is simulated;
- what is performed manually;
- what data is real, synthetic, or preselected;
- which scenarios are supported;
- which scenarios are deliberately excluded;
- what the customer is and is not being asked to evaluate.
These disclosures do not weaken a design-partner conversation. They make the evidence interpretable.
Build an evaluation environment, not a stage performance
A presentation is controlled by the seller. An evaluation environment allows the customer’s behavior to create evidence.
The environment may be hosted software, a private tenant, a local build, an interactive prototype, or a guided sandbox. Regardless of form, it should contain enough context for a representative customer to attempt the core value path.
Use a realistic starting condition
The evaluation should begin with a situation the customer recognizes. Give the evaluator a goal rather than a sequence of instructions.
Weak prompt:
Click “New Project,” select the second option, upload this file, and press “Analyze.”
Stronger prompt:
You have received this file before a meeting. Find the records that require action and prepare the information you would give the team.
The first tests whether the participant can follow the presenter’s directions. The second tests whether the product supports the customer’s work.
Representative data matters for the same reason. Perfectly clean sample data may conceal setup effort, ambiguity, missing information, error handling, and output-quality problems. Real customer data can improve realism, but it also raises privacy, security, contractual, and operational risks. The Federal Trade Commission advises businesses to collect only information they need, limit access, understand where data is stored and transmitted, and retain sensitive data only for a legitimate purpose. For early evaluations, synthetic or carefully de-identified data is often the safer starting point unless real data is essential to the hypothesis.
Let the customer control meaningful parts of the path
A narrated demonstration is useful for introducing the idea. It should not be the only test when the goal is readiness for design partners.
The evaluator should perform the actions that could expose misunderstanding or friction. That may include supplying an input, configuring a rule, interpreting a result, correcting an error, sharing an output, or deciding what to do next.
The team should observe before explaining. Immediate coaching can turn a failed workflow into an apparently successful demonstration. Assistance should be recorded so the team can distinguish unaided completion from completion that required support.
Make the result inspectable
The output should be something the customer can judge. Depending on the product, that may be:
- a completed task;
- a recommendation with supporting evidence;
- an identified error;
- a generated document;
- a reconciled dataset;
- an automated action;
- time saved on a known process;
- a decision made with greater confidence.
The result should not be defined as “the customer reached the final screen.” The customer must be able to say whether the output is correct enough, fast enough, understandable enough, and useful enough for the intended work.
Instrument the path
At this stage, analytics do not need to be elaborate, but the team should not depend on memory. Record the events required to reconstruct what happened:
- whether the evaluator started and completed the path;
- where they paused, failed, or requested help;
- elapsed time to the first meaningful result;
- errors and recovery attempts;
- inputs used and outputs generated;
- manual work performed by the company;
- support time before, during, and after the session;
- the customer’s assessment of the result;
- the next commitment the customer is willing to make.
Google’s HEART framework separates user-experience measures such as adoption, engagement, retention, satisfaction, and task success, and recommends mapping product goals to observable signals and metrics. For an early MVP, task success is usually more informative than broad engagement because the product may not yet have enough users or elapsed time to measure retention reliably.
Test the whole condensed journey
A prototype does not need to contain the eventual whole product. It does need to represent the smallest complete journey through which value can be judged.
This includes essential entry and exit steps. How does the customer know when to use the product? How do they get the required information into it? Where does the result go? What happens if the input is incomplete? Who acts on the output? What support is required?
The boundary should be narrow, but the path through that boundary should be complete.
Public examples show the difference between demonstration and evidence
Dropbox’s early demonstration remains useful because it showed a difficult-to-explain experience rather than describing a list of storage features. On April 4, 2007, Drew Houston posted “My YC app: Dropbox — Throw away your USB drive” to Hacker News. The discussion did not produce only encouragement. Potential users questioned installation restrictions, offline access, the USB-drive comparison, and monetization. Houston responded by explaining the importance of a local folder, background synchronization, and disconnected access.
The lesson is not that every startup should create a video. The video was appropriate because the early question was whether people understood and wanted the proposed file-synchronization experience. The public response generated concrete objections that could guide subsequent work. It did not, by itself, prove reliability, profitable delivery, long-term retention, or enterprise readiness.
Public-service assessments provide a second useful lesson: a narrow prototype can be either appropriately focused or dangerously incomplete.
In one assessment, a Prevent e-learning team presented only part of one user journey. The assessors asked it to test a condensed end-to-end experience, including how users found the service, moved through it, and received proof of completion, even if all training content was not yet developed.
A Civil Service Learning prototype had demonstrated that users could complete one new-starter flow, but it covered too little of the wider service and did not represent important needs involving learning records and managers. The service was not allowed to proceed to private beta on the strength of the polished narrow slice.
By contrast, the United Kingdom’s Waste Tracking service was judged ready for private beta after the team had tested multiple ideas, iterated prototypes using user and stakeholder feedback, covered different user journeys, and begun defining measures including adoption, completion, satisfaction, and transaction cost. The assessment did not treat the prototype’s existence as proof. It examined the evidence and operating preparation surrounding it.
Together, these cases show the distinction:
- A focused prototype removes unnecessary scope while preserving the complete value path.
- An incomplete prototype removes steps that are necessary to determine whether value can actually be delivered.
The correct boundary is defined by the customer result, not by the limits of the current interface.
Measure demo readiness and decide whether design partners can start
“Demo readiness” should not be reduced to whether the software launches successfully before a sales meeting. It is better treated as a structured judgment supported by several measures.
For this purpose, a design partner is a suitable early customer that agrees to evaluate the product in a realistic setting, provide structured evidence, and work within clearly stated limits. A design partner is not a substitute engineering department, an unlimited source of feature requests, or a customer receiving an undisclosed custom service.
The following scorecard can support the readiness decision:
| Readiness area | Evidence to review | Warning sign |
|---|---|---|
| Customer fit | Evaluators match the intended user, buyer, workflow, and problem | Feedback comes mainly from friends, investors, colleagues, or customers with unrelated needs |
| Core-path completion | Suitable users can reach the result with limited assistance | The founder must narrate, configure, or rescue every attempt |
| Result quality | Customers can assess the output and explain how they would use it | Reactions focus on appearance while the result remains hypothetical |
| Time to value | Time from starting the evaluation to receiving a meaningful result | Setup takes longer than the problem warrants |
| Technical stability | No unresolved failure prevents evaluation of the core path | The team dismisses critical failures as “only demo issues” |
| Delivery effort | Setup, support, manual processing, and customization are recorded | Labour is hidden or varies unpredictably by customer |
| Trust and data handling | Data boundaries, access, deletion, limitations, and responsibility are clear | Sensitive data is collected casually or customer expectations exceed actual controls |
| Learning design | Hypotheses, observations, metrics, and decision rules are documented | The team collects comments without knowing what would trigger a change |
| Commercial signal | The customer makes a meaningful next commitment | Interest consists only of compliments or a request to “keep me updated” |
A company is reasonably ready for design partners when the following conditions are true:
- A representative evaluator can complete the narrow end-to-end path.
- The product produces a result the evaluator can judge.
- The team knows which parts are functional, simulated, or manual.
- Known limitations are acceptable for the selected partners.
- No unresolved security, privacy, safety, regulatory, or reliability issue makes the test irresponsible.
- The company can onboard and support the initial cohort without creating a different product for every participant.
- The team has decided what evidence will cause it to proceed, revise, or stop.
These are decision criteria, not a universal industry benchmark. Appropriate readiness differs by product risk and customer environment. A scheduling tool evaluated with synthetic data may be ready much earlier than software that makes medical recommendations, moves money, controls industrial equipment, or processes regulated information.
The number of design partners should also follow the learning problem rather than a standard quota. The first group should be small enough for close observation and fast iteration, but varied enough to expose the important customer, workflow, and data differences. Adding partners faster than the team can analyze evidence merely creates support load and competing feature requests.
The most useful measures during the first evaluations are usually:
- Core task completion: Did the customer reach the intended result?
- Time to first value: How long did the complete path take?
- Assistance required: How much guidance or intervention was necessary?
- Critical failure rate: How often did a problem prevent evaluation?
- Result acceptance: Did the customer judge the output usable for its intended purpose?
- Setup and support effort: How much company time did each evaluation require?
- Customization effort: What had to change for this customer alone?
- Commitment: Did the customer agree to a pilot, provide data, assign employees, accept defined terms, or pay?
A verbal statement that the product is “cool” is weak evidence. A customer who supplies representative data, schedules users, agrees on success criteria, and accepts a defined commercial or pilot arrangement has made a stronger commitment.
Repeated use, retention, renewal, and expansion become more important later. They should not be claimed from a one-session demo. An early evaluation can establish that the value path is credible enough to test in ongoing work; it cannot establish a durable subscription business.
Avoid the common ways a demo looks finished when it is not
The most expensive errors at this stage often come from misreading the evidence.
The team builds a horizontal layer of many incomplete features. Customers can tour the product but cannot finish one valuable job. A narrower vertical slice—from input through processing to result—is usually more informative.
The demonstration tests presentation skill rather than customer success. The founder performs every action while the customer watches. This can establish comprehension or initial interest, but not usability, onboarding, or independent value.
The product stops before the result. It gathers information, displays a dashboard, or generates an intermediate artifact, but the customer must still perform the most important work manually. The prototype has demonstrated activity rather than value.
Hidden service work makes the software appear more capable than it is. Concierge work can be a valid experiment, but only when it is identified, measured, and tied to a later automation decision. Otherwise the company may carry the economics of custom delivery into a product price.
The company treats every request as product validation. Early customers often ask for capabilities that fit their own processes. The team should separate requests that strengthen the shared value path from requests that create a private branch for one customer. Research with a broad range of likely users helps prevent a single enthusiastic participant from defining the market.
The team changes the prototype without preserving the hypothesis. Constant iteration can feel productive while making evidence impossible to compare. Record the assumption, version, participant, result, and resulting decision for each meaningful test.
The test contains no failure rule. Without a pre-agreed threshold or decision standard, every outcome can be reinterpreted positively. The company should decide in advance which failures would require another iteration, a different customer segment, a different value claim, or termination.
The team confuses technical feasibility with willingness to buy. A functioning workflow proves that something can be built. It does not prove that the problem is important enough to fund, that the buyer has authority and budget, or that ongoing use justifies the cost.
The company invites design partners before it can constrain the relationship. Design partners need a defined use case, evaluation period, supported environment, responsibility split, data rules, feedback process, and change policy. Without those boundaries, each partnership can become a custom project.
The company waits for production perfection. The opposite error is building mature architecture, comprehensive permissions, broad integrations, and every edge case before testing the central value claim. The MVP literature emphasizes that early products require explicit choices among competing forms of completeness; attempting to minimize every risk at once defeats the purpose of an early experiment.
The final result of this task should therefore be more than a URL or a scripted presentation. It should be an evaluable product environment accompanied by a clear value path, known limitations, representative scenarios, basic instrumentation, a test protocol, and decision rules.
When that evidence exists, the company can make a defensible next decision:
- Proceed when suitable customers can complete the path, value the result, and enter a bounded design-partner engagement.
- Revise when the problem remains important but the workflow, output, customer, or delivery model fails the test.
- Stop or redirect when the core result is not valuable enough, cannot be delivered credibly, or requires so much custom work that it is not becoming a standalone product.
The purpose of the demo, prototype, or MVP is not to prove that the team has finished building. It is to make that decision clearer before the company depends on the product for growth.
Sources
Primary and official sources
- United Kingdom Government Digital Service, guidance on understanding user needs and conducting research during alpha development.
- United Kingdom Government Digital Service, Waste Tracking service alpha assessment.
- United Kingdom Government Digital Service, Prevent E-Learning alpha assessment.
- United Kingdom Government Digital Service, Civil Service Learning alpha assessment.
- Federal Trade Commission, guidance on protecting personal information and reducing unnecessary data collection.
- Google Research, Measuring the User Experience on a Large Scale: User-Centered Metrics for Web Applications.
- Drew Houston’s original Dropbox launch discussion on Hacker News.
Open research
- Stevenson, Burnell, and Fisher, The Minimum Viable Product: Theory and Practice, 2024.
- Camuffo and colleagues, A Scientific Approach to Entrepreneurial Decision-Making: Large-Scale Replication and Extension, 2024.
- Khanna, Nguyen-Duc, and Wang, From MVPs to Pivots: A Hypothesis-Driven Journey of Two Software Startups, 2018.
- Hokkanen, Kuusinen, and Väänänen, Minimum Viable User Experience: A Framework for Supporting Product Design in Startups, 2016.
- Catani and Biers, Usability Evaluation and Prototype Fidelity, 1998; Walker, Takayama, and Landay, High-Fidelity or Low-Fidelity, Paper or Computer?, 2002.
