Recruit Design Partners from Earned Trust

Task

Recruit design partners from best-fit consulting/productized-service customers.

Summary

Use the strongest existing relationships to test the product with customers who understand the problem.

Recruit Design Partners Who Can Become Paying Customers

Task ID: S3-06

A design-partner program should do more than collect opinions. It should recruit a small group of best-fit customers who will use the early product in a real workflow, commit time and data, judge it against agreed results, and reveal whether the product can be sold and delivered without becoming another custom consulting project.

The real problem is not finding friendly customers

A former consulting customer says the product idea sounds useful. Another offers to “help however we can.” A third asks for a demo, then sends a list of features unique to its organization. These conversations feel like progress, but none proves that a customer will adopt the product, change a workflow, allocate staff, accept a bounded scope, or pay for continued use.

The operating principle is simple: recruit design partners for the quality of evidence they can produce, not for the warmth of the relationship. A useful design partner is an early customer that can run the product against a real problem and make explicit commitments in return for early access, focused support, and some influence over how the product develops. A standard design-partner agreement therefore treats the relationship as two-sided: the customer receives product access, while the provider receives scheduled participation and usable feedback.

This work belongs after a company has learned from customer service work and built a first useful standalone product. The question is no longer merely, “Does this problem exist?” The question is, “Can this product solve enough of the problem, for the right customer, without rebuilding the old service around every account?”

Recruiting from existing consulting or productized-service customers is valuable because the company already has evidence about their workflows, buying behavior, results, delivery effort, and working relationship. Basecamp offers a clear historical example of the underlying pattern: it began inside a design firm that needed a better way to manage client work. The repeated service problem supplied the product opportunity; external product adoption still had to establish that the tool could stand on its own.

No formal dependency is listed for this task, but there are practical readiness conditions. The company needs a product that can demonstrate one end-to-end job, a reasonably clear target customer, an owner who can run the program, and enough product and support capacity to serve the cohort. Without those conditions, recruitment produces promises the team cannot test.

Choose partners for evidence, not enthusiasm

Research on “lead users” explains why some customers are more useful than an average sample. Lead users experience an important need earlier or more intensely than the broader market and often have already developed workarounds or solution ideas. Studies of industrial product development found that carefully identified lead users could contribute unusually useful information about emerging needs and product concepts.

That does not mean the company should simply choose its most demanding or innovative customer. Customer participation is not uniformly beneficial. A meta-analysis found that customer participation tended to help during ideation and launch, but intensive participation during development could slow time to market and harm financial performance. The practical lesson is to use customers to expose the workflow, test value, and challenge assumptions—without handing one customer control of the roadmap.

Start with the service-customer record. Identify accounts that bought the repeated work, reached a meaningful result, moved through the sale without extraordinary persuasion, and did not require unusual delivery exceptions. Add what the team knows about the buying group, the trigger that created urgency, the current workaround, the data or systems involved, support demand, margin, retention, and willingness to act as a reference. This turns account history into a ranked candidate list rather than a memory contest among sales, delivery, and product leaders.

A practical scorecard can use a zero-to-three scale for each criterion. The score is a discussion aid, not a claim of mathematical precision.

CriterionStrong evidenceWarning sign
Problem fitThe customer repeatedly experiences the exact job the product is built to handleThe customer likes the concept but has no urgent or recurring use
Product fitThe current product can complete one valuable workflow with limited configurationThe pilot depends on major new features, integrations, or manual work
Learning valueReal users, real data, observable work, and a measurable starting point are availableFeedback will come only from an executive sponsor or innovation team
ParticipationA named owner can provide users, data, access, and regular review timeNobody owns the pilot or meetings are “as available”
Commercial fitA budget owner, buying process, and plausible path to a paid subscription are knownThe customer can experiment but has no route to purchase
RepeatabilityThe use case resembles other target accounts and can become a standard setupSuccess would depend on customer-specific rules or consulting
Relationship qualityThe teams can discuss failures directly and make timely decisionsThe relationship is politically sensitive or driven by obligation
Operational fitSecurity, privacy, procurement, and implementation needs fit the pilot windowBasic access or approval will take longer than the proposed test

Two selection errors are common. The first is choosing only the largest or most prestigious customers. They may bring slow procurement, complex security requirements, and edge-case workflows that are poor tests of a simple first product. The second is choosing only easy, friendly customers. They may give generous feedback but weak commercial evidence. The better cohort combines strong problem fit with enough operational readiness to start and enough commercial realism to make the result meaningful.

The U.S. National Science Foundation’s I-Corps program separates customer discovery from product building and requires teams to conduct extensive interviews to evaluate product-market fit and the business model. In 2026, its national program expects at least 100 interviews over seven weeks. That number is not a design-partner target; it illustrates a useful distinction: broad discovery reduces the risk of building around a few familiar accounts, while the design-partner cohort performs a deeper operational test with a much smaller group.

Turn outreach into a two-sided pilot offer

Recruitment works best as a direct, demo-led sale rather than a general invitation to “co-create.” The company is not asking the customer to join an advisory board. It is offering a bounded way to solve a known problem early, with extra access to the product team, in exchange for real use and disciplined feedback.

The outreach should begin with evidence from the prior relationship: the problem the customer already paid to solve, the repeated step now handled by the product, and the result the pilot is intended to produce. The first message should make clear that this is not another open-ended consulting engagement.

A strong invitation contains five points:

  • why this customer appears to fit;
  • what workflow the product handles;
  • what the pilot will and will not cover;
  • what the customer must contribute;
  • what decision both sides will make at the end.

The demo should show the standard product path first. Do not begin by asking what the customer wants built. Show how the product handles the target job today, then ask the customer to compare it with the current workflow.

Useful discovery questions concern observed behavior: Who performs the work? How often does it occur? What triggers it? Which systems and approvals are involved? Where do errors or delays appear? What does the current process cost? What result would justify changing it?

The recruitment path should screen for commitment before the company allocates scarce product time:

flowchart LR
    A[Past service customers] --> B[Score problem, product, participation, and commercial fit]
    B --> C{Can the current product handle one real workflow?}
    C -->|No| D[Keep as a research contact]
    C -->|Yes| E[Demo and define a bounded pilot]
    E --> F{Will the customer commit users, data, time, and a decision date?}
    F -->|No| D
    F -->|Yes| G[Sign pilot commitment]
    G --> H[Onboard, observe use, and measure results]
    H --> I{Value delivered without custom service?}
    I -->|Yes| J[Convert and prepare to repeat]
    I -->|No| K[Narrow, change, or stop]

Text description: Service customers are screened for fit, shown a standard product workflow, and asked for concrete commitments. Only signed, measurable pilots proceed. Results then determine whether to convert, revise, or stop.

A one-page pilot brief can precede the contract. It should state the customer problem, users, workflow, starting condition, pilot scope, non-goals, duration, success measures, each side’s responsibilities, feedback schedule, support boundaries, price or discount, and the decision expected at the end.

Microsoft’s guidance for proof-of-concept work similarly recommends defining goals and success criteria, selecting a small set of representative scenarios, choosing participating users, setting a duration, gathering feedback, and making an explicit scale-or-stop decision.

The signed document should then make the commitment real. At minimum, it should cover product access, participation and feedback, term and termination, fees if any, confidentiality, intellectual-property treatment, and the company’s right to use feedback to improve the product for all customers. Depending on the product, it may also need data-protection, security, service-level, liability, and acceptable-use terms.

A public design-partner standard includes both paid and unpaid options and makes regular feedback, references, case studies, future discounts, and specific provider obligations explicit choices rather than assumptions.

Whether the pilot must be paid is a judgment call. Payment is useful evidence that the problem has budget and that the buyer can complete a commercial process. A free or deeply discounted pilot can still be rational when the product is immature, the learning value is unusually high, or the customer must contribute costly access, data, staff time, or reference rights.

The rule is not “always charge.” The rule is require a meaningful commitment and define the route to a normal commercial relationship before the pilot starts.

Run a product test, not a disguised consulting project

A design partnership fails when the customer believes it has purchased a custom development team, or when the provider quietly supplies enough manual labor to make weak software appear successful. The pilot must test the standard product, including the onboarding, configuration, support, and handoffs the company expects to repeat.

Three boundaries protect the test.

First, define one primary job and one first-value milestone. The first-value milestone is the earliest observable point at which the customer receives a useful result from the product. It might be a completed analysis, an approved workflow, a reconciled data set, a generated deliverable, or a task completed with materially less time or effort. Measure the time and human help required to reach it.

Second, separate product work from service work. Product work improves a capability that should help many similar customers. Service work configures, trains, migrates, or advises for this customer. Both may be legitimate, but they must be labeled and measured. If each partner requires different logic, repeated manual processing, or founder judgment to succeed, the company has not yet proved a standalone product.

Third, treat requests as evidence, not orders. Ask what event produced the request, which user is blocked, how often the problem occurs, what workaround exists, and whether the missing capability affects purchase or value.

Gong has publicly described using customers as design partners for unusual enterprise needs while avoiding hard commitments to every requested requirement or date and focusing instead on the customer outcome. That is the right posture: listen closely, but preserve product judgment.

Public programs also show that good participation criteria are concrete. GitLab’s Co-Create program, for example, limits its highest-touch support to certain existing customers and looks for engineering capacity, technical curiosity, and willingness to contribute. In return, it offers structured workshops, direct engineering support, legal help, and ongoing mentorship.

The setting is different from an early startup, but the lesson transfers: select partners that can do the work, state the required effort, and make the provider’s contribution equally clear.

Basecamp teaches the complementary lesson. A product can emerge from a service company’s own repeated operational pain, but that origin should not become permission to preserve every part of the service. The design-partner program must identify the smallest product workflow that customers can adopt independently, then expose where onboarding or support still relies on the old delivery model.

The operating team should review pilots on a fixed cadence. A short weekly internal review should examine what users actually did, whether the first-value milestone moved, new support or custom-work hours, unresolved risks, product changes affecting multiple partners, and the next customer decision. Customer feedback sessions should use product behavior and workflow evidence as the starting point, not a tour of a feature-request list.

Treat commitments and results as different kinds of evidence

The primary measure is the number of design partners, but a raw count can hide weak work. A company can report eight “partners” when three have only expressed interest, two have signed but never supplied data, one is waiting for procurement, and two are active. The measure therefore needs defined stages.

Evidence levelWhat should existWhat it proves
Qualified candidateNamed account, selection score, target workflow, sponsor, known risksThe account is worth pursuing
Active outreachContact history, demo date, objections, owner, next actionRecruitment is being managed
Committed partnerSigned agreement, pilot owner, users, data or access plan, start date, review cadenceThe customer has accepted specific obligations
Active pilotUsers onboarded, product events or workflow records, support log, baseline measuresThe product is being tested in real work
Evaluable resultSuccess criteria assessed, product and service effort separated, customer decision recordedThe pilot produced decision-quality evidence
Converted customerPaid commercial agreement or documented renewal or expansion pathThe product can progress beyond the pilot

The expected deliverable should therefore contain three linked records: a prioritized design-partner list, an outreach pipeline, and signed pilot commitments. Each signed commitment should connect to a pilot scorecard.

The scorecard should include:

  • activation;
  • time to first value;
  • completion of the core workflow;
  • customer outcome;
  • user participation;
  • support hours;
  • custom-work hours;
  • defects or blockers;
  • willingness to pay;
  • the end-of-pilot decision.

The working target of five to eight committed partners is a sensible capacity-planning starting point for many small teams, but it is not an industry benchmark. The right number depends on how similar the customers are, how complex implementation is, how much support each pilot consumes, how diverse the target market is, and how much evidence each account can provide.

Qualitative research offers a useful caution about interpreting small samples. The “information power” principle says fewer participants may be sufficient when the sample is highly specific, the question is narrow, and the dialogue is rich.

In a separate study of 25 interviews, researchers reached broad topic saturation at nine interviews but needed 16 to 24 to develop a richer understanding of several issues. Those studies concern research interviews, not software pilots, so they do not validate a five-to-eight target. They do explain why a focused cohort may reveal recurring product problems without proving that the company understands the whole market.

Use the target with explicit adjustment rules:

  • Fewer partners may be appropriate when each pilot requires deep integration, regulated data, substantial implementation, or high-touch executive access.
  • More partners or sequential cohorts may be needed when the company is testing several industries, buyer types, workflows, or pricing models.
  • A narrow, homogeneous cohort is stronger for testing repeatability within one segment.
  • A deliberately varied cohort is stronger for finding boundary conditions, but weaker for drawing conclusions about one standard setup.

The task is not complete merely because the agreements are signed. Signed commitments prove recruitment and seriousness. Active use proves participation. Measured value proves usefulness. Conversion proves commercial progress. Low support and low custom-work requirements provide early evidence that the product can be delivered repeatedly.

Know when the company has enough evidence to move on

The work is complete when the company can show a traceable path from account evidence to partner selection, outreach, signed commitment, active use, measured result, and commercial decision.

The following should be true:

The chosen partners share the target problem but are not all copies of one politically convenient customer. Each pilot tests a product workflow that exists now. The scope and non-goals are written. A customer sponsor and working owner are named. Users, data, access, and feedback time are committed. Success measures and a decision date are agreed. Product effort, implementation effort, support effort, and custom work are recorded separately. The company knows what would cause it to convert, extend, change, or stop each pilot.

Several failure modes can make the program look complete when it is not. A logo without active users is not a design partner. A non-binding expression of interest is not a pilot commitment. A successful outcome created by unrecorded founder labor is not proof of a repeatable product. A long feature list is not evidence of demand. A free pilot with no staff time, data, or decision process is not meaningful commitment. Five customers from five unrelated markets do not establish one repeatable offer.

The final decision is not whether the partners liked the team or the demo. It is whether a defined group of customers used the product for a real job, reached value with acceptable support, and showed a credible willingness to continue under normal commercial terms.

When that evidence exists, the company can invest in a repeatable sales, onboarding, and support system. When it does not, the result is still useful: narrow the customer, reduce the workflow, change the product, or stop before custom service quietly becomes the business again.

Sources

Primary sources

  • U.S. National Science Foundation, “About Teams” and national I-Corps participation requirements.
  • Common Paper, “Design Partner Agreement” standard and cover-page guidance.
  • Microsoft Learn, “Deliver a Proof of Concept.”
  • Basecamp, “Where We Came From,” and 37signals, “Scratch Your Own Itch.”
  • GitLab Handbook, “Co-Create Program.”
  • Gong, “Four Strategies for Building Best-in-Class Enterprise Products.”

Open research

  • Eric von Hippel, “Lead Users: A Source of Novel Product Concepts.”
  • Glen L. Urban and Eric von Hippel, “Lead User Analyses for the Development of New Industrial Products.”
  • Woojung Chang and Steven A. Taylor, “The Effectiveness of Customer Participation in New Product Development: A Meta-Analysis.”
  • Kirsti Malterud, Volkert Dirk Siersma, and Ann Dorrit Guassora, “Sample Size in Qualitative Interview Studies: Guided by Information Power.”
  • Monique M. Hennink, Bonnie N. Kaiser, and Vincent C. Marconi, “Code Saturation Versus Meaning Saturation: How Many Interviews Are Enough?”