On this page
In short
- Building got cheap. Deciding what to build didn't, so a wrong build now costs runway and focus far more than it costs code.
- Test the idea on six questions before you brief anyone: the problem, who has it, what they do today, evidence they'd pay, why now, and your unfair advantage.
- Decide the stop line before the build. A result you agreed to accept in advance can't be argued away afterwards.
- "Don't build it" is a legitimate verdict, and the cheapest one you'll ever get.
Most products that fail had a build. What they lacked was a decision: who the product was for, what it would prove, and what would count as failure. That decision is cheap to make before the build and expensive after it, because once something exists, every conversation turns into a conversation about saving it. What follows is the method we use to make the decision, written down so you can run it on your own idea.
01Why is deciding harder now that building is easy?
A working prototype that once needed a small team now needs one person and an AI coding tool. That's real progress, and it moved the risk. When building was slow and costly, a weak idea usually died in the budget meeting. Now it survives the budget meeting, gets built, demos well, and dies later, after the runway has gone into it.
The tools tell you what can be built. They say nothing about what should be. So scope loosens, because another feature costs an afternoon. The buyer blurs, because a product that does a little of everything impresses everyone in a demo and nobody in particular. Product-market fit drifts further away while the backlog grows.
Building is cheap. Seeing isn't.
The expensive part of a product is no longer the code. It's the months you spend finding out the code was aimed at the wrong person. Every method in this piece exists to move that discovery to the front, where it costs a conversation instead of a release.
02What does "worth building" actually mean?
Worth building is not the same as a good idea. Plenty of good ideas solve problems nobody pays to fix. We use a narrower test. A product is worth building when:
- a specific buyer has the problem, and you can name them by role and situation rather than by market size;
- the problem already costs them something they can feel: money, hours, customers or risk;
- they already work around it, which proves the cost is real;
- a small version of the product could show whether they'd switch;
- and you've agreed, before building, which result would make you stop.
The last point is the one teams skip, and it's the one that keeps the other four honest. Without a stop line, every result becomes a reason to keep going. A weak signal is "early days". A bad one is "the wrong segment". Nothing ever counts as a no, so nothing ever stops.
03Which six questions should you answer before you brief anyone?
We call this the Idea Diagnosis. It's six questions, and each one is answered with evidence rather than opinion. If an answer starts with "I think", it isn't finished.
- Is the problem real? People describe it without being prompted, in their own words, and it shows up in what they do, not only in what they say.
- Who has it? Name the buyer narrowly enough that you know where to find them. "Small businesses" is a market. "Salon owners who lose artists over commission disputes" is a buyer.
- What do they do about it today? The workaround is the most useful evidence you'll find: a spreadsheet, a group chat, a person whose job it has quietly become.
- Is there evidence they'd pay? Money already spent on the workaround, a paid pilot, or budget someone has moved. Compliments don't count.
- Why now? Something changed: a cost, a rule, a behaviour, or a capability the model now has that it didn't before. If nothing changed, ask why nobody has solved it already.
- What's our unfair advantage? Access to the buyer, data nobody else holds, distribution, or domain knowledge that took years to earn.
The output is one of three verdicts: build, test first, or kill. We come back to what each one means below.
04How do you know the problem is real, and not just interesting?
The most reliable test we know is to ask how far the pain has climbed up a ladder: mentioned, complained about, worked around, paid for, switched for. It's our adaptation of Jobs-to-be-Done thinking, and we credit it as such. Each rung needs evidence before you climb to the next one. Most ideas that feel strong are sitting on the second rung: people complain, loudly, and do nothing.
Two of our own ventures started from the third rung, where the workaround already existed. With otlo, community builders were running their communities on a group chat, a spreadsheet and an Instagram account. Nobody needed persuading that the problem existed; they had built their own tools out of whatever was to hand. With apprn, salon owners were losing money where nothing could be checked: commission disputes that cost them their best artists, services billed but never delivered, cash that went missing. Their workaround was trust, and trust was failing.
The quickest way to find the rung is to ask about the past, not the future. "Would you use this?" invites politeness. "When did this last happen, what did you do, and what did it cost you?" invites evidence. Listen for two things in the answer: the workaround and the money. If the answers stay vague after a handful of conversations, the pain is lower on the ladder than it feels from the inside.
Insight
The workaround is the evidence. If nobody works around the problem today, you're not looking at a market yet. You're looking at an opinion.
05What's the smallest version that proves it?
Once the problem clears the ladder, the question changes from "should this exist?" to "what's the least we can build to find out?" Two filters do most of the work.
The first is the Scope Cut. Every idea on the list gets one of four verdicts, build first, later, cut or don't build, and every verdict carries a reason. The first release is whatever survives. The discipline is in the reasons: "later" is a decision with a trigger, not a polite word for "never", and "don't build" is a decision you'll defend when someone asks for the feature again.
The second is the Remove-the-AI Test. Take the AI out entirely and look at what's left. If a simpler product, built on rules, still delivers most of the value, an AI-native architecture isn't justified yet. That isn't an argument against AI. It's an argument for spending the model where it changes the outcome, and nowhere else.
apprn shows why the order matters. The part of apprn that everything else depends on isn't a model. It's a rule: an artist can only start a service by entering a one-time code sent to the customer with that booking. That one verified event feeds everything downstream, from honest commission to feedback that reaches the right artist. Get the record right first, and the intelligence built on it has something true to work with.
06What would make you stop? Decide before you build
Write the stop conditions down while nobody is invested yet. Each one should name something you could observe, such as a behaviour, a payment or a return visit, so that someone who wasn't in the room could check it. "If it doesn't feel right" isn't a criterion. Then put the conditions somewhere the team can't quietly edit them: the proposal, the kickoff notes, or in public.
We hold our own ventures to this. apprn is in a live test with real salons, and its go or no-go decision is made against kill criteria set before the test began. We'll publish the criteria before we publish the result, so the call can't be rewritten afterwards.
A result you agreed to accept in advance can't be argued with afterwards.
Stop lines feel pessimistic. They're the opposite. A team that knows what failure looks like can take bigger swings, because it knows it will notice when a swing misses.
07What does an honest verdict look like?
Every idea should leave the process with one of three verdicts, in writing, with the reasoning attached.
| Verdict | When it applies | What happens next |
|---|---|---|
| Build | All six questions have evidence behind them, the smallest version is defined, and the stop line is written. | Scope it, fix the scope, and build it. |
| Test first | The problem is real, but one assumption carries the whole venture: usually payment or reach. | Run the cheapest test that could break that assumption, before paying for a build. |
| Stop | The pain isn't strong enough, the buyer can't be reached, or someone else holds the advantage. | Keep the evidence. It often points at the version that is worth building. |
"Test first" is the verdict founders like least and need most. A paid pilot, a version run by hand behind a simple interface, or a real offer with a real price will teach you more about payment than a polished build ever will. It also costs a fraction as much.
A diagnosis that stops a bad build is worth more than a build that ships. Stopping here costs a few conversations. Stopping after launch costs the runway, the team's belief, and usually the next idea too.
08Where do founders fool themselves most often?
Across the products we've torn down and the ventures we've built, the same patterns keep turning up. We label this as observation, because that's what it is.
- Speed mistaken for direction. Shipping every week toward a buyer nobody has named.
- Friendly design partners mistaken for product-market fit. People who like you are not the same as people who need the product.
- AI features that need a workflow change users won't make. The feature works; the habit it depends on never forms.
- Model quality treated as the product. The buyer is paying for an outcome, and the model is one part of how it arrives.
- Distribution left until after launch. If you can't say how your first buyers will hear about it, the build is early.
We apply the same strictness to ourselves. otlo has customers paying real money. That proves willingness to pay. It doesn't yet prove product-market fit, and we won't claim fit until retention data supports it. Being strict about your own labels is the habit that makes the method work on anyone else's idea.
09Questions founders ask
Can I validate a startup idea without building anything?
Usually, yes. The strongest evidence comes before code: a workaround people already pay for, a paid pilot, or a commitment in writing. A prototype is good for testing whether people can use the product. It's weak evidence of whether they'll pay for it.
Is a waitlist evidence that people will pay?
No. A waitlist sits on the first rung of the ladder: it shows interest at the cost of an email address. Evidence of payment is money, a signed commitment, or budget someone has already moved.
Should I build the first version with AI coding tools?
To learn, often yes: a prototype makes the idea concrete and cheap to change. As the product your first customers depend on, rarely without more work, because real users bring data, load and edge cases a demo never meets. We've written about what breaks first when real users arrive.
What if the answer is "don't build it"?
Then the process did its job. You keep the evidence, the assumptions you tested and the reasoning, and that is often enough to find the version of the idea that is worth building.
Where this fits
