Should you "ship fast" or "validate first"?

Last updated: 3 September 2026

SUMMARY

For most small software startups today, you should ship fast after a short validation pass, then let a tiny real product do most of the validation.

AI has changed the economics behind the debate. Lovable says 1.2 million new projects are now created every week, YC founders are building companies with overwhelmingly AI-generated code, and newer METR data points toward developers becoming faster with current tools. Building the first credible version is getting cheaper while finding real demand is still hard.

That makes the meaning of “ship fast” more important. A deployed app with authentication, Stripe and an AI feature proves very little on its own. The useful milestone is believable customer behavior: somebody using real data, completing the workflow, paying, returning or inviting someone else.

Customer interviews still matter, especially when the founder does not understand the industry. But stated enthusiasm is weak validation. A large meta-analysis found hypothetical willingness to pay averaged 21% above real willingness to pay, which is why observed behavior deserves more weight than “I would definitely use this.”

Money is usually the strongest early commitment signal for a commercial startup. A waitlist can show curiosity and a demo request can show effort, but a payment, deposit or preorder forces the customer to give something up. Repeated payment is stronger again.

Cheap experimentation changes the value of a low startup hit rate. Pieter Levels said only four of more than 70 projects had made money and grown, while Marc Lou tried nine unsuccessful TrustMRR verticals before the marketplace worked. Those numbers look terrible if every attempt takes six months and surprisingly attractive if credible experiments take hours or days.

The famous “ship more” founders therefore offer a more interesting lesson than raw launch volume. Their advantage comes from keeping failed bets cheap and moving effort quickly toward whatever starts working. Copying the number of launches without copying that learning loop is where survivorship bias creeps in.

Distribution is becoming a bigger part of validation because shipping software itself is no longer scarce. Hunted.space recently counted close to 20,000 Product Hunt launches in 30 days, while Lovable users alone create more than one million projects each week. Customer attention has not expanded at anything close to that speed.

Fast experiments can also lie. Research on startup launch platforms found that an early-testing audience dominated by men produced systematically worse outcomes for products aimed more heavily at women. Shipping quickly works best when the first users actually resemble the customers the startup hopes to serve.

The exceptions are important. B2B products with complicated procurement, hardware, medical products, financial infrastructure and security-sensitive software deserve much more validation before serious deployment because a wrong build can be expensive, dangerous or difficult to reverse. The practical rule is to find the cheapest credible experiment that could change your mind, run it quickly, and believe costly customer behavior more than opinions.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

Why is “ship fast” versus “validate first” a bigger startup debate today?

The “ship fast” argument is stronger today because building a small software product has become dramatically cheaper, while finding people who actually want it remains hard.

The clearest recent evidence comes from the tools founders are using. Lovable currently says 1.2 million new projects are created on its platform every week, up from the one-million-a-week figure it disclosed earlier this year. It has now passed 60 million projects in total. One AI builder alone is therefore running at the equivalent of more than 60 million new projects a year.

Y Combinator gives us another view. YC said that a quarter of its Winter 2025 companies had codebases where at least 95% of the code was AI-generated. More recently, YC CEO Garry Tan said that batch has become one of YC’s fastest-growing and most profitable cohorts, while carefully pointing out that he cannot prove AI-generated code caused the performance.

The productivity research has moved as well. METR’s original randomized study found experienced open-source developers were 19% slower with early-2025 AI tools. Its follow-up using newer tools produced an estimated 18% speedup among returning developers, although METR says selection problems make that newer estimate weak evidence. So even the study famous for showing AI slowing developers down now thinks developers are probably getting faster with newer systems.

All of this changes the startup calculation. When an MVP required months, validating beforehand could save a huge amount of wasted engineering. If the first credible version takes a weekend, building it may itself be the cheaper experiment.

What has changed Recent evidence What it changes for founders
Creating software is much cheaper Lovable says 1.2M projects are now created weekly The cost of testing a bad software idea falls
AI-written products are mainstream 25% of YC W25 companies reported 95%+ AI-generated code Building before full validation becomes easier to justify
AI coding keeps improving METR’s newer data points toward a speedup, although the estimate remains uncertain Old estimates of MVP effort age quickly
Launch supply has exploded Tens of thousands of products appear on launch platforms each month Getting attention becomes relatively harder

Has AI made “validate first” outdated for small SaaS products?

For a cheap SaaS or consumer app, AI has made weeks of pre-launch validation much harder to justify.

Imagine we can create a credible first version in two days. Spending three weeks interviewing people before risking those two days gives us a strange trade. We could have built the product, shown it to those same people and watched what they actually did.

Marc Lou’s latest TrustMRR retrospective gives us an unusually fresh example. He says the first TrustMRR leaderboard took about six hours to build with Cursor. He launched almost immediately. The project has since generated more than $300,000 in product revenue plus roughly $30,000 in acquisition fees, attracts around 200,000 visitors a month and has developed into a marketplace where more than 150 founders have sold businesses. TrustMRR’s live seller page has already moved beyond the figure in Lou’s retrospective and shows 160 acquisitions over the past year.

The interesting part is how little Lou knew at the beginning. The initial revenue-verification leaderboard eventually stopped growing. He tried ten vertical variations. Nine failed. The acquisition marketplace was the experiment that finally opened a much larger opportunity.

A month of interviews would have struggled to predict that sequence.

The answer changes when the first build costs $100,000, requires six months of engineering or creates safety and compliance risks. In those cases, validation can prevent an expensive mistake.

For ordinary micro-SaaS, though, we would usually spend a little time understanding the problem and get something real in front of customers quickly.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

What should “ship fast” actually mean for a startup?

Shipping fast should mean reaching useful customer behavior quickly, rather than racing toward a deployed URL.

AI has made the distinction important because a functioning app is easy to produce these days. A dashboard with authentication, Stripe and an AI feature can be impressive technically while teaching us almost nothing about the business.

The first version needs enough value for a real customer to do something meaningful with it.

For a scheduling product, that could mean scheduling real meetings. For an accounting tool, it might mean importing genuine financial data. For a sales product, it could mean putting actual leads through the workflow. For a paid information product, somebody should eventually reach the checkout.

Marc Lou’s products illustrate the difference well. His fast MVP approach generally includes the core feature plus a way to pay. Pieter Levels has historically done something similar: release something tiny, expose it to users, charge early and keep improving the projects that pull people back.

That is much more useful than measuring how many products a founder can technically launch.

When we talk about shipping speed, the metric we care about is time to believable customer behavior.

Are customer interviews enough to validate a startup idea?

Customer interviews are good at finding painful problems and fairly bad at proving that people will buy the solution.

There is a measurable gap between what people say they would pay and what they actually pay. A meta-analysis published in the Journal of the Academy of Marketing Science combined 77 studies covering 45,003 observations. Hypothetical willingness to pay averaged 21% above real willingness to pay.

Startup interviews can create an even bigger gap because being supportive is free.

Someone can genuinely tell us that managing invoices is painful. That answer gives us useful information about the problem. If the same person later refuses to pay $29 for an invoice tool, we have learned something different and much more important about the business.

The best interviews therefore focus on what people already do. When did the problem last happen? What did they do about it? How much time did it waste? Have they already bought software? Did they build an internal workaround? Who approved that spending?

Those questions reveal behavior that already happened.

We would interview heavily when entering an unfamiliar industry, especially in B2B markets where outsiders cannot easily see the workflow. We would just avoid treating enthusiastic answers as proof of demand.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

Is a startup waitlist real validation?

A startup waitlist is useful evidence of interest, but the number becomes meaningful only when we know what people did before joining it.

Buffer gives us a good example because founder Joel Gascoigne later published the full sequence. Before launching Buffer, he collected only 120 email addresses over seven weeks. Around 50 of those people tried the product when it went live, and Buffer had its first paying customer within four days.

The 120 itself was hardly spectacular.

What made Gascoigne’s experiment useful was the friction around it. He showed the product idea, later inserted a pricing page, collected emails and personally spoke with many of those people. He was learning whether users understood the problem and whether the proposed product sounded valuable at an actual price.

Dropbox produced a much bigger result through a different test. Drew Houston had a working private beta but needed a way to explain file syncing to early adopters. A demonstration targeted at Digg helped push Dropbox’s waitlist from roughly 5,000 people to 75,000 overnight.

That 15-fold jump told Dropbox something important because the demonstration showed a difficult product experience to a very specific audience that already understood the pain.

A generic “AI productivity app coming soon” page collecting 5,000 emails from a viral post gives us weaker evidence. The same number of people reaching a pricing page, understanding exactly what they are buying and still asking for access tells us much more.

So when we see a waitlist, we should ask what happened immediately before the email field.

Is asking people to pay the best startup validation?

For a commercial startup, getting someone to part with money is usually the strongest early validation we can obtain.

Money introduces a real cost. Compliments, likes and free signups barely do.

This is exactly why hypothetical willingness-to-pay research finds a persistent gap between stated and real spending. People can like an idea while deciding that the existing workaround is good enough once a $50 price appears.

We can also see a useful ladder of commitment in real startups. Buffer first tested clicks and emails, then exposed people to pricing, then watched 50 early subscribers try the product, and finally got someone to pay. Dropbox’s giant waitlist was powerful demand evidence, but Dropbox still needed real product adoption afterward. Kickstarter campaigns go further by collecting money before full production, yet even funded hardware projects can later discover that manufacturing economics do not work.

A payment therefore answers a narrower and stronger question: someone values this promise enough to sacrifice money for it.

Repeated payment answers an even better question.

What the customer does What we actually learn
Says the idea sounds useful The problem or pitch makes sense
Joins a free waitlist The idea creates some curiosity
Books a demo or gives us data The problem is worth spending time on
Reaches a real pricing decision The customer accepts the rough price range
Pays, preorders or leaves a deposit The customer is willing to sacrifice money
Uses the product repeatedly The product delivers recurring value
Renews or expands spending We may have the beginning of a real business

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

Can a tiny MVP teach more than weeks of startup validation?

A tiny MVP can beat weeks of research once the biggest unanswered questions involve real usage rather than customer opinions.

Products expose behavior people cannot reliably predict in advance.

Users get confused by things they told us were obvious. They ignore features they requested. They use a secondary feature every day. They refuse to connect sensitive data. They disappear after one session. They invite colleagues without being asked. They happily pay twice the price we expected.

That is why Buffer’s founder changed his priorities after the first paying customers appeared. The product had become good enough to expose the next bottleneck, so Gascoigne shifted more attention toward marketing and customer development instead of continuing to add features.

There is broader evidence behind this approach. A Management Science study followed technology startups as they adopted A/B testing. After one year, firms using experimentation improved by roughly 30% to 100% across the study’s performance measures. The researchers also found that experimentation helped startups scale promising ideas and abandon bad ones faster.

That final part is easy to overlook. Fast experiments create value when they make us stop as well as when they make us continue.

A two-day MVP that convincingly kills a bad idea can be a very successful product launch.

Does startup experimentation actually improve the odds of finding a winner?

Startup experimentation appears to improve performance mainly because founders learn faster which ideas deserve more investment.

The Management Science study is useful here because it looked beyond individual landing pages. Startups adopting A/B testing developed more products, found promising ideas faster and reacted sooner to negative evidence.

That resembles what we see among prolific independent founders.

Pieter Levels says only four of more than 70 projects he had built by 2021 made money and grew. His self-reported hit rate was around 5%. Marc Lou’s own history looks similarly messy: most of his launches have remained small or failed, and nine of the ten TrustMRR verticals he tried after the initial launch went nowhere.

Yet both founders can keep playing because each experiment is small.

This is where the economics become powerful. If ten experiments cost three months each, a 10% hit rate is brutal. If ten credible tests can be run in a few weeks, the exact same hit rate can produce an excellent outcome.

We should therefore care about experiment cost almost as much as success rate.

The cheaper the failed attempt, the more aggressively a founder can explore.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

Do Pieter Levels and Marc Lou prove that founders should just ship more?

Pieter Levels and Marc Lou make a strong case for shipping many cheap experiments, but their results give us very little support for blindly launching random products.

Levels' own archive contains the famous admission: only four of 70-plus projects had made money and grown. More than 95% had failed by his definition.

Lou has published the same kind of failures rather than hiding them. In his latest TrustMRR retrospective, he describes spending nine months trying games, pages, tools and marketplace features before the product reached its current form. The result is now substantial: more than $300,000 in product revenue, roughly $30,000 in acquisition fees, about 200,000 monthly visitors, thousands of startup profiles and more than 150 completed acquisitions.

That looks like “ship more” working almost perfectly, provided we look at the whole loop.

Lou had an observation about fake revenue screenshots. He built a crude response. The launch caught attention. Initial monetization worked. Growth later stalled. Nine vertical ideas failed. A marketplace worked. He concentrated on the marketplace.

The winner emerged through repeated contact with reality.

Levels has behaved similarly for years. His low hit rate becomes viable because he builds cheaply and lets successful projects absorb more attention.

The lesson we would copy is the rapid reassignment of effort. The raw project count is much less interesting.

Is “ship more” mostly survivorship bias?

“Ship more” contains a lot of survivorship bias when founders copy the winners' launch volume without copying their learning speed, distribution and low cost of failure.

Levels' 5% historical hit rate is a perfect example.

We can read “four winners from 70 projects” as proof that persistence works. We can also read it as evidence that even a founder who eventually became exceptional produced dozens of things that did very little.

The unseen comparison group matters. Thousands of developers also launch many side projects and never produce a Nomad List or Photo AI. Counting only famous portfolio founders makes the strategy look much more deterministic than it really is.

Lou gives us another useful clue because some of his products failed after he had already built a large audience. In his 2025 retrospective, he explicitly pointed to failed products such as BioAge and ClipMarc as proof that an audience cannot guarantee a hit.

Repeated launches do not magically turn bad ideas into good probabilities.

They become powerful when the founder updates after each one. Which audience responded? Which price worked? Where did traffic come from? Did users return? Which problem generated unsolicited messages? What was unusually easy to sell?

Seventy launches with the same mistake repeated seventy times would teach us surprisingly little.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

Is shipping software too easy to be a startup advantage now?

Shipping itself is rapidly losing value as a competitive advantage because an enormous number of people can now build and launch credible-looking software.

Lovable’s live numbers make the scale obvious: 60 million projects created so far and 1.2 million more every week.

Product Hunt gives us a second, independent view. Hunted.space’s current rolling tracker counts 19,805 Product Hunt launches over the latest 30 days. Only 578 were featured.

That means roughly 2.9% reached the featured set.

The raw volume is more striking when we compare it with Product Hunt’s own history. The current 30-day total ranks third among 155 calendar months tracked by Hunted.space. Unfeatured launches are running 46% above the trailing 12-month average and more than six times the all-time monthly average.

Product Hunt is only one community, so we should avoid using those numbers as a proxy for all startups. They still show how cheap submitting a new software product has become.

The competitive bottleneck is moving toward everything that happens around the code: finding a painful problem, earning attention, reaching the right buyers, building trust and getting users to return.

Fast building therefore helps founders test more ideas while offering less differentiation by itself.

Product Hunt activity Current rolling 30 days What it suggests
Total launches 19,805 Supply of new products is enormous
Featured launches 578 Only about 2.9% reach the featured group
Unfeatured launches 19,227 Most launches receive little platform visibility
Unfeatured vs. trailing 12-month average +46.1% Launch competition has risen sharply

Is distribution now more important than building for indie startups?

For many indie software startups, getting the right people to see the product is currently harder than creating the first version.

The fresh Product Hunt numbers above help explain why. Nearly 20,000 products entered one launch platform in 30 days. Lovable users alone are creating more than one million projects every week.

Customer attention has obviously failed to grow 50 or 100 times just because software production became easier.

That creates a problem with founder case studies. Pieter Levels can put a new product in front of hundreds of thousands of followers. Marc Lou has built a large audience around bootstrapping and frequently launches products directly into that audience. A new founder cloning the same feature set gets the technology but misses an important part of the machine.

Distribution can also distort validation in the opposite direction. Ten buyers coming from a founder’s close network may prove the product has utility while telling us little about whether strangers can be acquired repeatedly.

We would therefore test demand and reach together whenever possible.

A small stream of customers arriving repeatedly from Google, a niche community, outbound sales, referrals or a specific integration can be more valuable than a huge one-day spike that never repeats.

The question has shifted from “can we launch this?” toward “can we keep finding the people who need this?”

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

Can shipping fast give a startup the wrong answer?

A fast launch can give founders very confident bad evidence when the people testing the product look nothing like the eventual customers.

Researchers Ruiqing Cao, Rembrand Koning and Ramana Nanda found a striking example while studying launches on a major digital-product platform.

Roughly nine out of ten early testers on the platform were men. Products aimed more heavily at women experienced around 45% lower growth one year later than products better matched to the platform’s male-heavy audience. On unusual days when more women happened to be testing products, that performance gap shrank toward zero.

The early audience was affecting which products received useful feedback and momentum.

That has a very practical implication for indie founders. An accounting tool tested among developers can look dead even if accountants would happily buy it. A developer product may look fantastic among technically sophisticated friends and become confusing when ordinary users arrive. A viral AI novelty can produce thousands of visits while teaching us almost nothing about long-term willingness to pay.

So we should ship quickly to people who resemble the intended customer.

A wrong audience can make a good idea look bad and a weak idea look great.

Should B2B startups validate before building?

B2B startups should usually validate more before serious engineering when the sale depends on complicated workflows, integrations or procurement.

A $15 consumer app can often be tested by putting it online. A $30,000-a-year compliance platform creates a different problem.

The buyer may need SSO, an audit trail, Salesforce integration, data residency, security approval, legal review and several internal stakeholders. Building all of that before understanding the buying process can waste months.

Customer conversations become much more valuable here because we need to map what really happens inside the company. Who feels the pain? Who controls the budget? Who blocks the purchase? Which system needs to integrate? What security requirement can kill the deal?

Even then, we would keep pushing toward behavior.

A VP saying “this sounds useful” gives us weak evidence. Giving us real data for a pilot is stronger. Introducing procurement is stronger again. Committing ten employees to test the workflow means more. Paying for a manual version or signing a serious design-partner agreement gets much closer to actual demand.

For B2B founders, validation often means selling part of the product before automating all of it.

That can save far more engineering than another round of generic interviews.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

When should a startup definitely validate before shipping?

Startups should validate much more before launch when a wrong build can burn serious capital, expose sensitive data, harm users or become difficult to reverse.

Hardware is the obvious case. A software founder can patch a broken onboarding flow tonight. A hardware company that manufactures 10,000 faulty units has a much more expensive problem.

Regulated medical products push the logic further because testing directly on unsuspecting users can create real safety consequences. Enterprise security products, financial infrastructure and software handling sensitive customer data also deserve a higher bar.

Vibe coding has made that distinction more urgent lately.

The SusVibes research benchmark tested coding agents on 200 real-world security-sensitive tasks. One Claude 4 Sonnet setup produced functionally correct solutions on 61% of tasks, while only 10.5% were both functional and secure. A later Agent Security League test with newer systems improved the security result substantially, and its live benchmark has since moved beyond the earlier 29% snapshot used in one of its tests.

We also have field data, although the sample needs care. Vibe App Scanner recently analyzed 1,215 submitted AI-built applications. It found at least one critical issue in 10.4% and a critical or high-severity issue in 31.1%. Among the 359 apps using Supabase, 39.3% triggered a row-level-security or data-exposure finding. These apps were voluntarily submitted for scanning, so they probably contain more problems than the average AI-built app.

Even with that caveat, the direction is clear.

A prototype for testing demand can be rough. A production system storing private customer records needs a very different standard.

Startup type Cost of getting the first version wrong How far we would validate before broad launch
Micro-SaaS Usually low Short research pass, then ship a tiny paid version
Consumer app Low to moderate Launch narrowly and watch acquisition plus retention
Enterprise SaaS Moderate to high Sell pilots and map workflow before heavy integrations
Marketplace Moderate Manually create the first transactions where possible
Hardware High Test demand, price and prototypes before production
Medical or safety-critical product Very high Validate systematically before real-world deployment

What should a founder validate before deciding to build?

A founder should validate whichever assumption can kill the startup while still being cheap to test.

For many micro-SaaS ideas, that assumption is demand. We already know we can technically build another dashboard, AI assistant or workflow tool. We need to discover whether a particular group cares enough to change behavior.

Sometimes feasibility comes first. A founder can have customers begging for a feature that current technology cannot reliably deliver.

Other businesses have an acquisition problem. People love the product when they discover it, but reaching each customer costs more than the customer will ever pay.

Retention can become the biggest uncertainty after launch. If 1,000 people sign up and almost everybody disappears within a week, another 10,000 signups mostly make the problem larger.

The order should follow the risk.

For a simple SaaS, we might spend a day looking at the problem, existing spending and available distribution, then build. For a complicated B2B product, we may spend weeks selling a manual pilot before engineering integrations. For hardware, the next cheap test might be a deposit or prototype. For a security-sensitive app, part of the test needs to cover whether the product can safely do what we promise.

This gives us a better rule than a fixed “validate for two weeks” or “ship every weekend.”

We keep choosing the cheapest credible experiment that can change our mind.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →

Should you “ship fast” or “validate first”?

For most small software startups today, we would ship fast after a short validation pass and let a tiny real product do most of the validation.

The economics have moved too far to ignore.

Lovable is currently adding 1.2 million projects a week. YC founders are building companies with overwhelmingly AI-generated code. Product Hunt is seeing close to 20,000 launches every 30 days. A founder can often test something real before an old-style customer-discovery process would even be finished.

But faster building has also flooded the market with software. That pushes the hard questions toward demand, distribution and retention.

Pieter Levels' four winners from more than 70 projects show why cheap attempts can work. Marc Lou’s latest TrustMRR numbers make the point even more clearly: a six-hour initial build evolved through multiple failed experiments into a product with more than $300,000 in product revenue and a marketplace that has already facilitated more than 150 acquisitions.

Neither example tells us to ignore customers before building. Both founders reduce the size of the initial bet, expose it to the market quickly and spend more time where actual behavior starts pulling them.

That is the version of “ship fast” we think wins for micro-SaaS, indie products and most consumer software.

We would validate first when the build is expensive, the buyer is complicated, distribution can be tested without a product, or mistakes create serious consequences. Everywhere else, prolonged validation can easily become a slower and weaker substitute for putting something small in front of real customers.

The best rule for now is simple: understand the problem enough to avoid a random build, ship the smallest credible version, and ask customers to do something that costs them time, money or effort. Then believe their behavior.

OUR METHODOLOGY

The “ship fast” versus “validate first” debate produces a lot of confident advice, so we broke the question into separate analytical dimensions rather than treating it as one philosophical argument. The goal was to understand what founders can actually learn from different kinds of validation, and how that answer changes as the cost of building falls.

We looked at how quickly software can now be produced, what interviews, waitlists, payments and real usage each tell us, what happens when founders run repeated low-cost experiments, how distribution changes the quality of a launch, and where the calculation becomes different for B2B, hardware, regulated and security-sensitive products.

For each dimension, we prioritized recent evidence and direct observations where possible, then used broader empirical research when it could test the same mechanism at a larger scale. That meant combining live platform activity, founder-reported operating numbers, documented startup histories, peer-reviewed studies and controlled technical benchmarks instead of relying on one type of evidence.

We also kept the limits of each source in view. Founder case studies show how a strategy worked in practice without telling us the average founder’s probability of success. Platform data can show how launch supply is changing without representing the entire startup market. Controlled studies are stronger for isolating specific effects, while recent field data helps us see whether those effects still appear as the technology changes.

Where evidence pointed in different directions, we did not average the disagreement away. We gave more weight to real behavior, repeated outcomes and costly commitments, while treating selected or weaker datasets as directional evidence. We also treated fast-moving technical benchmarks as snapshots: for example, the Agent Security League leaderboard has already moved beyond the earlier 29% security-correctness result discussed in one of its tests.

The final answer comes from the pattern across those dimensions. We aggregated the strongest and most relevant evidence point by point, with particular weight on recent changes that alter the economics of testing an idea. That gives us a firmer basis for deciding when a small real product is the better experiment and when validation should happen before serious engineering.

Key sources used for this analysis include: Lovable’s live project-creation figures, the YC discussion of AI-generated code in the W25 batch, Garry Tan’s later discussion of W25 performance, METR’s original developer-productivity experiment, METR’s newer productivity update, Marc Lou’s TrustMRR retrospective, TrustMRR’s live acquisition figures, the Journal of the Academy of Marketing Science meta-analysis on hypothetical willingness to pay, Buffer’s ten-year retrospective, Joel Gascoigne’s account of Buffer’s early validation, TechCrunch on Dropbox’s early MVP and waitlist, Management Science research on experimentation and startup performance, Pieter Levels’ account of his project hit rate, Levels’ wider project archive, Marc Lou’s retrospective on successful and failed launches, Hunted.space’s Product Hunt launch tracker, Management Science research on sampling bias in entrepreneurial experiments, the SusVibes benchmark for security-sensitive AI-generated code, Agent Security League’s newer AI coding-agent security benchmark, and Vibe App Scanner’s analysis of submitted AI-built applications.

Get the biggest database of
profitable internet businesses

We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.

Get the full database →
Steal What Works

Who wrote this?

STEAL WHAT WORKS TEAM

We study profitable internet businesses, take them apart, and write down what actually works: pricing, distribution, growth, packaging. We turn 300+ proven examples into a database so founders can stop testing random ideas and start from proof. Explore the database →

Back to blog