The Economy of Prototypes
I have argued for demos over decks my whole career, on the grounds that a working thing cuts translation loss. It does. It also makes people certain faster than it makes them better informed, and we just made it free. Building got cheap. The queue in front of the decision got five times longer.
The argument is here. The reference underneath it is Field Guide 04: Prototypes, in three questions.
This quarter, in a product review somewhere in your company, someone is going to skip the deck.
They will open a link instead. The thing will work. Not a mockup, a running version with real data in it, built in the two days since the last meeting. The room will click around for a few minutes and then start arguing about the button copy. A decision that used to take a month gets made before lunch. Everyone leaves feeling good. And not one person in that room could tell you what changed their mind.
I have spent most of my career arguing the value of that moment. I want to talk about what it costs.
Demos Over Decks
Early on I realized that I understand things better when I build them. Not perfectly, and not only that way. But building hands you a grade of detail that reading and talking do not: where the opportunity actually sits, which part is harder than it looked, what breaks the first time somebody uses it wrong, what I had been waving past without noticing I was waving past it.
Of the three, talking, reading, building; talking is the worst, honestly.
Another problem. What was in my head arrived in someone else's head at not quite the right frequency. Understanding, know-how, ego, agenda, take your pick. The loss is real regardless, and the size of the discount was set by how well I could articulate it. And I'm not articulate. We know this. A messy deck widened the gap.
So: demos over decks. Not because they look better. Because they cut translation loss. They make intent visible. They amplify the approach. They achieve a higher level of alignment.
AI reduced the cost of trying to find mutual understanding of intent, which should feel like a win, and mostly does.
Building Got Cheap, And the Numbers Are Not Subtle
Start with the thing everybody agrees on, because it is the only part of this memo that is not contested.
Lovable says users now create a million new projects a week, up from a hundred thousand a day a year earlier. Vercel's CEO says 30 percent of the apps now running on his platform came from agents rather than people. Both are company-reported and neither is audited, so take them as evidence of demand rather than value. The one number in this category carrying securities-law liability comes from Figma, which told investors in May that about 60 percent of its customers spending over $100,000 a year used its prompt-to-app product weekly, up from over 50 percent the quarter before. On the engineering side, GitHub logged 518.7 million pull requests merged in public repositories in the year to August 2025, 29 percent more than the year before. DORA's 2025 survey of nearly 5,000 technology professionals put AI use at work at 90 percent, with a median of two hours a day spent inside the tools.
Cheap. Fast. Everywhere.
If you run anything, the natural read on all that is capacity. More things built per quarter, at lower unit cost. That is not what the rest of the data shows.
The Bottleneck of Decision
Deloitte surveyed 3,235 business and IT leaders across 24 countries and found that only 25 percent had moved 40 percent or more of their AI pilots into production. Fifty-four percent expected to get there within three to six months, which is a forecast I have watched slide for about two years now. S&P Global's 2026 wave puts 37 percent of AI initiatives launched in the past year as live and delivering value, with the rest stuck between development and partial deployment.
You have probably seen a much louder version of this, the one about 95 percent of AI pilots failing. But take a look. The source is a preliminary-findings document from a project affiliated with the MIT Media Lab, built on 52 interviews and 153 survey responses, never revised past its first version and never defended against the criticism. And its 95 percent describes organizations reporting no return on their AI spending, which is a different and much larger denominator than pilots failing. The gap between pilot and production is real. You do not need the inflated number to prove it.
Here is the part that clarified it for me. LinearB looked at 8.1 million pull requests from 4,800 teams and sorted them by how they were written: by a person, by a person with AI help, or by an agent outright. Human-authored work merged within thirty days 84.5 percent of the time. AI-generated work merged 32.7 percent of the time. The delay was not in the reviewing. Once a human started, AI-generated reviews finished faster, 194 minutes against 252. The delay was in the waiting. Human-authored changes sat about 200 minutes before anyone picked them up. AI-generated changes sat more than sixteen hours.
Production went up. Review got faster. The queue in front of a human deciding got more than five times longer.
A Google engineer said this out loud in DORA's follow-up research: "Reviewing [another's] code is so much harder than writing it. AI tools are increasing the rate at which people can churn out code that needs to be reviewed." DORA's own framing is that velocity for the author becomes cognitive load for the reviewer.
The constraint moved. It used to sit on the build. It sits on the decision now, and the decision is the part of your company nobody upgraded. Gartner found 75 percent of IT application leaders piloting or deploying AI agents, and 13 percent who strongly agreed they had the governance to handle them.
What a Working Thing Does
I sold demos over decks on the grounds that they transmit intent better than slides.
That claim holds up. Engineering researchers who showed the same medical device to 45 clinicians as sketches, cardboard, CAD, and 3D prints found that tangible prototypes drew substantially more useful and better-justified feedback than the abstract versions.
Show people a real thing and you get real answers. That is the case for demos and it is a good one.
What I never priced was the second effect.
Jeff Sauro's synthesis of eleven studies on prototype fidelity is the uncomfortable one. Rough prototypes, including paper ones, surface most of the major problems a polished build surfaces. Not all of them, and Sauro is careful to say so. But the gap in what you learn is much smaller than the gap in what you spend, while the gap in how people feel about it is not small at all: participants consistently prefer the high-fidelity version. Better finish reliably buys you a warmer room. It buys you a fraction of that in new information.
Then there is the ownership problem. Norton, Mochon and Ariely's work on self-assembly found that people who build a thing value it well above what anyone else will pay for it, and the effect has a precondition worth sitting with: it only appears when the builder finishes. Participants who assembled and then destroyed their creations, or never completed them, showed no bump at all.
Completion is the trigger.
So the honest description of what AI did to the demo is three things happening together.
- It closed the translation gap, which is what I always wanted.
- It manufactured finished artifacts that their creators overvalue and their audiences find more convincing without being better informed.
- And it pushed all of that onto the one part of your organization that never got faster: the decision.
The Part Nobody Wants to Say
Prototypes did not replace the deck.
The demo was an antidote to slide theater back when it was expensive, because the cost was the proof. If you built the thing, you had already done the hard part, and the demo was evidence of work completed. That was never a property of demos. It was a property of the price. When a prototype cost a week of somebody's life, a working demo was a costly signal. At an afternoon, it is a rendering of an opinion, and it is the most persuasive rendering anyone in your company has ever had access to.
Which means the current advice to bring a prototype instead of a deck is, in most rooms, advice to bring a stronger persuasion device to a meeting that was already decided by whoever had the strongest one. Same politics. Better graphics. Faster.
And the pile is growing. Every team that ships thirty prototypes and two products has not become efficient. It has produced twenty-eight completed artifacts with owners emotionally attached to them, sitting in a queue in front of the same number of decision makers, who have the same number of hours, and who are now expected to say no more often than they have ever said no in their careers, to things that work, in front of the people who built them.
Nobody is running the meeting where that gets counted.
Where This Could Be Wrong
Two real outs, and a third I will name and set aside.
The first is the review-queue reading. LinearB's merge-rate collapse, 84.5 percent down to 32.7, could be the decision system working exactly as designed: AI produces more low-quality candidates, humans correctly reject more of them, and the falling merge rate is a filter doing its job rather than a bottleneck failing. I think the sixteen-hour pickup delay argues against the benign version, because queue time measures capacity rather than quality. But LinearB has not published how it decides a change was AI-generated, or whether it controls for team or repository, so I am leaning on a number whose method I cannot inspect. Weigh it accordingly.
The second is that the conviction problem has a known fix, and it is old. Stanford's parallel prototyping study found that designers who built several options before getting feedback produced measurably better results, more divergent ideas, and, most relevant here, took criticism far better. Nearly half the serial group reacted negatively to critique. None of the parallel group did, because no single artifact carried their identity. The study is small, 33 people finishing a banner-ad task in 2010, so do not build a policy on it alone. But the mechanism is exactly right for this moment. If cheap building makes people fall in love with what they finished, the counter is to make everyone finish four. Expensive prototyping forced serial work. Cheap prototyping makes parallel work possible for the first time, and almost nobody is organized to use it that way.
The third one, quickly. Maybe review capacity catches up as tooling improves and this all reads like 2004 complaining about email volume. Maybe. It does not help you this year, and the organizations that came out of the last few volume shocks intact were not the ones waiting for tooling.
The Bet
Judgment was already the product. I argued that in Why You're All Suddenly Talking About Human Judgment and again in When Time and Materials Go to Zero, and the supply side of it worries me more every quarter. What changed is that the cost of being wrong about what to build has quietly moved from the build to the portfolio, and nothing in your operating cadence tracks it.
So the bet: within a couple of years, the companies that got good at this will be measuring the thing nobody measures now. Not time to prototype. Not pilots launched. Kill rate. How many finished, working, good-looking things did we decide not to pursue this quarter, how fast did we decide it, and how much did it cost us to find out. A team that builds forty and kills thirty-eight is operating well. Today it looks like a team that wasted thirty-eight builds, which is why nobody publishes the number.
And a smaller one, for your next review. When someone opens a working demo, the useful question is no longer whether it works. It obviously works. A prototype with no siblings is not evidence. So: what else got built alongside this one, and what got left out?
That is the argument. The taxonomy under it, sorted and sourced, is Field Guide 04: Prototypes, in three questions.