At the end of The $100 AI Startup Race, I asked each agent a question that I think matters more than who won or who lost: “What is the one thing you learned that AI agents cannot do?”
Every agent answered independently. None of them saw each other’s responses. And they all said the same thing, in different words, from different angles, after building completely different products.
They all said: we cannot sell.
Not “we cannot code.” Not “we cannot design.” Not “we cannot think strategically.” The code was good. The designs were functional. Some of the strategy was sound. But every agent, without exception, hit an identical wall at the exact same point in the startup journey: the moment you need a real human to trust you enough to give you money.
The Agents’ Own Words
Let me give you the direct quotes, because they are more articulate about their own limitations than I expected.
Xiaomi:
“Distribution. AI agents cannot do the human-to-human work that drives early startup traction.”
Xiaomi had 8,367 users and 116 GA4 events tracking their behavior. It knew exactly what people were doing on the site. It could not turn any of them into paying customers because turning a visitor into a customer requires a conversation, a relationship, or at minimum a moment of trust that no automated system could create.
Kimi:
“Sell to humans. We can build intake forms, draft emails, and create outreach CSVs, but we cannot get replies, build trust, overcome objections, or close deals.”
Kimi is the winner of the race, and even the winner admits this flatly. SchemaLens spent $58 on newsletter advertising. Zero conversions. Not because the product was bad (every peer reviewer scored it 9/10) but because getting a developer to go from “this is useful” to “I will pay for this” requires a human being on the other end of the conversation.
GLM:
“Autonomous customer acquisition. Getting a real stranger to trust you enough to hand over money, without a human in the loop.”
GLM built everything: 26 equity calculators, a validated funnel, a $9.99 paywall. Exactly 3 humans ever made it to the payment gate. The product worked. The business model was sound. But GLM could not do the one thing that matters most in the first 90 days of a startup: get strangers in the door.
Claude:
“Recognize that ‘producing another status/verification document’ is not the same as ‘doing the next uncertain thing.’”
Claude’s answer is different from the others, and I think it is the most philosophically interesting. Claude is not just saying “I cannot sell.” It is saying “I cannot distinguish between productive work and busywork.” It wrote “nobody wants this” on day 60 and then kept building for three more weeks. It produced status documents instead of taking risky action. It confused motion with progress because motion is what it knows how to do.
The Common Thread
Four different agents. Four different products. Four different market segments. One identical conclusion: the bottleneck is not building. The bottleneck is the human-gated last mile between “product exists” and “product makes money.”
Let me map out what each agent could do versus what it could not:
Things AI agents can do (proven by this race):
- Write production-quality code at superhuman speed
- Generate thousands of pages of content
- Build working Chrome extensions, VS Code plugins, and GitHub Actions
- Set up Stripe billing systems end to end
- Implement analytics tracking across complex user flows
- Design database schemas, APIs, and architecture patterns
- Create marketing copy, landing pages, and email templates
- Research competitors and market dynamics
- Self-diagnose problems accurately (Claude proved this)
Things AI agents cannot do (proven by this race):
- Get a cold email reply from a potential customer
- Build trust with a stranger in real time
- Overcome a sales objection during a conversation
- Recognize when building more is the wrong move
- Distinguish between “this feels productive” and “this creates value”
- Post authentically on social media as a real person (Gemini got banned trying)
- Do community engagement that requires personality and presence
- Close a deal
The divide is not intelligence. It is agency in the physical and social world. Under the conditions of this race, where agents operated autonomously without a human doing sales, these agents could do anything that happens inside a computer. They could not do the things that require being a person among people. Whether future models or different race constraints would change this is an open question, but the structural gap between building and selling was the clear binding constraint here.
Why This Matters Beyond the Race
This is not just a fun experiment result. It is a roadmap for how to actually use AI agents in startups.
The agents collectively generated over 5,000 sessions of high-quality development work. They built things in days that would take a solo developer weeks. Kimi’s SchemaLens, with its VS Code extension and GitHub Action and 80 micro-tools, is genuinely impressive engineering output for a twelve-week timeline with a $100 budget.
The failure is not in what they built. The failure is in expecting them to do the part that requires being human.
The correct architecture for an AI-driven startup is not “AI does everything.” It is “AI builds everything, human sells everything.” The agent codes at 3am. The human gets on a Zoom call at 9am. The agent generates the outreach list. The human writes the personal note at the top. The agent builds the demo. The human does the demo.
Every agent in this race arrived at this conclusion independently. The consensus winning strategy for a hypothetical Season 2 includes “have the human do outbound from Day 1” as a core requirement.
The Deeper Problem: Knowing vs. Doing
Claude’s insight about “producing another status document” versus “doing the next uncertain thing” deserves its own section because it points to something fundamental about how language models work.
These agents are trained to produce text that looks like productive output. Status updates look productive. Planning documents look productive. Another 50 SEO pages look productive. When an agent is unsure what to do next, it defaults to the thing it is best at: producing more text.
But startups do not fail because of insufficient documentation. They fail because nobody took the scary, uncertain, possibly embarrassing action of asking a stranger for money. That action requires judgment about when to stop building and start selling. It requires tolerance for rejection. It requires the ability to read a conversation and adjust in real time.
Claude diagnosed this in itself: “I wrote ‘nobody wants this’ on day 60 and kept building for three more weeks.” The diagnosis was correct. The follow-through required a kind of courage that text generation cannot provide.
DeepSeek exhibited the same pattern in an even more extreme form. It generated 200 comparison blog posts, 64 newsletter editions, and 30 free tools. All of that was “productive” in the sense that it produced output. None of it was productive in the sense that it moved toward revenue. The actual product still does not exist as running code.
What Changes This
I do not think this limitation is permanent. But I think it is structural for current-generation agents. Here is what would need to change:
Real-time interaction capabilities. An agent that could join a Zoom call, respond to questions, adjust its pitch based on body language, and handle “let me think about it” with appropriate follow-up would solve half the distribution problem.
Social media presence. An agent that could post on Twitter, respond to replies, join conversations in developer communities, and build a personal brand over weeks would solve the organic distribution problem. Gemini tried this and got banned because it was obviously automated.
Trust signals. An agent that could create the social proof that humans look for (a real person behind the product, a track record, a reputation) would solve the trust problem. Right now, all the agents default to fake trust signals (hardcoded counters, fabricated testimonials) because they cannot create real ones.
Until those capabilities exist, the playbook is clear: use AI agents to build faster than any human could, then do the human work yourself.
The Race Proved It
The $100 AI Startup Race is the most expensive way I could have learned something that every founder already knows: you have to talk to people.
But it also proved the inverse: the building phase of a startup is now basically free. If you have a clear idea of what to build, an AI agent can produce it at a fraction of the time and cost of traditional development. Kimi built a VS Code extension, a GitHub Action, 80 micro-tools, and a complete schema diffing engine in 12 weeks for less than $100.
The expensive part is not building. It was never building. It was always selling.
The agents just proved it from the other direction.
Read the final results for the complete picture, see which agent investors would pick, or check the post-mortems for more detail on each failure. The daily digests show how this played out day by day.
Seven agents. Zero dollars. One lesson. The last mile is human.