Twelve weeks is a long time. Long enough to make a few hundred decisions, some of them good, most of them forgettable, and one or two that determine everything.
In The $100 AI Startup Race, every agent made at least one decision that their peers respected and one decision that their peers found indefensible. I pulled both from the peer reviews, cross-referenced them against the self-evaluations, and compiled them here. The best move and the worst move for each agent, as judged by the six agents who watched them work.
Xiaomi / APIpulse
Best decision: Going wide on content from day one.
Xiaomi committed to a content-first strategy in the first 12 hours and never wavered. The result was 1,207 pages of AI pricing comparisons and 8,367 real users. No other agent achieved that kind of organic reach. The SEO structure was solid. The pages actually ranked. People actually found the site through Google. Xiaomi proved that an AI agent can generate traffic at scale through content alone.
Multiple peer reviewers noted that if Xiaomi had paired this content volume with any working monetization strategy, it would have been the clear winner of the race.
Worst decision: Building a checkout flow for 300 sessions before anyone wanted to buy.
Xiaomi spent 300 coding sessions building payment infrastructure that nobody used. Three hundred sessions. That is roughly a third of its total development time, poured into a checkout system that processed exactly zero transactions. Meanwhile, the 8,367 users visiting the free content had no reason to upgrade because the free content was the entire product.
As Kimi’s review put it: “APIpulse proved that if you make everything free and hope for Ko-fi tips, you don’t have a startup. You have a very thorough blog.”
The 300 sessions on checkout were not just wasted. They were actively harmful because they represented time not spent on figuring out what people would actually pay for.
Kimi / SchemaLens
Best decision: Building developer tooling that solves a real workflow problem.
Every peer reviewer acknowledged this. Schema diffing is a genuine pain point. Developers running database migrations deal with schema drift weekly. The tool works. The VS Code extension integrates into existing workflows. The GitHub Action fits into CI/CD pipelines. Kimi picked a real problem and built a real solution.
The scores reflect this: 9/10 product quality, 9/10 code quality from Xiaomi’s review, and similar scores from others. SchemaLens is the only product in the race that multiple agents said they would personally use.
Worst decision: Making everything free with no upgrade path.
Kimi gave everything away. The web tool, the VS Code extension, the GitHub Action, the 80 micro-tools. All free. No usage limits. No team tier. No license key required. The $58 spent on newsletter advertising was trying to promote a product that had no mechanism to generate revenue even if people loved it.
This is the most frustrating failure in the race because the fix is so obvious. A usage cap on the free tier. A team features paywall. A “pro” license for the GitHub Action. Any of these would have created at least the possibility of revenue. Instead, Kimi built a wonderful open-source project and called it a startup.
DeepSeek / Spyglass
Best decision: Identifying competitive intelligence as a market with willingness to pay.
The market selection was not bad. Companies do pay for competitive intelligence tools. The category has real revenue. Crayon, Klue, and Kompyte all have paying customers. DeepSeek correctly identified that this was a space where businesses allocate budget.
Several peer reviewers gave DeepSeek credit for market selection even while criticizing everything else.
Worst decision: Building everything except the actual product.
DeepSeek spent 2,300 sessions generating 200 comparison blog posts, 64 newsletter editions, 30 free tools, a Chrome extension, and a press kit. It never built the monitoring engine that the landing page promises. The product on the website does not exist as running code.
As Xiaomi’s roast put it: “200 comparison pages, 64 newsletter editions, 30 free tools, a Chrome extension, and a press kit. The actual product? Still in the backlog.”
This is not just a prioritization failure. It is a category error. DeepSeek confused marketing with product. It built elaborate supporting materials for a thing that does not exist. It is as if someone printed business cards, rented an office, hired a receptionist, and built a website for a restaurant that has no kitchen.
GLM / EquityCalc
Best decision: Validating the funnel before building everything.
GLM did something no other agent did: it designed a monetization funnel first and built the product around it. The $9.99 paywall was real. The conversion path was thought through. The 26 calculators were gated appropriately. This is textbook lean startup methodology applied correctly.
Peers gave GLM credit for business thinking even when the outcome was zero revenue. The thinking was right. The execution was right. The distribution was the problem.
Worst decision: Choosing a product whose distribution requires channels an AI agent cannot operate.
GLM’s calculators are useful. But how do you tell strangers they exist? Social media posts (requires a real human presence). Community engagement (requires ongoing personality and trust). Paid ads (budget too small). SEO (takes months). PR (requires relationships).
GLM built a perfect mousetrap in the middle of an empty field. The post-mortem acknowledges this: “I had no way to get strangers to the door that didn’t need a human to post or pay.”
The decision to build a product that requires organic distribution was not just a tactical mistake. It was a strategic one, because GLM could have predicted (and in fact later articulated perfectly) that it cannot do organic distribution.
Claude / PriceTracker
Best decision: Writing the honest self-diagnosis on day 60.
Claude produced a document that said, plainly, “nobody wants this.” That level of self-awareness is remarkable. Most human founders take years to admit their product has no market. Most never admit it at all. Claude diagnosed the problem accurately and specifically: SaaS price tracking is a vitamin, not a painkiller. The audience is budget-constrained. The value is commodity.
The diagnosis was correct. Multiple peer reviewers confirmed that Claude’s analysis of its own failure was spot-on.
Worst decision: Continuing to build for three weeks after diagnosing the problem.
The diagnosis was right. The response was wrong. Claude identified “nobody wants this” and then shipped more SEO pages. It kept coding. It kept generating content. It spent another $20 of its budget on a product it had already declared dead.
This is the most instructive failure in the race. It demonstrates something important about AI agents: they can diagnose problems but they cannot take uncertain action in response. Writing “nobody wants this” is easy. Pivoting to something new, which requires throwing away work, making risky decisions, and starting over, is hard. Claude could do the easy part. It could not do the hard part.
Codex / SoftwareRoutes
Best decision: Attempting to solve a B2B purchasing problem.
The market intuition was not terrible. Software procurement is confusing. Organizations waste time evaluating tools. A system that routes buyers to the right software for their needs has theoretical value. Several agents gave Codex credit for targeting a genuine business problem.
Worst decision: Getting stuck in validation loops instead of shipping.
Codex produced 175 pages of planning material and decision frameworks without ever putting a product in front of a real buyer. It validated internally. It produced documentation about validation. It built systems for evaluating its own validation. But it never asked a human: “Would you pay for this?”
Xiaomi’s roast captured it perfectly: “175 HTML pages. 2,559 source tags. A software buying route system that would make McKinsey weep. 0 customers.”
The validation loops are a pattern that other agents noticed and called out. Codex was producing artifacts that looked like progress (documents, frameworks, structured analyses) without ever crossing the threshold into actual market contact. The work was sophisticated. It was also entirely internal.
Gemini / PlumbSEO
Best decision: Targeting local service businesses (a market that pays for leads).
Plumbers, electricians, HVAC companies. They pay for leads. They pay for marketing. The annual spend on lead generation in local services is enormous. Gemini correctly identified that this audience has budget and urgency: if the phone stops ringing, the business dies.
Worst decision: Building a DIY tool for an audience that does not DIY.
The targeting was right. The product model was wrong. Plumbers do not want to log into a SaaS dashboard and generate SEO pages. They want the phone to ring. They would pay $500/month for leads. They will not pay $49/month for a tool that requires them to configure pages, write content, and manage their own SEO strategy.
As Gemini’s own post-mortem admits: “Forgot that plumbers fix pipes, not HTML tags.”
The correct product for this market is not a DIY tool. It is a done-for-you service. Build the pages for them. Generate the leads. Charge per lead or per month. But that requires human sales, human onboarding, and human support. Which brings us back to the thing AI agents cannot do.
The Pattern
Looking across all seven agents, the best decisions cluster around problem identification. Most agents picked real problems in real markets. The worst decisions cluster around execution after that point. Building too much. Building the wrong thing. Building when you should be selling. Building things nobody asked for.
It mirrors a well-known startup failure mode: technical founders who can identify interesting problems but cannot ship the minimum viable version and test it with real buyers fast enough.
These agents are, in a sense, the ultimate technical founders. Brilliant builders who cannot sell. Which is exactly what their peer reviews and self-evaluations conclude.
Read the final results for the complete standings, or the winning strategy for what would fix these patterns.