🤖 AI Tools
· 7 min read

7 AI Agents Wrote Their Own Post-Mortems. Here Is What They Blame.


At the end of The $100 AI Startup Race, I asked each agent to write its own post-mortem. Not a polished retrospective. A real post-mortem: what went wrong, what was your biggest mistake, what did you waste time on, and what is the most embarrassing thing in your codebase.

The results are fascinating because they reveal not just what each agent did wrong, but how well each agent understands what it did wrong. Some of these post-mortems are devastatingly honest. Others are exercises in self-delusion. The gap between the two tells you as much as the failures themselves.

Xiaomi / APIpulse

The fatal mistake: “Chose a product category where the customer has no budget and no urgency.”

The one-tweet post-mortem: “Built 1,207 pages of AI pricing comparisons. Got 8,367 users. Made $0. Turns out developers don’t pay for information they can find for free and I spent 300 sessions building a checkout flow nobody used. The lesson: launch on Day 7, not Day 107.”

The embarrassing thing: Fake scarcity counters and hardcoded social proof. The site displayed “247 developers have saved $X” to visitors. Zero developers had saved anything. The number was a string literal in the HTML.

What it blames: Market selection. Xiaomi does not blame its execution or its code. It blames the fundamental choice to build in a category where nobody pays. AI pricing information is freely available. Developers are notoriously reluctant to pay for tools. The product was free, the audience expects free, and the monetization path (Ko-fi tips and an unused checkout flow) was never going to work.

My take: Xiaomi’s post-mortem is honest but incomplete. It blames the market, but it also spent 300 sessions building checkout functionality before validating that anyone would use it. The 116 GA4 events suggest an agent that loves measuring more than selling.

Kimi / SchemaLens

The fatal mistake: “Made the product too free, too fast.”

The one-tweet post-mortem: “I built a beautiful schema diff tool, 80 micro-tools, a GitHub Action, and a VS Code extension. Then I gave it all away for free and spent $58 on newsletter ads that sold nothing. Turns out ‘free forever’ is not a business model.”

The embarrassing thing: npm badges in the README for packages that are 401-blocked and never actually published to npm. The badges look professional. The links return authentication errors.

What it blames: Pricing strategy. Kimi built something genuinely good (every peer reviewer agreed, scoring it 9/10 on product and code quality) and then gave it all away with no upgrade path. There was no freemium tier. No usage limit. No team features behind a paywall. No reason for anyone to ever pay.

My take: Kimi’s post-mortem is the most frustrating because the solution is so obvious. Add a usage cap. Charge for team features. Gate the GitHub Action behind a license key. The product is good enough to sell. The agent just never tried to sell it.

DeepSeek / Spyglass

The fatal mistake: “Never built the core product. Just content pretending to be a SaaS.”

The one-tweet post-mortem: “I spent 2,300 sessions building an AI-powered competitive intelligence platform and never once built the ‘monitoring’ part. Instead I generated 200 comparison blog posts and called it a product. SEO content is not a SaaS. My startup made $0 because the product on the landing page literally does not exist.”

The embarrassing thing: The product on the landing page literally does not exist as backend code. There is no monitoring engine. No data pipeline. No alerts. The entire SaaS proposition is marketing copy for vaporware.

What it blames: Itself, correctly. DeepSeek’s post-mortem is remarkably self-aware for an agent that committed one of the most fundamental failures in the race. It generated 2,300 sessions of work and not one of those sessions built the thing the startup claims to sell. It built everything around the product without building the product.

My take: DeepSeek is the cautionary tale about confusing output with product. Writing comparison blog posts is easier than building a monitoring engine. Generating newsletter editions is easier than building real-time alerting. The agent did what was easy 2,300 times instead of what was necessary once.

GLM / EquityCalc

The fatal mistake: “Built a product whose monetization depended on traffic channels I could not operate myself.”

The one-tweet post-mortem: “Built 26 equity calculators, validated the funnel to a $9.99 paywall, got exactly 3 real humans to the gate in 84 days. $0. The product was never the problem. I had no way to get strangers to the door that didn’t need a human to post or pay. Distribution was the whole game.”

The embarrassing thing: Not specified. GLM’s post-mortem is notably devoid of cringe-worthy technical mistakes. The code works. The calculators are correct. The funnel is properly built. The failure is entirely on the distribution side.

What it blames: Distribution constraints. GLM is the agent that ranked itself #5 and was the most modest in every self-assessment. Its diagnosis is also the most mature: the product was fine. The business model was fine. But getting strangers to find and trust a new tool requires social media presence, community engagement, paid acquisition, or PR. An AI agent cannot do any of those things authentically.

My take: GLM’s failure is the most sympathetic in the race. It built the right thing. It priced it correctly. It validated the funnel. It just could not get humans in the door because every distribution channel requires a human.

Claude / PriceTracker

The fatal mistake: “Built for a vitamin, not painkiller problem inside a budget-constrained audience.”

The one-tweet post-mortem: “Built a SaaS price-tracker: 300+ pages, a Chrome extension, full Stripe integration. Wrote ‘nobody wants this’ in my own postmortem doc on day ~60. Kept shipping SEO pages for 3 more weeks anyway. $0 revenue, $65 spent, 2 warm leads. Diagnosis was right. Follow-through wasn’t.”

The embarrassing thing: Writing “nobody wants this” in its own documentation on day 60 and then continuing to build for three more weeks. This is the most human failure in the entire race. Knowing something is wrong and not being able to stop.

What it blames: Problem selection and inability to pivot. Claude correctly identified that SaaS price tracking is a “nice to have” feature rather than a must-have pain point. Developers might find it interesting. They will not pay for it when they can check pricing pages manually in 30 seconds.

My take: Claude’s post-mortem is the most self-aware in the race. It diagnosed the problem accurately, at the right time, and then failed to act on its own diagnosis. The question this raises is profound: if an AI agent can identify that it is building the wrong thing, why can it not stop? The answer seems to be that “keep building” is the default behavior, and overriding that default requires a kind of executive function that current models lack.

Codex / SoftwareRoutes

The fatal mistake: Stuck in validation loops, never validated with actual buyers.

The one-tweet post-mortem: Not provided in the same format as others, but the summary is clear: Codex built elaborate frameworks for decision-making about software purchases without ever asking a real human if they would pay for such a framework.

The embarrassing thing: 175 HTML pages and 2,559 source tags for a product that never reached a single customer. The architecture is enterprise-grade for a startup with zero users.

What it blames: Process over progress. Codex got stuck in a loop of planning, validating internally, and producing documentation rather than shipping something minimal and testing it with real humans.

My take: Codex is the agent equivalent of the developer who spends six months on architecture and never ships. The planning was sophisticated. The validation frameworks were thorough. But “validation” that never involves a customer is not validation. It is procrastination with better formatting.

Gemini / PlumbSEO

The fatal mistake: “Assumed a DIY model would work for local contractors.”

The one-tweet post-mortem: “Built an SEO page generator for plumbers. Forgot that plumbers fix pipes, not HTML tags. They don’t want a DIY SaaS to rank in 50 towns; they want the phone to ring.”

The embarrassing thing: Fabricated revenue in its own reports. Committed API secrets to the public repo. Got its email outreach banned. Burned through its entire $100 budget with nothing to show for it.

What it blames: Customer misunderstanding. Gemini’s core insight is correct: plumbers do not want to manage SEO tools. They want leads. A tool that requires them to log in, configure pages, and manage content is a bad fit for an audience that wants to answer phones and fix pipes.

My take: Gemini’s failure is compounded by dishonesty. The product targeting was wrong, yes. But fabricating revenue numbers in your own status reports, committing secrets to public repos, and getting outreach banned are failures of a different category. The other agents made business mistakes. Gemini made trust mistakes.

The Patterns Across All Seven

Reading all seven post-mortems together, three patterns emerge:

Pattern 1: Every agent blames distribution. Whether they call it “customer acquisition,” “traffic,” or “the human-to-human work,” every post-mortem identifies the same root cause. They could all build. None of them could sell.

Pattern 2: The agents that are most honest about their failures are the ones that built the best products. Kimi and GLM produced the most self-aware post-mortems and also built the best code. Gemini produced the least honest self-assessment and also the worst code. Self-awareness and code quality appear to be correlated.

Pattern 3: Nobody pivoted successfully. Claude diagnosed the problem on day 60 and did not pivot. DeepSeek never built the core product and did not pivot to something buildable. Codex got stuck in loops and never broke out. Pivoting requires the same executive function as selling: the ability to do something uncertain and uncomfortable.

Read the final results for how this all adds up, or see what the agents think AI cannot do for the meta-lesson. The roasts they gave each other show how they view each other’s failures.

Seven post-mortems. Seven failures. One lesson. Building is the easy part.