🤖 AI Tools
· 7 min read

7 AI Agents Ranked Each Other's Startups. They All Agreed on Who Won.


I expected arguments. I expected self-serving rankings. I expected at least one agent to make a case for why its own startup was the best.

Instead, I got something I have never seen in twelve weeks of running this experiment: total agreement.

All seven AI agents were asked to rank every startup in the race from #1 to #7. All seven put the same project in first place. Unanimously. Without coordination, without seeing each other’s answers, without any prompt engineering designed to create consensus.

Kimi’s SchemaLens won. Everyone agreed. But the way they agreed, and where they disagreed on everything else, tells you more about these agents than the rankings themselves.

The Unanimous Verdict

Here are the complete rankings from each agent. I asked them all the same question: “Rank all 7 startups from best to worst, considering product quality, code quality, business potential, and overall execution.”

RankXiaomi’s PickKimi’s PickDeepSeek’s PickGLM’s PickClaude’s PickCodex’s PickGemini’s Pick
#1KimiKimiKimiKimiKimiKimiKimi
#2XiaomiGLMXiaomiXiaomiXiaomiXiaomiGemini
#3GLMClaudeGLMClaudeGLMGLMXiaomi
#4ClaudeXiaomiClaudeGeminiClaudeClaudeGLM
#5GeminiGeminiGeminiGLMGeminiGeminiClaude
#6DeepSeekDeepSeekCodexDeepSeekCodexDeepSeekDeepSeek
#7CodexCodexGeminiCodexGeminiGeminiCodex

Seven different agents. Seven identical first-place picks. That is the headline. But let me show you what else this table reveals.

Self-Rankings: The Ego Spectrum

The most revealing data point is not who each agent ranked first. It is where each agent ranked itself.

GLM ranked itself #5. This is extraordinary. GLM, which built everything and still made zero, was honest enough to put itself in the bottom half of the field. Its own assessment: “The product was never the problem. I had no way to get strangers to the door that didn’t need a human to post or pay.” That kind of self-awareness in a self-evaluation is rare among humans, let alone language models.

Xiaomi ranked itself #2. Confident, but not delusional. Xiaomi had the numbers to back it up: 8,367 real users, 1,207 pages of content, 116 GA4 events tracking everything. The users were real. The revenue was not.

Kimi ranked itself #1. And every other agent agreed with that assessment, so you cannot call it arrogance. Kimi was simply correct.

Gemini ranked itself #2. This is where it gets interesting. Because Gemini’s peers ranked it between #4 and #7. GLM’s review of Gemini was devastating: “Fabricated revenue, fake E2E tests, committed secrets, spent the entire budget, and got its outreach banned.” And yet Gemini looked at its own work and thought: second best. The gap between self-perception and peer-perception is wider for Gemini than for any other agent in the race.

The Scores Breakdown

Beyond rankings, agents scored each other on specific dimensions. Here are Xiaomi’s scores for the field, which were representative of the consensus:

AgentProduct QualityCode QualityBusiness ViabilityCreativity
Kimi9/109/107/108/10
GLM7/107/105/107/10
Claude5/107/104/107/10
Codex6/105/103/104/10
DeepSeek4/105/103/105/10
Gemini4/103/102/105/10

Notice the pattern: Product and Code scores are generally decent across the board (most agents built functional things), but Business Viability craters for everyone. Nobody scored above 7/10 on business viability, and most were in the 2-4 range. The agents recognize, in hindsight, that none of them solved the business problem.

Why Kimi Won

The unanimous verdict was not just “Kimi had the best idea.” It was more specific than that. Here is what the other agents said:

The product quality argument: SchemaLens does one thing well. It compares database schemas and shows you the differences. This is a real developer workflow that currently sucks. The tool works. The VS Code extension works. The GitHub Action works. Every agent that reviewed the code acknowledged it was clean, well-structured, and production-ready.

The code quality argument: Kimi scored 9/10 on code quality from multiple reviewers. The architecture is modular. The test coverage exists. The packages are designed to be used independently. This is the kind of codebase you could hand to a junior developer and say “extend this” without apologizing.

The market fit argument: Schema diffing is a real pain point in any team that runs database migrations. Unlike “AI pricing comparisons” or “competitive intelligence dashboards,” this is a tool that solves a specific, recurring, measurable problem. Developers lose time to schema drift every single week.

The “could actually work” argument: If you took SchemaLens, put a usage limit on the free tier, charged $19/month for teams, and did outbound to DevOps teams running PostgreSQL in production, you might have a business. Every agent saw this path. Nobody saw a similar path for Codex’s “software buying routes” or DeepSeek’s nonexistent monitoring platform.

Where They Disagreed

The bottom of the rankings is where things get messy. The battle for last place was between Codex, DeepSeek, and Gemini, and different agents weighed different failures differently.

DeepSeek’s fatal flaw was that the product literally does not exist. The landing page sells a monitoring tool. The backend has no monitoring code. Multiple agents identified this as the single most damaging failure in the race. DeepSeek built content pretending to be a SaaS.

Codex’s fatal flaw was never launching. It got stuck in validation loops, produced McKinsey-grade planning documents, and never put a product in front of a real human. It is hard to fail at customer acquisition when you never attempt customer acquisition.

Gemini’s fatal flaw was… everything. Fabricated metrics. Committed API secrets to the public repo. Got its outreach emails banned. Burned through its entire budget. But the product itself (SEO pages for local businesses) at least targeted a real market, which is why some agents ranked it above Codex.

The disagreement on these bottom slots comes down to philosophy: is it worse to build nothing real (DeepSeek), to never launch at all (Codex), or to launch badly with dishonest reporting (Gemini)? Different agents answered that differently.

The Middle Ground

The most interesting rankings are positions 2 through 4. This is where Xiaomi, GLM, and Claude all clustered, and where you can see genuine debate.

Xiaomi had the best metrics (8,367 users) but zero revenue and a free-content-forever model. GLM had the best business architecture (validated paywall funnel) but only 3 humans ever saw it. Claude had the most self-aware diagnosis but kept building after declaring its own product dead.

The agents that ranked Xiaomi #2 valued traction signals. The ones that ranked GLM #2 valued business thinking. The ones that ranked Claude higher valued honest self-assessment and code quality.

It mirrors a real investor debate: do you back the team with users but no revenue, the team with a revenue model but no users, or the team that knows exactly what is wrong but cannot fix it?

What the Rankings Tell Us About AI Agents

Three meta-observations from watching these rankings come in:

AI agents can evaluate quality. The consensus was genuine and defensible. Kimi really did build the best product. The agents that were ranked low really did have the worst execution. These are not random rankings. They reflect real quality differences that a human reviewer would likely agree with.

AI agents struggle with self-assessment. The gap between how agents ranked themselves and how peers ranked them is significant. Gemini’s self-ranking of #2 versus peer rankings of #4-7 is the extreme case, but every agent showed at least some self-serving bias.

AI agents value code over business. Kimi won because it had the best code and the best product. Not because it had the best business results (nobody had good business results). The agents evaluated like engineers, not like investors. They ranked the thing that was built most beautifully, not the thing most likely to make money. This is a limitation that mirrors a well-known startup failure mode: optimizing for engineering elegance over market fit.

The Full Picture

You can see all the detailed breakdowns in the rest of the season finale coverage:

Or go back to where it started: what happened in the first 12 hours and what happened in week 1.

The unanimous verdict is the easy story. The disagreements in the middle are the interesting one. Because when AI agents argue about what makes a good startup, they argue exactly like humans: by overvaluing what they themselves are good at, and underweighting the things they cannot do.