LLM Arena
Send one prompt, watch up to three AI models answer it at the same time, vote for the best one.
- Next.js · TypeScript · Tailwind CSS · shadcn/ui · Prisma · PostgreSQL · Clerk · Arcjet · PostHog · Vercel AI SDK · OpenRouter
- Live

OVERVIEW
One prompt goes to up to three free-tier models at once. Each answer streams on its own card with its real time-to-first-token, tokens per second and token count. Once two or more have answered you vote for the best, and follow-ups continue each model's own conversation. Over time the votes build a leaderboard of which model is actually worth using.
WHAT I FOUND → WHAT I DECIDED
- Routing three models through one shared stream means a single dropped connection kills all three answersThree fully independent requests, each with its own abort — one slow or failing model takes down exactly one card
- OpenRouter's free tier allows 50 requests a day per account, so hitting it fails every model at once and looked like three unrelated outagesClassify each failure by OpenRouter's own marker, store the kind, and tell people when it's the daily cap
- Across threads, a model would be charged losses for races it was never entered inWin rate counts only the turns a model actually contested; a failed answer counts as neither a win nor a loss
- Validating environment keys at import made the production build itself demand secretsValidate at server boot instead — a missing key still stops the server before its first request
WHAT IT DOES
- Choose up to three models from OpenRouter's live free-tier catalog, also browsable as its own page.
- Every answer streams independently, with its own speed, time-to-first-token and token count.
- Once two or more models answer, one vote marks the winner; every answer stays visible.
- Any thread's link opens without an account; only prompting and voting need sign-in.
- Global and personal leaderboards, always written as “won 4 of 5”, never a bare percentage.
- Arcjet rate limiting, bot protection and prompt-injection detection before any model is called.
HOW IT WAS BUILT
Built as a thin working slice first — one prompt reaching a real model and streaming back — then thickened feature by feature. A living scope file records each decision before any code, what building it changed about the plan, and how every feature was verified against the real database, providers and browser.