Skip to main content
new · Live HTML Preview Rendering Evals active

The World's #1
Vibe Coding Benchmark

We built the premier evaluation platform for AI models. 57+ frontier models, head-to-head scoring, and an AI judge that grades it all. No vendor hype. Just raw, real results.

4,400+ prompts judged·57+ frontier models·10 vibe coding categories
~/projects/acme · Verdict AI Benchv0.9.3
$ bench run --model claude-opus-4-5 --suite full
→ resolved 14 benchmarks · 18,540 graded samples
→ concurrency=64 · est. cost $42.18 · est. time 14m
✓ authenticated · org acme · cluster us-east-2
spawning 64 workers...
✓ mmlu-pro · 12,032 / 12,032 · 88.4
✓ humaneval+ · 164 / 164 · 92.7
✓ swe-bench · 500 / 500 · 71.2
! aime-2025 · 30 / 30 · retried 2 timeouts
✓ aime-2025 · 30 / 30 · 76.1
✓ gpqa · 198 / 198 · 64.8
✓ tau-bench · 500 / 500 · 60.5
→ 3 disagreements flagged for human review
✓ run complete · 14/14 benchmarks · composite 79.4
→ report saved to runs/run_a7f2c.json
$ _▋
57+
Frontier Models Benchmarked
85.2
Top Composite Score
4.4K
Prompts AI-Judged
10
Vibe Coding Categories
Anthropic logo
Claude Fable 5·85.2
OpenAI logo
GPT-5.6 Sol·82.4
Moonshot logo
Kimi K3·80.9
OpenAI logo
GPT-5.6 Terra·78.4
OpenAI logo
GPT-5.5·77.7
xAI logo
Grok 4.5·77.4
Anthropic logo
Claude Opus 4.8·77.6
Google logo
Gemini 3.6 Flash·72.3
Qwen logo
Qwen 3.7 Max·71
DeepSeek logo
DeepSeek V4 Pro·69.5
Anthropic logo
Claude Sonnet 4.6·69.7
OpenAI logo
GPT-5·70
Anthropic logo
Claude Fable 5·85.2
OpenAI logo
GPT-5.6 Sol·82.4
Moonshot logo
Kimi K3·80.9
OpenAI logo
GPT-5.6 Terra·78.4
OpenAI logo
GPT-5.5·77.7
xAI logo
Grok 4.5·77.4
Anthropic logo
Claude Opus 4.8·77.6
Google logo
Gemini 3.6 Flash·72.3
Qwen logo
Qwen 3.7 Max·71
DeepSeek logo
DeepSeek V4 Pro·69.5
Anthropic logo
Claude Sonnet 4.6·69.7
OpenAI logo
GPT-5·70
01 · the leaderboard

The whole frontier, ranked by what you actually ship.

Sort by composite score, narrow to coding or reasoning, filter by price or open-weight. Every cell links to the exact test cases — no black boxes.

verdict · leaderboardlive · 4,400+ prompts scored
#ModelComposite
01
Anthropic logo
Claude Fable 5
Anthropic
85.2
02
OpenAI logo
GPT-5.6 Sol
OpenAI
82.4
03
Moonshot logo
Kimi K3
Moonshot
80.9
04
OpenAI logo
GPT-5.6 Terra
OpenAI
78.4
05
OpenAI logo
GPT-5.5
OpenAI
77.7
why Verdict Bench

Numbers you can defend in a design review.

Built by a team that got tired of inconsistent evals, vendor-curated benchmarks, and "trust us, it's better." Three principles, no exceptions.

01CAN I RUN

Instant model compatibility checks.

Paste any model ID and we'll tell you which benchmarks it can run, what it'll cost per prompt, and where it might fail.

$ bench can-i-run claude-opus-4
provider: anthropic
categories: 10/10 supported
est. cost: $0.84 / prompt
✓ ready to benchmark
02PROMPT LIBRARY

3,900+ vibe-coding prompts.

The largest curated prompt set for evaluating frontier models. 10 categories — battle-tested across 15+ models.

Categories3,900+ prompts
Frontend UI
Game Dev
SVG Art
Creative
Agentic
03AI JUDGE

Multi-agent scoring that matches human consensus.

Our AI judge panel uses weighted rubrics to grade functionality, design, code quality, and creativity. 5 dimensions. Every score is auditable.

judge-panel gemini-3.1-pro
rubric: v3.2 · 5 dimensions
✓ functionality94/100
✓ design91/100
✓ code quality88/100
✓ creativity92/100
composite:91.3
THE PLATFORM

Every model. Every category. One dashboard.

Run any frontier model against our 3,900+ prompt library. Get AI-judged scores in 10 vibe-coding categories, compare results head-to-head, and share benchmarks to the community showcase.

  • —Run benchmarks across 16+ frontier models — OpenAI, Anthropic, Google, xAI, DeepSeek
  • —AI-powered judge panel scores every response on 5 weighted dimensions
  • —Head-to-head arena mode for direct model comparisons
  • —Community showcase — share & browse the best AI-generated creations
●bench.config.yaml14 lines
# Production model gate — run on every PR
suite: code-suite-v2
fail_if:
composite: < 82.0
swe-bench: < 68.0
latency_p50: > 2.0s
 
models:
- id: claude-sonnet-4-5
- id: gpt-5-mini
- id: deepseek-v3-2
 
benchmarks:
- humaneval+
- swe-bench
- livecodebench
- custom: ./internal/refactor-suite
 
grader: llm-judge
budget: $25

No Vendor Bias

Provider marketing benchmark claims are often cherry-picked. Verdict runs independent multi-judge scoring panels with auditable logic.

Auditable Judge Panel

Every score includes written rationale across Functionality, Craft, Design, Creativity, and Fidelity. Disagreements trigger human review.

Free & BYOK Architecture

Zero subscription fees. Bring your own API keys to run custom evaluation suites, or deploy locally via Docker Compose.

Start benchmarking free today.

Join thousands of developers who use Verdict to pick the right model for every task. No credit card required.

$docker compose up -d