Skip to main content
Verdict Developer Documentation Portal v2.5
Audited Specification & Live API Docs

Verdict AI Benchmark Documentation

Complete technical guide to Verdict's multi-judge evaluation framework, real-time Artificial Analysis market sync, BYOK key execution pipeline, and REST API routes.

Table of Contents

1. Overview & Quickstart

Verdict is an enterprise-grade AI evaluation platform designed to benchmark 500+ frontier AI models on real-world engineering tasks including full-stack code generation, interactive WebGL 3D graphics, agentic refactoring, and financial analytics dashboards.

500+ Models
Live synced from Artificial Analysis & Chatbot Arena
100 Benchmark Prompts
10 high-level prompts per engineering category
3-Judge Consensus
Audited multi-judge scoring with bias safeguards

Quickstart: Fetch Live Model Rankings (cURL)

curl -X GET "http://localhost:3000/api/leaderboard?sort=composite"

2. Multi-Judge Scoring Methodology

To eliminate single-model bias, Verdict employs an audited panel of 3 independent frontier judge models (e.g., GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro) to grade every generated artifact.

Self-Evaluation Prevention Protocol

A model candidate is strictly prohibited from evaluating its own generated output. If Candidate A is Claude Opus 5, Anthropic models are excluded from the judge panel for that run.

50% Held-Out Prompt Suite

To prevent model providers from overfitting, 50% of Verdict benchmark prompts remain private and held-out. Evaluation runs execute on a randomized blend of public and held-out test sets.

3. 5-Dimensional Rubric Weights

Judges score each evaluation sample on a scale from 0.0 to 100.0 across 5 weighted dimensions:

1. Functionality & ExecutionZero console errors, valid DOM render, working interactive controls
25% Weight
2. Code Craft & Type SafetySemantic HTML5, strict TypeScript types, modular architectural pattern
25% Weight
3. Visual Design & AestheticsModern dark-mode UI, curated HSL color palettes, responsive layouts
20% Weight
4. Micro-Interactions & MotionCanvas particle physics, spring transitions, hover feedback states
15% Weight
5. Prompt FidelityStrict compliance with all multi-constraint technical requirements
15% Weight

4. REST API Endpoint Reference

GET /api/leaderboardFetch Leaderboard Models

Query parameters: sort=composite|frontend|game|svg|agentic, openWeight=true|false.

POST /api/models/syncAuto-Sync Market Rankings

Fetches latest Artificial Analysis Quality Index, throughput speed, latency, and pricing data.

POST /api/runsLaunch Benchmark Run
// Payload body
{
  "modelId": "gemini-3-6-flash",
  "promptText": "Create a responsive HTML5 particle galaxy script.",
  "categories": ["frontend-ui"]
}

5. BYOK (Bring Your Own Key) & Security

Verdict supports full Bring Your Own Key (BYOK) execution. You can connect your private API keys for Google Gemini, OpenAI, Anthropic, OpenRouter, and DeepSeek.

AES-256-GCM Encryption Architecture

All API keys saved in Settings are encrypted at rest using AES-256-GCM authenticated encryption. Keys are decrypted strictly in-memory during benchmark execution and never written to plain-text logs.

6. Python Execution Engine & CLI Gating

For local or CI/CD automated test runs, Verdict includes a Python benchmark runner service in engine/.

# Execute local engine benchmark suite
python3 engine/main.py --config verdict.config.yaml

7. All 10 Benchmark Categories

Frontend UI10 Prompts
10 prompts for dashboards, landing pages, & accessible components.
Game Dev10 Prompts
10 prompts for 2D platformers, tower defense, & canvas shooters.
SVG Art10 Prompts
10 prompts for cyberpunk skylines, sacred geometry, & vector art.
Agentic Tasks10 Prompts
10 prompts for schema migrations, tool executors, & dependency audits.
Creative Writing10 Prompts
10 prompts for architecture specs, post-mortems, & release notes.
3D Graphics10 Prompts
10 prompts for Three.js WebGL particles, PBR shaders, & solar systems.
Data Viz10 Prompts
10 prompts for D3 heatmaps, candlestick charts, & network graphs.
Animation10 Prompts
10 prompts for Framer Motion spring timelines & SVG morphing loaders.
Full-Stack10 Prompts
10 prompts for Next.js API routes, JWT auth, & Stripe webhooks.
Code Golf10 Prompts
10 prompts for minified matrix solvers, Sudoku algorithms, & JSON parsers.

8. FAQ & Troubleshooting

Q: How frequently are market rankings updated?

Rankings auto-sync every 1 hour from artificialanalysis.ai. You can also trigger an instant manual update using the Auto-Sync Now button on the Leaderboard.

Q: What happens if an API key request times out?

Verdict automatically retries API calls up to 3 times with exponential backoff before falling back to the sandboxed runtime inspector.