Verdict AI Benchmark Documentation
Complete technical guide to Verdict's multi-judge evaluation framework, real-time Artificial Analysis market sync, BYOK key execution pipeline, and REST API routes.
1. Overview & Quickstart
Verdict is an enterprise-grade AI evaluation platform designed to benchmark 500+ frontier AI models on real-world engineering tasks including full-stack code generation, interactive WebGL 3D graphics, agentic refactoring, and financial analytics dashboards.
Quickstart: Fetch Live Model Rankings (cURL)
curl -X GET "http://localhost:3000/api/leaderboard?sort=composite"2. Multi-Judge Scoring Methodology
To eliminate single-model bias, Verdict employs an audited panel of 3 independent frontier judge models (e.g., GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro) to grade every generated artifact.
Self-Evaluation Prevention Protocol
A model candidate is strictly prohibited from evaluating its own generated output. If Candidate A is Claude Opus 5, Anthropic models are excluded from the judge panel for that run.
50% Held-Out Prompt Suite
To prevent model providers from overfitting, 50% of Verdict benchmark prompts remain private and held-out. Evaluation runs execute on a randomized blend of public and held-out test sets.
3. 5-Dimensional Rubric Weights
Judges score each evaluation sample on a scale from 0.0 to 100.0 across 5 weighted dimensions:
4. REST API Endpoint Reference
Query parameters: sort=composite|frontend|game|svg|agentic, openWeight=true|false.
Fetches latest Artificial Analysis Quality Index, throughput speed, latency, and pricing data.
// Payload body
{
"modelId": "gemini-3-6-flash",
"promptText": "Create a responsive HTML5 particle galaxy script.",
"categories": ["frontend-ui"]
}5. BYOK (Bring Your Own Key) & Security
Verdict supports full Bring Your Own Key (BYOK) execution. You can connect your private API keys for Google Gemini, OpenAI, Anthropic, OpenRouter, and DeepSeek.
All API keys saved in Settings are encrypted at rest using AES-256-GCM authenticated encryption. Keys are decrypted strictly in-memory during benchmark execution and never written to plain-text logs.
6. Python Execution Engine & CLI Gating
For local or CI/CD automated test runs, Verdict includes a Python benchmark runner service in engine/.
# Execute local engine benchmark suite
python3 engine/main.py --config verdict.config.yaml7. All 10 Benchmark Categories
8. FAQ & Troubleshooting
Q: How frequently are market rankings updated?
Rankings auto-sync every 1 hour from artificialanalysis.ai. You can also trigger an instant manual update using the Auto-Sync Now button on the Leaderboard.
Q: What happens if an API key request times out?
Verdict automatically retries API calls up to 3 times with exponential backoff before falling back to the sandboxed runtime inspector.