Skip to main content
Back to Leaderboard
N

Hermes 4 - Llama-3.1 70B (Non-reasoning)

Open Weight

Provider: Nous Research • ID: hermes-4-llama-3-1-70b

Detailed evaluation metrics and judge breakdown for Hermes 4 - Llama-3.1 70B (Non-reasoning) across all 4 launch categories.

Compare in Blind Arena
Pricing: In $0.13 per 1M tokens • Out $0.40 per 1M tokens
Composite Verdict Score
6.9
0
.
0
/100

Category Score Breakdown

Frontend UI7.0

Interactive dashboards and landing pages

Game Dev6.8

Canvas 2D physics and browser games

SVG Art7.0

Vector graphics and generative math art

Agentic Tasks6.7

Multi-step execution planning

Auditable Multi-Judge Dimension Scoring

Grades evaluated across 3 independent judge models using published rubrics.

Functionality— Executes without runtime JS exceptions
7.1
Craft— Clean semantic markup and clean TypeScript
7.0
Design— Visual aesthetics, contrast, layout harmony
6.8
Creativity— Original micro-interactions and layout depth
6.7
Fidelity— Strict adherence to complex prompt constraints
6.8