
JevBench
Ranking and benchmarking Jev-class decision models
Benchmark Heaven provides an independent evaluation framework and leaderboard specifically for Jev-class decision models. JevBench v1.3.0 tests 52 systems across 534 standardized decisions, helping organizations compare open-source, self-hostable, and hosted options based on intelligence, calibration, speed, and cost.
Built and run independently rather than aggregating third-party leaderboards, the benchmark solves the problem of opaque model comparisons by providing transparent, reproducible metrics. It is designed for developers, researchers, and enterprise buyers looking to evaluate and select the best decision-making models for their specific technical and cost requirements.
Key Features
Independent leaderboard for Jev-class decision models
JevBench is Benchmark Heaven's own benchmark: a state and a bounded rubric go in, a typed answer comes out. Version 1.3.0 measures 52 systems on 534 decisions (72 easy, 96 standard, 146 judge, 220 hard), scored 21 Sept 2026 under protocol jevbench::v1.2. Jev 1.13.0 leads with a score of 74.4.
Four-factor JevBench Score
Ranks systems by the geometric mean of Intelligence, Calibration, Speed and Cost, weighted 25% each. Below 50 Intelligence receives a growing near-chance penalty, so systems that barely beat coin-flipping can't rank high on cheapness alone.
Cost and speed measured, not just quality
Every entry shows sub-scores plus cost per 1,000 decisions (e.g., $0.040 for Jev 1.13.0), with clear flags distinguishing measured bills from estimated ('est.') and announced-but-uncharged ('ann.') prices. Self-hosted and demo endpoints get a documented x2 +0.15 s latency adjustment to approximate p
Reproducible methodology
One request at a time from a server in Germany; harness, public tasks and scoring rules published under MIT; results JSON released with a sha256 checksum; v1.0 results kept for comparison. Runs are built and executed by Benchmark Heaven, not collected from someone else's leaderboard.
Per-system caveats and typed categories
A legend separates closed Jev (TypeSafe), open rebuilds, instruction/JSON-schema models, small tool-calling models, rerankers and zero-shot classifiers, and closed API-only decision models. Notes disclose specifics per entry, such as djev's announced preview price, reflex 4B's development-gate discl
Filtering across providers, regions and data policies
The surrounding site lets you filter by hosting and company/lab region (China, EU, US, other), data confidentiality (whether the provider trains on or keeps your prompts), open-weights only, deprecated models, and price basis across input/output blends. EU hosting has a strict definition: inference
Comparison and cost tooling beyond the leaderboard
Benchmark Heaven also offers Benchmaxxing, Compare, Charts, Cost vs Capability, Providers per Model, a Provider explorer, Gateway and EU & Sovereign views — plus custom evals on your own data ('Need a custom eval?').
How It Works
- 1
Run the harness
Each of the 52 systems is asked all 534 decisions — 72 easy, 96 standard, 146 judge and 220 hard — one request at a time, serially, from Benchmark Heaven's own server in Germany.
- 2
Score four dimensions
Intelligence (accuracy), Calibration (whether stated probabilities hold), Speed and Cost are each scored 0-100 and combined via a 25% geometric mean into the JevBench Score; low Intelligence is penalized progressively.
- 3
Publish and document
Results, per-system methodology notes, a sha256 checksum of the results JSON and the MIT-licensed harness, tasks and scoring rules are published so entries can be inspected and reproduced.
Pros & Cons
Pros
- Runs are performed in-house from a German server rather than aggregated from third-party leaderboards, with checksummed results JSON and prior versions retained
- Balances Intelligence with Calibration, Speed and Cost, so label-only systems without probabilities and expensive-but-accurate models sort honestly
- Pricing assumptions are explicit: estimated, announced and measured costs are visually distinguished per entry
- Detailed methodology notes disclose unverified claims, dev-set usage and training-data gaps for individual entrants
- MIT-licensed harness and public tasks mean anyone can inspect or re-run the evaluation
Cons
- The site is explicitly in BETA and described as work in progress
- The x2 +0.15 s latency adjustment applied to self-hosted and demo endpoints is an assumption, not a measurement
- Most cost figures for self-hostable entries are estimated from hosted inference tariffs rather than measured bills, and some (e.g., djev) rely on announced prices that have not yet been charged
- Results describe the tested configurations, not every application, so rankings may not transfer to other workloads
- Some entries carry caveats that can't be fully independently verified — for example, Winnow-12B's private training corpus was not released and its author's audit cannot be independently reproduced
Who It's For
Best for
- Teams choosing among Jev-class decision models — open-weight rebuilds, rerankers, small tool-calling models and closed APIs — on quality, ca
- Buyers comparing self-hostable vs hosted options, or filtering by EU-hosted inference and provider data-retention policies
- Anyone who wants an independent, reproducible alternative to vendor-published benchmark numbers
Not ideal for
- General-purpose chat or coding model rankings — the benchmark targets typed-answer decision systems specifically
- Teams looking for an inference service or deployment tool; JevBench measures and ranks, it doesn't serve models
- Readers who need plug-and-play integration guidance — the page reports measurements and methodology, not implementation tutorials
Use Cases
- Evaluating open-source decision models for enterprise workloads
- Comparing hosted versus self-hostable Jev-class alternatives
- Analyzing cost-to-capability ratios for custom AI evaluations
- Benchmarking model speed, calibration, and intelligence scores
FAQ
What exactly does JevBench measure?
Jev-class decision models: each system receives a state and a bounded rubric and must return a typed answer. Version 1.3.0 tests 52 systems on the same 534 decisions, including 220 hard ones.
How is the JevBench Score computed?
It is the geometric mean of Intelligence, Calibration, Speed and Cost, each scored 0-100 and weighted 25%. Below 50 Intelligence, a growing near-chance penalty (multiplied by (I/50)^2) is applied. The weighting is adjustable on the page.
Who runs the benchmark and where?
Benchmark Heaven builds and runs it themselves — the results are not collected from someone else's leaderboard. Scoring was done 21 Sept 2026, one request at a time, from a server in Germany.
Is it open source?
The harness, public tasks and scoring rules are MIT-licensed, and results JSON is published with a sha256 checksum. Earlier results (v1.0) remain available for comparison. Individual ranked systems range from open-weight to closed API-only.
How are costs handled for self-hosted or unpriced models?
Entries marked 'est.' have no measured bill and are priced like a large inference provider; 'ann.' means the provider's announced price was used but nothing has been charged yet, as with djev's free-preview endpoint.
Why is classifier.dev's 83.6 score not ranked first?
It is shown as an honorable mention rather than ranked because it runs another entrant's model. Partial runs (Qwen3.8 27B, Needle 3 variants) are likewise displayed separately from the full ranking.
Can I get the benchmark run on my own data?
Yes — the page solicits custom evals ('Need a custom eval on your data?'), though no pricing for that service is stated on this page.