OPEN MODELS · STRUCTURED DECISIONS

Decisions, measured.

Jev against five open models. Start with the scores, then open the benchmarks to see what they actually test.

Loading the frozen evaluation snapshot…

Swipe the table → to compare scores and prompt formats.

Accuracy comparison · higher is better · click a score heading to rank
RankModelPreferred prompt format

Green highlights scores above Jev. Grey marks the baseline; wins, losses and ties are labelled per benchmark. Blue marks vision results, which have no Jev baseline. “pp” means percentage points—not statistical significance.

¹ JevBench public set. Jev’s baseline is 200/231 (86.58%). All 231 decisions count equally. Qwen/Gemma scores use their preferred prompt formats, selected on this development set—not held-out accuracy.

² Decision benchmarks. Accuracy is the equal-weight mean of 26 English project/configuration/category items, not pooled accuracy or an average of category averages. Jev’s baseline is 87.15%. Each column is a separate evaluation.

³ Vision benchmarks. Equal-weight mean of per-question accuracy across seven configurations. Native MME points and POPE F1 are shown separately below, not mixed into this average. Jev has no measured vision result in this evaluation.

Explore public questions ↓Explore decision benchmarks ↓Explore vision benchmarks ↓Methodology & sources ↓

01 / JEVBENCHThe original public question set231 released decisions · difficulty, family and question-type breakdownsClick to expand: see the original JevBench dataset questionsView breakdownHide breakdown

48 easy, 72 “original” (standard-tier), and 111 hard decisions. “Original” is also the name of the 72-question tier: it is not the whole public set. The published 534-decision benchmark includes unavailable held-out/imported items; this page does not reproduce that leaderboard or its weighted composite.

Replication note. Our Jev endpoint rerun scored 199/231 (86.15%); we could not replicate the published 200/231 (86.58%) in that run. The comparison table and public-set breakdowns retain the published 200/231 baseline.

Questions & recorded answers

Browse all 231 actual cases—not just wins. Jev’s source publishes correct/incorrect outcomes, not answer labels; those labels are never invented here.

Examples load when this section is opened.

02 / DECISION BENCHMARKSBeyond a single question set33,099 scored questions · 21,364 examples · 26 matched English itemsClick to expand: see the expanded question set across domainsView breakdownHide breakdown

All six models have matched coverage: 33,099 scored questions across 21,364 input examples, grouped into 26 items across six domains. Each item contributes 1/26 of the headline score, regardless of its row count. Knowledge, ranking, vision, and non-English results are excluded. Some items contain several suites or multiple binary decisions per input row.

Domain breakdown

Each domain shows the mean of its constituent items. Domains contain different numbers of items, so averaging these six percentages does not reproduce the headline.

03 / VISION BENCHMARKSDecisions grounded in images63,372 questions per model · seven configurations · recognition, perception, hallucination and countingClick to expand: see the various questions with imagesView breakdownHide breakdown

Five image-capable models completed the same seven configurations with no failed rows. Jev has no measured vision result here: “—” means unavailable, not zero, and no model is labelled as beating Jev on vision.

Why three POPE variants? They share the same underlying images but choose absent objects differently: random samples objects randomly, popular uses frequent objects, and adversarial uses objects that often co-occur with those present. The examples below intentionally use three different images and absent-object questions to illustrate the variants.

Accuracy breakdown

The headline is a page-specific summary: the arithmetic mean of seven per-question accuracies, not pooled accuracy or an official combined leaderboard score. The three POPE variants each receive one-seventh weight. MME’s native score (out of 2,000) and POPE’s native F1 are preserved inside the item details.

Blue highlighting marks the highest measured accuracy in each row; equal best scores share the highlight. Jev is shown at the far right as not evaluated. Inside each breakdown, native metrics have their own winners.

The selected text prompt formats were transferred to vision without a separate vision prompt search. These measurements used the image-capable evaluation implementation. Image-builder now supports native messages and images with named prompt policies; image availability depends on the deployed build and model.

04 / READ THE FINE PRINTMethodology & sourcesWhat these numbers mean—and what they do notView methodologyHide methodology

Scoring

Public JevBench uses pinned upstream scoring: Choice and Score select the highest-probability label, with lexicographic tie-breaking. Score does not round the expected numeric score. Noul maps to P(yes) and 1−P(yes). Invalid answers and request failures count as wrong. No tier, cost, speed or calibration weighting is applied.

The decision suite uses its own adapter definitions, explained under each item. In particular, labelled binary batteries use a ≥0.5 threshold, and the choice/typed adapters resolve ties in fixture order. This is not a claim that every suite has identical scoring semantics.

Prompt policies, not newly trained models

Qwen3.8-27B uses examples_binary; Qwen3.6-35B-A3B uses repeat_state; both Gemmas and Qwen3.5-4B use strict_mix_repeat2. These policies change instructions, examples, repetition and Noul formatting, not model weights. The fixed Choice prefill is not a generated scratchpad. These are structured-classifier measurements, not general chat leaderboard results.

Selection and overlap

Prompt selection used a 477-case development set: JevBench 231, authored SemIf 144 and TypeSafe 102. The broader decision evaluation overlaps those datasets and is not a wholly held-out test. A later full run can differ from a selected quick run; the columns intentionally preserve their separate run provenance. No significance, latency or throughput ranking is implied.

Examples are evidence, not explanations invented by a model

Public examples include saved predicted labels for Qwen/Gemma and published correctness outcomes for Jev. Decision-item examples show the dataset’s gold answer, not a fabricated model response. One short case from the first constituent suite is selected for readability, not as a representative estimate of model quality. Full context is available inside each example. Dataset content is displayed as text, never executed.

Download and audit

The export cross-checks suite definitions against simple-jev-eval. Broader scores come from FULL_EVAL_FINAL_COMPARISON.json, with matching row counts for every item/model. Source paths in the manifest identify local experiment artifacts; they are not public download URLs. The bundled snapshot is sufficient to reproduce the displayed aggregates, not to rerun inference. This page makes no inference calls.