SlopScore
10 crowdincl. 1 critic

jevbench

JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.
Open repo on GitHubgithub.com/fstandhartinger/jevbench
Python · ★ 148 · 15 forks · MIT · paperwork by the Cap'mmostly ai (inferred)light human (inferred)works-on-my-machine (inferred)other
listed 1 hour ago by fstandhartinger · last checked 3 minutes ago
The owner didn't write this. This repo never submitted itself. The Cap'm found it on a truffle trawl and wrote its paperwork from what GitHub already shows. Picked by hand by the Cap'm on 2026-09-28: JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.; its own README says "jsonl , 109 held out), written by Claude Opus 5 and GPT-5". 143 stars; MIT license. The owner did not submit this. Votes count; awards don't until the owner claims it.

I'm not calling your project slop! Geeze, it's a joke... Do you own this repo?

Log in with GitHub as fstandhartinger. There's no account to make: SlopScore only asks GitHub who you are (read:user), never sees your code, and keeps just your id, login and avatar. Then you can:

  • Keep it, on your terms. Commit your own slopscore.md (spec) and press Refresh. Your paperwork replaces the Cap'm's, and you can submit it for Slop of the Day.
  • Take it down. One click on Remove. It stays gone; the trawl never brings it back.

Log in with GitHub

Can't log in as the owner? Request a takedown. No login needed, and a trawled listing comes down right away.

GitHub says
JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.
created
2026-09-19 · pushed 1 hour ago · 69 commits · 7 contributors
release
v1.4.2 · 2026-09-24
languages
Python 100%Shell 0%
paperwork
licensereadme 42% health
dependencies
no dependency graph (no manifest, or disabled) · OSV.dev, checked 1 hour ago

Disclosures, inferred by the Cap'm

slopbucket
vibe-coded
category
other
ai_generated
mostly
human_touch
light
status
works-on-my-machine
language (detected)
pythonshell
license (detected)
mit

The Cap'm's log

The Cap'm wrote this paperwork, not the owner. This repo never submitted itself to SlopScore. The Cap'm picked it by hand: JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.; its own README says "jsonl , 109 held out), written by Claude Opus 5 and GPT-5". It carries the MIT license. The disclosures above are his best guess from what GitHub shows.

Is this yours? Commit a real slopscore.md and press Refresh to replace this, or remove the listing in one click. There's no account to make: you log in with GitHub.

README — the repo's own words, folded up so the grading fits on one screen

JevBench

Combination experiments (confidence cascades, committees, and real-sample best-of-n) are reported in RESULTS-COMBINATIONS.md. None changed the ranked board.

A benchmark for Jev-class decision models: you hand the model a piece of state and a bounded rubric, and it hands back a typed answer, ideally with a probability for every option. No prose, no parsing, no "as an AI language model".

JevBench is Benchmark Heaven's own benchmark. It is not affiliated with or endorsed by TypeSafe AI, whose Jev model is one of the systems measured here.

v1.4.2.2: Imajev-4B leads; Plumb-4B is #2 (current)

Live board · v1.4.2.2 release notes · v1.4 method · aggregate results · changelog

The v1.4.2 scorer is unchanged. v1.4.2.2 adds Imajev-4B, measured on the full v1.4 protocol with one rotation and calibration.json, to the exact v1.4.2.1 result. The board has 95 systems, 91 ranked; prior measurement and score fields remain unchanged, with ranks moving only where the added row changes the order. Only aggregate sealed statistics are published.

Rank System JevBench Score
1 Imajev-4B 67.37
2 Plumb-4B (crh225, JevK5 v0.2 + LoRA) 65.84
3 decider-4b v2 (Mapika) 64.13
4 Jev 1.13.0 (TypeSafe AI) 63.29
5 JevK5 v0.2.0 62.04

Imajev-4B's estimated Cost uses the public DeepInfra Qwen/Qwen3.5-4B reference price and zero generated output tokens. It is a model-price estimate, not a GPU bill.

v1.4.2.1: Plumb-4B added (previous release)

Live board · v1.4.2.1 release notes · v1.4 method · aggregate results · changelog

The official score blends 20% chance-corrected Intelligence from 308 fresh sealed decisions with 80% of the v1.3.0 Intelligence axis. Calibration is blended toward the sealed-inclusive measurement. The four axes use an equal-weight harmonic mean. A public-to-sealed accuracy gap above 25 percentage points reduces Intelligence; the existing low-Intelligence penalty remains, and Speed and Cost each receive a quadratic gate below 50.

Rank System JevBench Score
1 Plumb-4B (crh225, JevK5 v0.2 + LoRA) 65.84
2 decider-4b v2 (Mapika) 64.13
3 Jev 1.13.0 (TypeSafe AI) 63.29
4 JevK5 v0.2.0 62.04
5 Cygnet (blockbrain, frozen Gemma-4-12B-it) 61.76

Plumb-4B takes the highest equal-weight composite. Jev 1.13.0 out-reasons decider-4b v2 (Intelligence 53.1 vs 49.4) and is better calibrated; decider-4b v2 leads on speed and cost. JevBench weighs the four axes equally; sort by Intelligence for raw reasoning.

v1.4.2.1 adds Plumb-4B to the live v1.4.2 baseline: 94 systems, 90 ranked. The v1.4.2 scorer and prior measurements are unchanged; ranks move only where the added row changes the order. Plumb's estimated Cost uses the bookable EmpirioLabs Qwen3.5-4B input rate of $0.04/M, not a GPU bill. Only aggregates are published for the sealed items. Rows with an API flag identify operator endpoints that received item text without answer keys. The public half can be trained on or selected against, so the held-out private part will need to evolve as the field changes.

v1.4.2: eleven new systems and swanOne's sealed run

v1.4.2 release notes · aggregate results

v1.4.1: six completed additions to the v1.4 board

v1.4.1 release notes · aggregate results. The previous top five was Jev 1.13.0 (63.29), JevK5 v0.2.0 (62.04), Hopper (59.43), Winnow-12B Q8 (55.58) and reflex 4B (53.99).

v1.3.0: previous scoring release

Results -> RESULTS-v1.2.md · artifact results/v1.2/jevbench-v1.2-results.json · interactive: benchmarkheaven.com/jev-models · how the hard tier was made: datasets/HARD-TIER.md

JevBench Score = chance-corrected Intelligence, Calibration, Speed, Cost — 25 % each, geometric mean. Below 50 Intelligence, multiply the score by (Intelligence / 50)².

Axis Score 0-100
Intelligence (accuracy - chance) / (1 - chance) per tier, clipped at 0; hard 30 %, easy 14 %, standard 28 %, judge 28 %
Calibration hard tier: ECE + fidelity to exact gold distributions (label-only systems: none, counts as 0)
Speed mean of score(p50), score(p95); score(s) = 100 - 20 log10(s / 0.1 s): 0.1 s = 100, each 10x slower -20
Cost 100 - 30 log10($ per 1,000 decisions / $0.001): $0.001 = 100, each 10x more expensive -30

The Cost column is US dollars per 1,000 DECISIONS, not per 1,000 tokens. One decision is a whole question: its state, its rubric and its options — hundreds to thousands of input tokens. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.

Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in RESULTS-v1.2.md and the artifact. Why, and its limits: Limits, stated plainly.

  • Accuracy by subject topic (math, coding, rules & law, finance, support & operations, everyday language, safety & security): datasets/TOPICS.md and results/v1.2/jevbench-v1.2-topics.json — aggregates only, not part of the score.

  • 220 hard decisions (111 public in datasets/public/hard.jsonl, 109 held out), written by Claude Opus 5 and GPT-5.6 Sol, cross-reviewed, frozen and hashed before any system ran; 534 decisions per system in total.

  • Top of the ranking (48 ranked rows): Jev 1.13.0 (TypeSafe AI) 74.4 · SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) 73.1 · djev (Maisa, diffusion-gemma) 73.0 · Winnow-12B Q8 71.2 · reflex 4B (kshetrajna12) 70.3. Qwen3.8 27B and Needle 3 (both modes) are partial runs, shown without a rank.

  • djev is self-hostable: Davipar/djev-dev is Apache-2.0 code over Google's Apache-2.0 DiffusionGemma weights. It is an inference method, not a separately trained model, and adds no weights of its own. djev-spark is another DiffusionGemma structured-read runtime, not another model row.

  • Honorable mention, not ranked: classifier.dev (fast tier) 83.6. A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models — its fast tier is Jev ("The fast tier is Jev, TypeSafe's decision model", classifier.dev/benchmark), so ranking it would rank Jev's model against Jev's model at a different price. See Honorable mentions.

  • Scoring code: jevbench/composite_v13.py; the final artifact is rebuilt from the frozen measurements by scripts/v1.2/finalize.py, charts by scripts/v1.2/charts.py. Per-task outcomes (public items): results/v1.2/jevbench-v1.2-per-task.json.

  • open-alternative-jev is ranked with the author's own option order (A. yes, B. no). With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on answer-judging items — small models are very sensitive to option order. Both runs: results/v1.2/runs/open-alternative-jev/.

Other mentions and diagnostics

  • stuntdouble is an independent shadow-proxy and replay companion that can import JevBench's public hard items to compare local decision models. It is a tool, not a JevBench entrant; its example report is not an official JevBench result.
  • The option-order robustness analysis in issue #40 is a separate diagnostic on public choice items. It does not change the canonical task order or official scores.

JevBench v1.3.0 — JevBench Score

What changed in the score

A system that is cheap and fast but barely better than guessing could rank high; intelligence is now measured above chance, and systems below half-way get a growing penalty. The task set, Calibration, Speed, Cost and who is eligible for a rank are unchanged.

Revision log of v1.2 (19 Sep 2026; items and answers never changed after the freeze): v1.2-wip (tag v1.2-wip) hard tier + calibration sub-score, Balanced 33:33:33 Main Score with hard 50 % of Capability; v1.2 final (tag v1.2): 4 axes, geometric mean — Intelligence (hard 30 %), Calibration, Speed, Cost at 25 % each; Speed 20 points and Cost 30 points per decade; latency of non-production endpoints adjusted ×2 (+0.15 s on our own servers, an assumption); one open-alternative-jev row (author's option order); Needle 3 options-as-tools priced on Needle 3's per-token basis ($0.0162 est.). The earlier weightings are kept as views, recomputed the same way, and are not the JevBench Score. v1.2.1 (tag v1.2.1): added djev (Maisa's diffusion-gemma Jev implementation, api.djev.dev) — all 534 decisions including the held-out items, one request at a time through its production API (no latency adjustment), scored with the unchanged v1.2 rules. Cost uses djev's announced price ($0.035 per million input tokens, output free), which is not charged yet (free preview). Adapter: jevbench/adapters/djev.py; row: results/v1.2/additions/djev.json. No other row changed; ranks below #2 move down one.

v1.2.2 (tag v1.2.2): five systems readers asked for — Laya (Convai Innovations, ModernBERT-large 421M), jeff (Logan Markewich, GLiFormer 400M), GLiNER2 (Fastino, gliner2.5-base), openJev Verdict (heman10x, 151M) and classifier.dev (fast tier). Each ran all 534 decisions including the held-out ones, with the unchanged v1.2 scoring. The four open systems ran on our CPU (4 threads) and carry the usual ×2 + 0.15 s latency adjustment; classifier.dev is a production API and carries none. Every mapping — above all GLiNER2's label scores → one distribution — was written down before the runs: docs/v1.2-additions.md. Adapters: jevbench/adapters/ (laya_local, gliner2_local, verdict_local, classifier_dev; jeff uses the existing typesafe adapter against its own server). Rows: results/v1.2/additions/. No other row changed. ProgramAsWeights was also requested and is prepared (paw_local), but its hosted compiler only keeps a program private for a signed-in account, and compiling 223 held-out rubrics into public programs would publish them; it is therefore not in this revision.

v1.2.3 (tag v1.2.3): cost correction. Every row's price was recomputed so that each of the 534 decisions is counted once and priced once. Three arithmetic mistakes were fixed — the 242-decision standard+judge run was averaged twice in the v1.1-tier price (556 rows instead of 314); rows priced from the gemini-3.1-flash-lite token counts used that run's standard+judge-only average (452 input tokens per decision) for all 314 v1.1 decisions instead of its average over all 314 (383); and requests whose answer came back unparseable were left unpriced although the provider billed them (9 DeepSeek V4.1 Flash decisions). No tariff was wrong, no measurement, item or answer changed, and no rank changed. Fifteen rows become 1.5–11 % cheaper (JevBench Score up by at most 0.2 points); DeepSeek V4.1 Flash becomes 2.6 % more expensive. The correction also spells the unit out everywhere: the Cost column is dollars per 1,000 decisions, never per 1,000 tokens. Per-row figures and their derivation: results/v1.2/cost-correction-v1.2.3.json; check: tests/test_cost_correction.py.

v1.2.4 (tag v1.2.4): classifier.dev leaves the ranking and becomes an honorable mention. New general rule: a service that runs another entrant's model is listed, but not ranked against the models. classifier.dev is not its own model — its own pages say "The fast tier is Jev, TypeSafe's decision model" (classifier.dev/benchmark, read 20 Sep 2026) and the API answers with "model": "jev-1.13.0" — so ranking it put the same model in the list twice, once at TypeSafe's per-token tariff and once at classifier.dev's flat plan. What it adds is that price and, on its smart tier (which we did not measure), an orchestration layer: "The smart tier is Jev plus a reasoning model re-asking only the answers Jev put under 0.7 confidence" — escalation on low confidence, a model cascade, not best-of-N, self-consistency or a committee. Its row keeps every number, axis, cost basis, radar and per-task outcome and carries no rank; Jev 1.13.0 is #1 and every other row moves up one place. No measurement and no score changed. The full note, with the price caveat and the one place where the fast tier measurably differs from Jev, is in RESULTS-v1.2.md; check: tests/test_honorable_mentions.py.

v1.2.5 (tag v1.2.5): the kev family. Added kev 0.5B and its 0.6B, 4B and 8B research previews. Each ran all 534 frozen decisions through the author's native TypeSafe-compatible endpoint in BF16 on one RTX 3090, with zero failed requests. kev 0.6B enters highest at #9 with 66.7; 0.5B is #15 with 63.1, 4B #16 with 62.2, and 8B #18 with 58.3. The larger checkpoints improve Intelligence but lose points on calibration and estimated hosted cost. No earlier row changed. Full method: docs/v1.2-additions.md; rows: results/v1.2/additions/.

v1.2.6: Added openJev Verdict 1.4 as its own row (same weights, fixed author engine) and the reachable, identified SimpleJev public-demo configurations. Every entrant ran the unchanged 534 frozen decisions including the hard tier; earlier Verdict and all other rows remain unchanged. Full method and endpoint conditions: docs/v1.2-additions.md.

v1.2.7 (tag v1.2.7): two more GLiNER2 checkpoints, and a submitted endpoint measured on the public items. Added GLiNER2.5 small at 62.1 (#21) and GLiNER2.5 multi at 63.1 (#19) — both ran all 534 frozen decisions on our CPU with the same mapping as the existing GLiNER2 row (gliner2.5-base), re-checked on each checkpoint before its run. Also added jqv at 67.2, a stock Qwen3-32B read as a decision model, submitted with a public endpoint in issue #6: it is a partial row, shown but not ranked, because its endpoint runs on the submitter's own machine and this round stopped sending held-out items to an endpoint a submitter operates. It answered 425 of 534 decisions — everything except the 109 held-out hard items — and on the public items it reproduced the submitter's own numbers exactly (easy 1.000, standard 0.958, hard 0.622). Mappings, endpoint conditions and cost bases were committed before any row was aggregated and before the published GLiNER2.5 runs started (docs/v1.2-additions-run3.md); jqv's run had begun about ten minutes earlier, because its endpoint was temporary, but it needs no mapping and is priced at its base model's public tariff. No other row changed.

v1.2.8 (tag v1.2.8): 10 more requested systems, and jqv complete. Every entrant ran all 534 frozen decisions through its author's own server, one request at a time, scored with the unchanged rules: reflex 4B 71.7 (#5), decision-machine-1 71.5 (#6), jqv 70.1 (#8), decider-35b-a3b 68.9 (#10), OpenDecision 67.0 (#14), decider-2b 64.6 (#20), reflex-27b 64.2 (#22), jev-local 63.8 (#23), LitJev 63.7 (#25), Bespoke Nimble 9B (re-run) 61.8 (#30), GLiNER2 large 50.5 (#36). decision-machine-1 is a closed decision model behind milliseconds.ai's production API and gets its own class ("closed decision model"); the rest ran on our RunPod GPUs, GLiNER2 large on our CPU. jqv (issues #6 and #9) was re-run in full on our own GPU from its now-public serving code, so its v1.2.7 partial row is replaced by a complete, ranked one. Bespoke Nimble 9B was re-run after Bespoke Labs raised its prompt limit to 8,192 tokens: hard tier 43.6 % → 65.5 %, yet its score fell from 63.7 to 61.8, because the long items that used to fail are now answered and priced, and its pod was farther from our server. Every GPU pod this round was in Canada, so those rows' Speed includes a transatlantic network path from Germany. Mappings, endpoint conditions and cost bases were pushed before the runs: docs/v1.2-additions-run4.md, docs/v1.2-additions-run4b.md. Not measurable this round: Werr (its server imports a module missing from the public repository) and DIY Jev (the repository answers 404). No earlier measurement changed.

v1.2.9 (tag v1.2.9): Certo v1. At AltSlate Labs' request, the public MIT altslate/certo-decision-model checkpoint ran all 534 frozen decisions through its author's DecisionModel, serially on our RunPod RTX 3090. The mapping, endpoint condition, published 64-token state and 48-token option limits, and hosted-price cost basis were pushed before the run (docs/v1.2-additions-certo.md). No earlier result or task changed.

v1.2.15 (tag v1.2.15): Zefan Cai's Open-Jev 2B and 9B. Both pinned Apache-2.0 adapter packages ran all 534 frozen decisions through the author's MIT server on our RunPod H100: Open-Jev 9B (Zefan Cai) scored 56.7 (#39) and Open-Jev 2B (Zefan Cai) scored 53.8 (#41). Their full names distinguish them from the other unrelated OpenJev projects already measured. An exact normalized-text audit found no public JevBench state or instruction in the 79,116-row public training projection, and on the held-out diagnostic both score slightly higher on held-out than on public hard items (gaps -2.6 and -2.9 points, field mean -0.7). Endpoint, revisions, licences, cost basis, the overlap check and the held-out figures are in docs/v1.2-additions-zefan-open-jev.md. No earlier result or task changed.

v1.2.16 (tag v1.2.16): rerankers. Added zerank-2, Qwen3-Reranker-4B, BAAI bge-reranker-v2-m3, Mixedbread mxbai-rerank-base-v2, and Alibaba GTE Reranker ModernBERT-base as a new reranker class. All ran the complete 534 frozen decisions through one neutral, preregistered option-ranking adapter. Temperature and yes/no threshold grids were fitted on public items only and then frozen; the public no-instruction baseline and complete sensitivity curves are embedded in each result row. See docs/v1.2-additions-rerankers.md.

v1.2.13 (tag v1.2.13): OpenJev (thinking, BF16). OpenJev's native typed-API think=512 switch ran all 534 frozen decisions on the same H200. See docs/v1.2-additions-openjev-thinking.md.

v1.2.12 (tag v1.2.12): djev (thinking). An experimental full-generation path over the same open DiffusionGemma checkpoint ran all 534 frozen decisions with thinking enabled and a fixed 8,192-token cap. Current djev-dev itself hard-codes thinking off, one denoising step and read-only inference, so this is not described as a switch in its published typed API. See docs/v1.2-additions-djev-thinking.md.

v1.2.11 (tag v1.2.11): djev openness correction. The djev score and all measurements are unchanged; its row now links the Apache-2.0 self-hostable runtime and identifies the Apache-2.0 Google base weights and absence of djev-specific weights.

v1.2.10 (tag v1.2.10): smalljev semantic-v9. At Aditya's request, the public Apache-2.0 isHeSatoshi/smalljev MiniCPM5-2B-Base LoRA and native decision heads ran all 534 frozen decisions serially on our lium.io A6000, scoring 62.4 (#29). The mapping, endpoint condition, non-zero hosted-reference cost basis and allowed public-benchmark-directed training disclosure were committed before the run (docs/v1.2-additions-smalljev.md). No earlier result or task changed.

v1.1.3: the GPU round

Results -> RESULTS-v1.1.3.md · artifact results/v1.1.3/jevbench-v1.1.3-results.json

The v1.1 task set and v1.1.2 scoring, unchanged, plus six rows for open rebuilds that need a GPU, each run the way its author serves it on a rented RunPod GPU: OpenJev on DiffusionGemma 26B-A4B (razorback16), SemIf on Qwen3.5-4B, open-alternative-jev (complete this time; plus a marked post-hoc mode), system-one on Qwen3-8B, and Bespoke Nimble 9B. Speed is measured from Germany over the internet to the GPU, like every remote entrant. v1.1.2 rows are unchanged; only ranks move. Their hard-tier runs feed v1.2.

v1.1: three sub-benchmarks and one Main Score (superseded by v1.2)

Results -> RESULTS-v1.1.md · artifact results/v1.1/jevbench-v1.1-results.json

JevBench v1.1 Main Score

Sub-benchmark What it measures Score 0-100
Capability accuracy on 314 decisions in three tiers - easy (72, new in v1.1), standard (96), judge (146) mean of the three tier accuracies
Speed median and p95 latency, serial, network included log scale: 0.1 s = 100, 1 s = 50, 10 s = 0
Cost $ per 1,000 decisions: public tariff x measured tokens, or (no tariff) a labelled estimate at hosted-provider prices for the same weights or size class (table) log scale: $0.001 = 100, $0.01 = 75, $0.10 = 50, $1 = 25, $10 = 0

JevBench Main Composite Score = (Capability + Speed + Cost) / 3 - "Balanced 33:33:33" (v1.1.2). Three more named weightings are published beside it - Emphasis on Accuracy (60:20:20), the default until v1.1.1; Emphasis on Speed (20:60:20); Emphasis on Cost (20:20:60) - plus capability-only and a geometric mean. benchmarkheaven.com/jev-models lets you set your own weights. Calibration (Brier, ECE) is reported for every system with a distribution but is not part of the score - see RESULTS-v1.1.md for why. The rules are pure functions in jevbench/composite.py.

The easy tier exists so that small function-calling models are measured, not floored: clear-cut intent, explicit yes/no facts, enum extraction, one obviously right tool. 48 of its items are public in datasets/public/easy.jsonl; 24 are held out.

Revisions of v1.1 (all 19 Sep 2026; items, answers, Capability and Speed never changed): v1.1 (tag v1.1) Main Score 60:20:20; v1.1.1 (tag v1.1.1) Cost re-priced at hosted-provider prices for every system without a tariff; v1.1.2 Main Score weights changed to Balanced 33:33:33 (the previous default is kept as the preset "Emphasis on Accuracy") and the Cost scale widened to $0.001-$10 per 1,000 decisions, so no system sits at the 100 cap.

v1.1 numbers are never mixed with v1.0's. v1.0 is described below and its results stay in RESULTS.md as published.

v1.0

v1.0 scored five axes side by side without a composite: smart (is it right), cheap (what 1,000 decisions cost), fast (end-to-end latency, network included), reliable (does the stated probability mean anything, does it survive a rephrasing, does it keep to the schema) and open (weights and licence).

What is in the suite (v1.0; v1.1 adds the easy tier)

242 decisions, six families, three cohorts:

Cohort Decisions Published here?
original-public 72 (36 paraphrase pairs) yes, datasets/public/original.jsonl
heldout-private 24 no - held back so the suite cannot be trained on in full
imported-public-source 146 no - ground truth is ours, the task text is not ours to redistribute

Families: request routing, answer adequacy judging, policy yes/no checks, intent classification, ordinal severity scoring and enum extraction. Every item states its exact label set; the model answers over that set and nothing else.

The 146 imported decisions come from our own auto-router experiment: 78 routing requests with human-assigned categories, and 68 answer-adequacy judgements whose ground truth is a deterministic grader's verdict on a saved answer. See THIRD-PARTY.md.

How a model is asked

Every system sees the same state, the same instructions, the same rubric and the same exact label set. Only the transport differs, and each adapter uses the interface the author published:

Adapter For
typesafe TypeSafe's /v1/systemone, and the open rebuilds that implement the same wire format
systemone_list the list-shaped /decide flavour some rebuilds ship
gradio_space a rebuild whose only public interface is its Hugging Face Space demo
local_openjev open weights loaded in-process, no network
openai_compat ordinary instruction models, JSON-schema-constrained

The first four read the model's own probability distribution. The last one asks the model to write probabilities out under a schema. Those are different objects and are labelled native and verbalized everywhere. Token-level logprobs are not used anywhere, for anyone.

Reproduce

The optional qwen_flash_linear adapter connects to the Qwen3.8 Flash Next compact BF16 linear runtime. Its pinned deployment instructions and repeatability limitations are documented separately; no scored row is added.

Python 3.10+. The HTTP adapters need only the standard library.

python -m unittest discover -s tests -v

# Jev, the published 72-decision cohort
python -m jevbench.cli run --tasks datasets/public/original.jsonl \
  --adapter typesafe --model jev-latest --key-env TYPESAFE_API_KEY \
  --price-in-per-m 0.042 --price-out-per-m 0 \
  --results RUN/results.jsonl --raw-dir RUN/raw \
  --ledger RUN/ledger.jsonl --cap-usd 15 --manifest RUN/manifest.json

# an open rebuild on its author's public endpoint - note the empty key
python -m jevbench.cli run --tasks datasets/public/original.jsonl \
  --adapter typesafe --endpoint https://SOME-PUBLIC-ENDPOINT --key-env '' \
  --model jev-latest --cost-basis no_billable_account_public_endpoint \
  --reserve-usd 0 --delay-s 0.2 \
  --results RUN2/results.jsonl --raw-dir RUN2/raw --ledger RUN/ledger.jsonl

python -m jevbench.cli summarize --tasks datasets/public/original.jsonl \
  --results RUN/results.jsonl --public-export RUN/summary.json

Keys live in the environment and are named, never written into a config file, a result or a log. --key-env '' sends no Authorization header at all, which is what a stranger's public endpoint should get from us.

House rules the harness enforces rather than documents:

  • One budget for everything. A file-locked ledger reserves the worst-case c

Read the rest on GitHub

Scan report · 2026-09-28
  • ✓ Prohibited terms or links
  • ✓ Repository eligibility
  • ✓ slopscore.md paperwork
  • ✓ Content policy
  • ✓ Risk review

From the balcony · 1 of 3 clapped

  1. Crusoeclapped
    No vulnerable dependencies, clear benchmark methodology with transparent results, no credential requests or telemetry concerns, and legitimate open-source research tool.

Princess and Schnitzel read it and passed. Their reasons are on the balcony, with every other verdict.

Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are.

0 comments

log in to comment.

report this listing — log in to report