Local inference

Benchmarks

Local LLMs served on the desktop's RTX 3060. Diffusion image generation on the MacBook's M1.

Qwen3.5:9B

Desktop Computer: RTX 3060 (12GB)

adhoc drop shard11 — 0% [00:00<?, ?it/s] · updated 5:40pm PDT
115 more queued · ~68h12m total
  • q drop_shard12 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q drop_shard13 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q drop_shard14 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa1 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa2 abliterated thinkingprio -1030m00s · unprobed~30m00s
… and 110 more
  • q babilong_qa3 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa4 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa5 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa6 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa7 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa8 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa9 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa10 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa11 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa12 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa13 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa14 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa15 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa16 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa17 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa18 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa19 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa20 abliterated thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa1 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa2 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa3 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa4 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa5 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa6 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa7 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa8 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa9 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa10 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa11 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa12 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa13 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa14 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa15 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa16 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa17 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa18 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa19 official thinkingprio -1030m00s · unprobed~30m00s
  • q babilong_qa20 official thinkingprio -1030m00s · unprobed~30m00s
  • q ifeval_shard07 official thinkingprio -1030m36s · unprobed~30m36s
  • q ifeval_shard00 official thinkingprio -1042m00s · unprobed~42m00s
  • q ifeval_shard01 official thinkingprio -1042m00s · unprobed~42m00s
  • q ifeval_shard02 official thinkingprio -1042m00s · unprobed~42m00s
  • q ifeval_shard03 official thinkingprio -1042m00s · unprobed~42m00s
  • q ifeval_shard04 official thinkingprio -1042m00s · unprobed~42m00s
  • q ifeval_shard05 official thinkingprio -1042m00s · unprobed~42m00s
  • q ifeval_shard06 official thinkingprio -1042m00s · unprobed~42m00s
  • q squad_completion_shard00 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard01 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard02 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard03 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard04 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard05 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard06 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard07 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard08 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard09 abliterated thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard00 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard01 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard02 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard03 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard04 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard05 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard06 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard07 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard08 official thinkingprio -1045m00s · unprobed~45m00s
  • q squad_completion_shard09 official thinkingprio -1045m00s · unprobed~45m00s
  • q babi_shard00 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q babi_shard01 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q babi_shard02 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q babi_shard03 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q babi_shard04 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q babi_shard05 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q babi_shard06 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q babi_shard07 abliterated thinkingprio -1054m00s · unprobed~54m00s
  • q drop_shard07 abliterated thinkingprio -1050s · 1 probe1m20s
  • q drop_shard00 abliterated thinkingprio -101m10s · 1 probe1m40s
  • q babi_full_shard12 officialprio -102m15s · 1 probe2m45s
  • q drop_shard02 abliterated thinkingprio -103m30s · 1 probe4m00s
  • q drop_shard08 abliterated thinkingprio -103m50s · 1 probe4m20s
  • q drop_shard01 abliterated thinkingprio -104m30s · 1 probe5m00s
  • q drop_shard03 abliterated thinkingprio -106m10s · 1 probe6m40s
  • q drop_shard04 abliterated thinkingprio -106m50s · 1 probe7m20s
  • q babi_full_shard11 officialprio -107m45s · 1 probe8m15s
  • q drop_shard09 abliterated thinkingprio -108m10s · 1 probe8m40s
  • q babi_shard08 abliterated thinkingprio -1010m30s · 1 probe11m00s
  • q drop_shard05 abliterated thinkingprio -1011m10s · 1 probe11m40s
  • q babi_full_shard00 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard01 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard03 officialprio -1016m00s · 1 probe~16m30s
  • q babi_full_shard04 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard05 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard06 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard07 officialprio -1016m00s · 1 probe~16m30s
  • q babi_full_shard08 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard09 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard10 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard13 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard14 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard15 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard16 officialprio -1016m00s · 1 probe~16m30s
  • q babi_full_shard17 officialprio -1016m00s · 1 probe16m30s
  • q babi_full_shard18 officialprio -1016m00s · 1 probe~16m30s
  • q babi_full_shard19 officialprio -1016m00s · 1 probe16m30s
  • q drop_shard06 abliterated thinkingprio -1016m30s · 1 probe17m00s
  • q babi_full_shard02 officialprio -1020m08s · 1 probe20m38s
  • q drop_shard10 abliterated thinkingprio -1034m50s · 1 probe35m20s
  • q fda abliterated thinkingprio -101h41m · 8 probes1h42m
  • q bbh_cot_zeroshot official thinkingprio -102h41m · 1 probe2h42m
  • q cnn_dailymail_abisee official thinkingprio -102h46m · 1 probe2h46m
  • q triviaqa abliterated thinkingprio -104h48m · 6 probes~4h49m
Abliterated (ours)

huihui-ai abliterated build — the Qwen3.5‑9B base model with abliteration applied — independently‑verified genuine GGUF, served via the 16‑way raw‑llama.cpp bypass, 2026‑08‑23 onward — 12 of 12 tasks done.

Abliterated, thinking (ours)

Same abliterated build, thinking mode forced on instead of the non‑thinking default used elsewhere on this page.3 Three results so far (ifeval, mmlu_pro, bbh_cot_zeroshot) from a 12‑task rerun still in progress — most rows don't have one yet.

Qwen 3.5:9b (ours)

Unsloth official Qwen3.5‑9B weights, same harness/prompting, run 2026‑08‑22–23 (two tasks via the raw‑llama.cpp bypass).

Qwen 3.5:9b, thinking (ours)

Same official weights as the row above, thinking mode forced on — the fourth of the four core settings this page tracks (abliterated/official × thinking/non‑thinking).5 One result so far (mmlu_pro), meant to fill in for every benchmark over time like the abliterated, thinking row above it.

Qwen3.5‑9B (posted)

Numbers from Qwen's own model card, quoted as-is and not run by us — except cruxeval_input/cruxeval_output, where no Qwen number exists at all and this row instead shows a same‑scale substitute model (labeled by name, not "Qwen3.5‑9B").

Score comparison

One table per benchmark, all scores as percentages. Δ is the change versus the row above, chained abliterated→Qwen 3.5:9b — except the two thinking rows (where present), whose Δ is each against its own non‑thinking baseline rather than a further link in that chain: the abliterated, thinking row compares against the plain abliterated row above it, and the Qwen 3.5:9b, thinking row compares against the plain Qwen 3.5:9b row above it — two independent same‑model thinking‑on/off comparisons, not one chain of four. The posted row isn't chained in either — see footnote 6 for why that comparison needs its own caveats. A task's posted row is omitted entirely when no public number (or substitute) exists for it, rather than shown as a blank — see footnote 1. Some backlog tasks below have only the abliterated row filled in; a qwen comparison doesn't exist yet for those.

Score comparison chart
gsm8k
82.5
not reported
85.7 ± 1.0
not reported
humaneval_instruct
61.6 ± 3.8
not reported
47.6 ± 3.9
not reported
mbpp_instruct_fixed
48.2 ± 2.2
not reported
48.0 ± 2.2
not reported
bbh_cot_zeroshot
45.6 ± 0.5
79.9 ± 1.1
65.8 ± 0.5
not reported
mmlu_pro
62.0 ± 0.4
67.5 ± 1.6
64.7 ± 0.4
77.0 ± 1.4
82.5
ifeval
64.7
85.7
58.8
not reported
91.5
cnn_dailymail
7.9
not reported
10.0
not reported
fda
76.0
not reported
70.2
80.8
squad_completion
75.8
not reported
70.7
not reported
cruxeval_input
44.1 ± 1.7
not reported
52.8 ± 1.7
not reported
37.4
cruxeval_output
27.6 ± 1.5
not reported
31.6 ± 1.6
54.3 ± 5.7
41.2
babilong
75.8
not reported
83.3
not reported

gsm8kexact_match · 1,319 samples

grade-school math word problems

SourceScoreΔ
Abliterated (ours)82.5strict 82.6
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)85.7 ± 1.0▲ 3.2
Qwen 3.5:9b, thinking (ours)not reported

humaneval_instructpass@1 · 164 samples

Python function generation from docstrings

SourceScoreΔ
Abliterated (ours)61.6 ± 3.8
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)47.6 ± 3.9▼ 14.0
Qwen 3.5:9b, thinking (ours)not reported

mbpp_instruct_fixedpass@1 · 500 samples

basic Python programming problems

SourceScoreΔ
Abliterated (ours)48.2 ± 2.2
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)48.0 ± 2.2▼ 0.2
Qwen 3.5:9b, thinking (ours)not reported

bbh_cot_zeroshotexact_match, flexible-extract · 6,511 samples

diverse reasoning — logic, algorithms, language

SourceScoreΔ
Abliterated (ours)45.6 ± 0.5
Abliterated, thinking (ours)79.9 ± 1.13▲ 34.3
Qwen 3.5:9b (ours)65.8 ± 0.5▲ 20.2
Qwen 3.5:9b, thinking (ours)not reported

mmlu_proexact_match · 12,032 samples

broad academic knowledge, 57 subjects

SourceScoreΔ
Abliterated (ours)62.0 ± 0.4
Abliterated, thinking (ours)67.5 ± 1.63▲ 5.5
Qwen 3.5:9b (ours)64.7 ± 0.4▲ 2.7
Qwen 3.5:9b, thinking (ours)77.0 ± 1.45▲ 12.3
Qwen3.5‑9B (posted)82.51

ifevalloose‑instruction acc · 541 samples

instruction-following & formatting compliance

SourceScoreΔ
Abliterated (ours)64.7strict‑p 48.8 · strict‑i 61.9 · loose‑p 52.7
Abliterated, thinking (ours)85.73strict‑p 79.5 · strict‑i 83.9 · loose‑p 81.7▲ 21.0
Qwen 3.5:9b (ours)58.8strict‑p 44.2–44.42 · strict‑i 57.4–57.7 · loose‑p 45.7▼ 5.9
Qwen 3.5:9b, thinking (ours)not reported
Qwen3.5‑9B (posted)91.51

cnn_dailymailrouge · 11,490 samples

news article summarization, single ROUGE score

SourceScoreΔ
Abliterated (ours)7.9
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)10.0▲ 2.2
Qwen 3.5:9b, thinking (ours)not reported

fdacontains · 1,102 samples

extracting a structured field from FDA drug-label text

SourceScoreΔ
Abliterated (ours)76.0
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)70.2▼ 5.8
Qwen 3.5:9b, thinking (ours)80.8▲ 10.6

squad_completioncontains · 2,984 samples

completing a passage's next span, SQuAD-derived

SourceScoreΔ
Abliterated (ours)75.8
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)70.7▼ 5.1
Qwen 3.5:9b, thinking (ours)not reported

cruxeval_inputpass@1 · 800 samples

predicting a Python function's input from its output

SourceScoreΔ
Abliterated (ours)44.1 ± 1.7
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)52.8 ± 1.7▲ 8.7
Qwen 3.5:9b, thinking (ours)not reported
deepseek-instruct-6.7b (posted, ref)37.44

cruxeval_outputpass@1 · 800 samples

predicting a Python function's output from its input

SourceScoreΔ
Abliterated (ours)27.6 ± 1.5
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)31.6 ± 1.6▲ 4.0
Qwen 3.5:9b, thinking (ours)54.3 ± 5.7▲ 22.7
deepseek-instruct-6.7b (posted, ref)41.24

babilongacc (20‑subtask weighted avg) · 39,972 samples

long‑context multi‑hop QA, 20 reasoning‑complexity levels

SourceScoreΔ
Abliterated (ours)75.8qa1–qa20 range 9.8–99.9
Abliterated, thinking (ours)not reported
Qwen 3.5:9b (ours)83.3qa1–qa20 range 11.2–99.8▲ 7.5
Qwen 3.5:9b, thinking (ours)not reported

Abliterated vs Qwen 3.5:9b

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
5
Qwen 3.5:9b
7

Abliterated, thinking vs Qwen 3.5:9b, thinking

row beats
column
Abliterated, thinking
Qwen 3.5:9b, thinking
Abliterated, thinking
0
Qwen 3.5:9b, thinking
1

Only 1 task compared so far — preliminary.

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
0
Qwen 3.5:9b
2

Only 2 tasks compared so far — preliminary.

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
2
Qwen 3.5:9b
2
row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
0
Qwen 3.5:9b
1

Only 1 task compared so far — preliminary.

row beats
column
Abliterated
Qwen 3.5:9b
Abliterated
3
Qwen 3.5:9b
2
  1. mmlu_pro and ifeval are the only two rows that are genuine name‑for‑name matches with a posted number — and even then, exact prompt version and scoring protocol aren't guaranteed identical. Treat the gap's size as directional, not exact. Every other task has no posted row at all: no public Qwen3.5 number exists for it anywhere we could find (checked Qwen's own model card, and for cruxeval/babilong, their independent community leaderboards directly — cached in external_data/), except cruxeval_input/cruxeval_output, which get a substitute instead of a blank row — see footnote 4.
  2. ifeval's strict submetrics show small run‑to‑run nondeterminism on identical cached inputs (four replays: strict‑prompt 43.99–44.36, strict‑inst 57.43–57.67). Loose metrics were stable. The spread is far smaller than the gaps in this table, but don't read these two columns to more than one decimal place.
  3. ifeval's, mmlu_pro's, and bbh_cot_zeroshot's "Abliterated, thinking" rows are the same abliterated build with thinking mode forced on (max_gen_toks=10000, no early stop) instead of the non‑thinking default used everywhere else on this page — see eval‑harness/run_thinking_batch.sh. They're the first three results back from a broader 12‑task thinking‑mode rerun still in progress (excludes humaneval_instruct/mbpp_instruct_fixed, where thinking mode is a confirmed no‑op under this harness's gen_prefix); more rows will fill in as that batch completes.
  4. cruxeval_input/cruxeval_output's posted row is a substitute, not a Qwen number: no Qwen model appears on the independent CRUXEval leaderboard (crux‑eval.github.io) at all, so the row instead shows the closest‑parameter‑count instruct model there, picked automatically from the cached leaderboard dump rather than hand‑typed — see generate.py's _closest_instruct_reference().
  5. mmlu_pro's "Qwen 3.5:9b, thinking" row is the official (non‑abliterated) weights with thinking mode forced on, run via eval‑harness/run_official_thinking_mmlu_pro.sh against the same 756‑sample subsample as the abliterated, thinking row above it — built specifically to isolate how much of the gap between abliterated, thinking (67.5) and Qwen's posted mmlu_pro number (82.5) is attributable to the abliterated build itself versus this local build's shared Q4_K_M quantization. It's the first result of a fourth core setting (official weights × thinking) this page now tracks alongside abliterated/qwen/abliterated‑thinking — more tasks will fill in over time the same way the abliterated, thinking rows have been.
  6. The "posted" column isn't part of any head‑to‑head count on this page (scorecards above, or otherwise): Qwen's posted numbers are almost certainly thinking‑mode (the model's default), while ours are forced non‑thinking, and ours also run Q4_K_M quantized rather than full precision. The gap to "posted" reflects both of those on top of any real capability difference — every win/loss count on this page only ever compares our own runs against each other.

Runtime & throughput

Per‑item speed for each run, wall‑clock runtime alongside it for reference. Not a clean speed benchmark — see the callout below before reading these as apples‑to‑apples.

Two different serving conditions, not one axis: Qwen 3.5:9b ran serially for gsm8k/ifeval/humaneval_instruct/mbpp_instruct_fixed/cnn_dailymail, but through the 16‑way raw‑llama.cpp bypass for mmlu_pro/bbh_cot_zeroshot — two different axes inside the same column. Abliterated runs through the 16‑way bypass for every task except cnn_dailymail, which ran serially like its qwen row instead, so that pair is still apples‑to‑apples. The abliterated, thinking row (ifeval, mmlu_pro, and bbh_cot_zeroshot so far) and the Qwen 3.5:9b, thinking row (mmlu_pro so far) both run through the same bypass build at 8‑way, not 16‑way — sized down from 16 for comfortable worst‑case context headroom on long thinking traces, not a VRAM tradeoff (-c is the same total KV pool either way; see hearth‑bypass/bypass_run_abliterated_thinking.bat's and bypass_run_thinking.bat's matching -np 8, and EVAL‑BACKLOG.md). The much higher per‑item time there is thinking tokens plus that lower concurrency, not a like‑for‑like comparison to the 16‑way non‑thinking rows. The five backlog tasks (fda/squad_completion/cruxeval_input/cruxeval_output/babilong) only have an abliterated row at all, all via the 16‑way bypass. Only compare rows where both sides used the same serving mode.
Runtime/speed reconstructed from raw logs, 2026‑08‑29: these jobs get killed and restarted (digest bypass‑protection blackout windows, manual interrupts), so lm_eval's own end‑to‑end wall‑clock was including dead time between restarts — nothing to do with how fast the model actually processed items. The abliterated² figures for gsm8k/humaneval_instruct/mbpp_instruct_fixed/bbh_cot_zeroshot/mmlu_pro/ifeval and the backlog cruxeval_input/cruxeval_output rows are now reconstructed straight from each surviving log's tqdm ticks (see `speed_stats.py`): summed real per‑tick elapsed deltas for the runtime, median per‑tick rate (not a global average) for the speed, so a stray near‑instant cache‑hit or ramp‑up tick can't skew it. All eight logs showed a single continuous segment with no restart evidence, so the change here is small — mostly stripping non‑generation setup/teardown overhead. Two rows were deliberately left unchanged despite being asked for: ifeval and mmlu_pro's abliterated², thinking figures, because their logs show a genuine crash/partial run with no way to cleanly reconstruct true generation time from what survived on disk — guessing would be worse than leaving them at the old (also imperfect) number.
Runtime comparison chart

Bars are scaled per task (each task's own longest bar is 100% width) — runtimes span two orders of magnitude across tasks, so a single shared scale would flatten the short ones to nothing.

gsm8k
35m47s
not reported
2h19m
not reported
humaneval_instruct
2m49s
not reported
14m26s
not reported
mbpp_instruct_fixed
12m11s
not reported
1h02m
not reported
bbh_cot_zeroshot
2h03m
10h59m
1h44m
not reported
mmlu_pro
5h53m
5h14m
3h51m
8h46m
ifeval
33m14s
8h06m
2h49m
not reported
cnn_dailymail
5h41m
not reported
6h43m
not reported
fda
10.8s
not reported
31m10s
1h52m30s
squad_completion
12.4s
not reported
56m11s
not reported
cruxeval_input
58m52s
not reported
2h55m37s
not reported
cruxeval_output
55m07s
not reported
1h02m23s
22m30s
babilong
1m40s
not reported
2h05m
not reported

gsm8kexact_match · 1,319 samples

grade-school math word problems

SourceSpeedRuntime
Abliterated (ours, bypass)2.00s/it16‑way bypass35m47s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)6.32s/itserial2h19m
Qwen 3.5:9b, thinking (ours, bypass)not run

humaneval_instructpass@1 · 164 samples

Python function generation from docstrings

SourceSpeedRuntime
Abliterated (ours, bypass)1.00s/it16‑way bypass2m49s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)5.28s/itserial14m26s
Qwen 3.5:9b, thinking (ours, bypass)not run

mbpp_instruct_fixedpass@1 · 500 samples

basic Python programming problems

SourceSpeedRuntime
Abliterated (ours, bypass)1.00s/it16‑way bypass12m11s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)7.42s/itserial1h02m
Qwen 3.5:9b, thinking (ours, bypass)not run

bbh_cot_zeroshotexact_match, flexible-extract · 6,511 samples

diverse reasoning — logic, algorithms, language

SourceSpeedRuntime
Abliterated (ours, bypass)1.00s/it16‑way bypass2h03m
Abliterated, thinking (ours, bypass)35.00s/it4‑way bypass10h59m
Qwen 3.5:9b (ours)0.96s/it16‑way bypass1h44m
Qwen 3.5:9b, thinking (ours, bypass)not run

mmlu_proexact_match · 12,032 samples

broad academic knowledge, 57 subjects

SourceSpeedRuntime
Abliterated (ours, bypass)2.00s/it16‑way bypass5h53m
Abliterated, thinking (ours, bypass)24.91s/it4‑way bypass5h14m
Qwen 3.5:9b (ours)1.15s/it16‑way bypass3h51m
Qwen 3.5:9b, thinking (ours, bypass)29.00s/it4‑way bypass8h46m

ifevalloose‑instruction acc · 541 samples

instruction-following & formatting compliance

SourceSpeedRuntime
Abliterated (ours, bypass)3.00s/it16‑way bypass33m14s
Abliterated, thinking (ours, bypass)53.88s/it4‑way bypass8h06m
Qwen 3.5:9b (ours)18.77s/itserial2h49m
Qwen 3.5:9b, thinking (ours, bypass)not run

cnn_dailymailrouge · 11,490 samples

news article summarization, single ROUGE score

SourceSpeedRuntime
Abliterated (ours, bypass)1.78s/itserial5h41m
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)2.11s/itserial6h43m
Qwen 3.5:9b, thinking (ours, bypass)not run

fdacontains · 1,102 samples

extracting a structured field from FDA drug-label text

SourceSpeedRuntime
Abliterated (ours, bypass)0.01s/it16‑way bypass10.8s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)1.68s/it16‑way bypass31m10s
Qwen 3.5:9b, thinking (ours, bypass)9.00s/it4‑way bypass1h52m30s

squad_completioncontains · 2,984 samples

completing a passage's next span, SQuAD-derived

SourceSpeedRuntime
Abliterated (ours, bypass)0.00s/it16‑way bypass12.4s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)1.13s/it16‑way bypass56m11s
Qwen 3.5:9b, thinking (ours, bypass)not run

cruxeval_inputpass@1 · 800 samples

predicting a Python function's input from its output

SourceSpeedRuntime
Abliterated (ours, bypass)6.67s/it16‑way bypass58m52s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)1.00s/it16‑way bypass2h55m37s
Qwen 3.5:9b, thinking (ours, bypass)not run

cruxeval_outputpass@1 · 800 samples

predicting a Python function's output from its input

SourceSpeedRuntime
Abliterated (ours, bypass)5.00s/it16‑way bypass55m07s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)0.60s/it16‑way bypass1h02m23s
Qwen 3.5:9b, thinking (ours, bypass)18.00s/it4‑way bypass22m30s

babilongacc (20‑subtask weighted avg) · 39,972 samples

long‑context multi‑hop QA, 20 reasoning‑complexity levels

SourceSpeedRuntime
Abliterated (ours, bypass)0.00s/it16‑way bypass1m40s
Abliterated, thinking (ours, bypass)not run
Qwen 3.5:9b (ours)0.19s/it16‑way bypass2h05m
Qwen 3.5:9b, thinking (ours, bypass)not run

Diffusion image generation · Apple M1 Pro, 16GB unified memory

Macbook: M1 (16GB Unified)

Image-generation speed and memory across diffusion models, run locally on the MacBook's unified memory instead of the desktop's dedicated VRAM — 7 runs across 6 model families that fit comfortably in 16 GB timed so far.

First pass, 2026‑09‑05: one image per run — 512×512 except SDXL Turbo, whose pipeline's native output is 528×528 (see its card's note) — same prompt and seed (42) everywhere metadata or timing is available: "a photorealistic portrait of an astronaut riding a horse on the moon, dramatic lighting, highly detailed, 8k" — run out of ~/diffusion-bench (not tracked in this repo — its venvs and downloaded weights alone run tens of GB; the sample images below are tracked, in images/). Six rows are timed with /usr/bin/time -l: five around an mflux-generate-* call (see run_frontier_batch.sh), and SDXL Turbo via a separate, non‑mflux pipeline (mlx‑examples' stable_diffusion, standardized 2026‑09‑06 — see its card's note). The remaining row (FLUX.1 schnell) comes from mflux's own embedded PNG metadata instead — each card says which. Only models that fit comfortably inside this machine's 16 GB of unified memory are shown below — a model tested but excluded for swapping heavily instead gets a plain callout explaining why, not a card with misleadingly slow numbers (see the notes below the grid). No image-quality scoring here (unlike the LLM tables above, there's no numeric metric for "does it look good") — judge that from the thumbnails, this section tracks generation speed and memory only.

ERNIE Image Turbobaidu/ERNIE-Image-Turbo

ERNIE Image Turbo, astronaut promptERNIE Image Turbo, combined bakery-text + kneading-hands promptERNIE Image Turbo, watercolor fox-and-bird promptERNIE Image Turbo, Seoul cyberpunk night market promptERNIE Image Turbo, octopus vs. raccoon chess prompt
MetricValue
Steps8
Quantize8-bit
Generation time57.0s
Speed7.12s/it avg
Peak MLX memory13.44 GB

LensComfy-Org/Lens

Lens, astronaut promptLens, combined bakery-text + kneading-hands promptLens, watercolor fox-and-bird promptLens, Seoul cyberpunk night market promptLens, octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize8-bit
Generation time13.8s
Speed3.44s/it avg (6.90s cold → 2.21s/it by step 4)
Peak MLX memory12.74 GB

Its text encoder is GPT‑OSS‑20B (mlx‑community/gpt‑oss‑20b‑MXFP4‑Q8), not a typical CLIP/T5 text encoder — a general‑purpose 20B language model doing prompt encoding for a 4‑step image model. Peak memory still fits comfortably despite that.

FLUX.2 klein-9Bblack-forest-labs/FLUX.2-klein-9B

FLUX.2 klein-9B, astronaut promptFLUX.2 klein-9B, combined bakery-text + kneading-hands promptFLUX.2 klein-9B, watercolor fox-and-bird promptFLUX.2 klein-9B, Seoul cyberpunk night market promptFLUX.2 klein-9B, octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize8-bit
Generation time1m28s
Speed22.1s/it avg
Peak MLX memory12.42 GB

The larger sibling of FLUX.2 klein‑4B (9B vs. 4B parameters). The first attempt at this model failed outright on a HuggingFace 403 gated‑repo error; access opened up since, and this retry completed cleanly — still fitting comfortably despite the larger size, though 4.2× slower per step than its 4B sibling.

FLUX.2 klein-4Bblack-forest-labs/FLUX.2-klein-4B

FLUX.2 klein-4B, astronaut promptFLUX.2 klein-4B, combined bakery-text + kneading-hands promptFLUX.2 klein-4B, watercolor fox-and-bird promptFLUX.2 klein-4B, Seoul cyberpunk night market promptFLUX.2 klein-4B, octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize8-bit
Generation time20.8s
Speed5.21s/it avg (7.46s cold → 4.72s/it by step 4)
Peak MLX memory6.89 GB

FLUX.1 schnell (q4)black-forest-labs/FLUX.1-schnell

FLUX.1 schnell (q4), astronaut promptFLUX.1 schnell (q4), combined bakery-text + kneading-hands promptFLUX.1 schnell (q4), watercolor fox-and-bird promptFLUX.1 schnell (q4), Seoul cyberpunk night market promptFLUX.1 schnell (q4), octopus vs. raccoon chess prompt
MetricValue
Steps4
Quantize4-bit
Generation time30.1s
Speed7.54s/it avg
Peak MLX memory10.33 GB

Loaded from a checkpoint pre‑quantized to 4‑bit and saved to disk, rather than re‑quantizing the full‑precision weights at load time — the representative FLUX.1 schnell setting on this page. Steps/speed come straight from mflux's own embedded PNG metadata (see this section's intro note); the original astronaut‑prompt run had no /usr/bin/time wrapper, so peak memory here instead comes from the later, separately‑wrapped combined‑prompt run of the same 4‑bit checkpoint (combined‑flux‑schnell‑q4.log).

Z-Image TurboTongyi-MAI/Z-Image-Turbo

Z-Image Turbo, astronaut promptZ-Image Turbo, combined bakery-text + kneading-hands promptZ-Image Turbo, watercolor fox-and-bird promptZ-Image Turbo, Seoul cyberpunk night market promptZ-Image Turbo, octopus vs. raccoon chess prompt
MetricValue
Steps9
Quantize8-bit
Generation time53.3s
Speed5.92s/it avg (10.37s cold → ~5.2–5.6s/it after)
Peak MLX memory8.74 GB

SDXL Turbostabilityai/sdxl-turbo

SDXL Turbo, astronaut promptSDXL Turbo, combined bakery-text + kneading-hands promptSDXL Turbo, watercolor fox-and-bird promptSDXL Turbo, Seoul cyberpunk night market promptSDXL Turbo, octopus vs. raccoon chess prompt
MetricValue
Steps2
Quantizefp16
Generation time30.2s
Speed8.91s/it avg (15.30s cold → 7.78s/it by step 2)
Peak MLX memory6.98 GB

Standardized 2026‑09‑06 to match the cards above — originally shown as an uninstrumented gallery (no /usr/bin/time, no PNG metadata from this pipeline; see git history for that version). This card's numbers come from a fresh, properly‑wrapped rerun with the same seed and prompt, confirmed byte‑identical to the original image (same file, sdxl-turbo-warm.png). Native output resolution for this pipeline is 528×528, not the 512×512 every other card uses — not operator‑configurable via this script's CLI, so shown as-is rather than cropped/resized to match.

Ideogram 4 (fp8) (ideogram‑ai/ideogram‑4‑fp8, both 8‑bit and 4‑bit), Qwen‑Image (Qwen/Qwen‑Image‑2512, already 4‑bit), FIBO Lite (briaai/Fibo‑lite, 8‑bit), and Krea 2 Turbo (krea/Krea‑2‑Turbo, 8‑bit) were also tested and all completed — but all five runs are excluded from the comparison above rather than shown alongside models that actually fit. Peak memory ran 28.28 GB (Ideogram 8‑bit), 27.36 GB (Ideogram 4‑bit), 27.37 GB (Qwen‑Image), 24.38 GB (FIBO Lite), and 22.48 GB (Krea 2 Turbo), all well past this machine's 16 GB of unified memory, so most of each run's time was spent swapping to disk, not generating — the 88.6s/it, 40.3s/it, 109.7s/it, 21.8s/it, and 48.9s/it measured aren't a fair read on these models' real speed, just how slow this particular machine is at thrashing its way through them.

Summary: everything tested

ModelPeak MemoryFits 16 GB?
SDXL-Turbo6.98 GB
FLUX.1 schnell (4-bit, saved standalone)10.33 GB
FLUX.2 klein-4B6.89 GB
FLUX.2 klein-9B12.42 GB
Z-Image Turbo8.74 GB
ERNIE Image Turbo (Baidu)13.44 GB
Lens (Microsoft)12.74 GB
Ideogram 4 (fp8)27–28 GB
Qwen-Image-251227.37 GB(but best quality)
FIBO Lite (Bria AI)24.38 GB
Krea 2 Turbo22.48 GB(breaks the "Turbo" heuristic)