Synthiq Labs · Leaderboard

The on-device efficiency frontier.

Quality per gigabyte, per watt, per second — every number measured or cited. How we measure →

Intelligence per GB

Cited quality ÷ on-disk size · Higher is better

29.7pts/GB

Leader: Qwen3-0.6B (bf16)

Intelligence per Watt

Quality ÷ avg decode draw · Higher is better

Pending first energy runs — never estimated.

“Runs on” RAM class

From measured peak RAM only — never file size

Badges land with the first device runs.

The board.

updated Jul 23, 2026 · raw JSON

9 models · 0 runtimes measured · 0 devices measured — preview. Quality is publisher-cited for now. Efficiency cells are pending, never zero. Results land as runs land.

How to read this board
  • Quality (cited) — the publisher’s own bf16 figure, exact variant labeled per record. Variants differ across publishers, so ordering is provisional until our one self-run harness replaces them. No published MMLU-family figure → “—”, deliberately unranked.
  • Intelligence/GB — cited quality ÷ the same weights’ size. One precision per row; a bf16 score is never paired with a Q4 file size.
  • Efficiency columns — pending until our device runs land. A dash means not-yet-measured, never zero.
  • record → opens the full record: every cited score, source, and reproduction config. Full method: methodology.
Familyruntime / device / RAM-class filters appear with the first measured runs

Qwen3-0.6B

0.6B · Apache-2.0 · bf16

record →
Quality
44.6MMLU-Redux
Size
1.50 GB
Intel/GB
29.7
Efficiency
pending runs

Llama 3.2 1B-Instruct

1.23B · Llama 3.2 Community License · bf16

record →
Quality
49.3MMLU
Size
2.47 GB
Intel/GB
20.0
Efficiency
pending runs

Qwen3-1.7B

1.7B · Apache-2.0 · bf16

record →
Quality
64.4MMLU-Redux
Size
4.06 GB
Intel/GB
15.9
Efficiency
pending runs

Llama 3.2 3B-Instruct

3.21B · Llama 3.2 Community License · bf16

record →
Quality
63.4MMLU
Size
6.43 GB
Intel/GB
9.9
Efficiency
pending runs

Qwen3-4B

4B · Apache-2.0 · bf16

record →
Quality
77.3MMLU-Redux
Size
8.04 GB
Intel/GB
9.6
Efficiency
pending runs

Phi-4-mini-instruct

3.8B · MIT · bf16

record →
Quality
67.3MMLU
Size
7.67 GB
Intel/GB
8.8
Efficiency
pending runs

Gemma 3 1B-it

1B · Gemma Terms of Use · bf16

record →
Quality
Size
2.00 GB
Intel/GB
Efficiency
pending runs

Gemma 3 4B-it

4B · Gemma Terms of Use · bf16

record →
Quality
Size
8.60 GB
Intel/GB
Efficiency
pending runs

SmolLM3-3B

3B · Apache-2.0 · bf16

record →
Quality
Size
6.15 GB
Intel/GB
Efficiency
pending runs

Full records.

Every score, source, and config — same data as the JSON.

Qwen3-0.6B bf16frontier

0.6B params · Apache-2.0 · Qwen/Qwen3-0.6B · added Jul 23, 2026

cited
Benchmark (as printed)ModeShotsScoreMetric
MMLU-Reduxheadlineno-think44.6as printed
MMLU-Reduxthink55.6as printed
GPQA-Diamondno-think22.9as printed
GPQA-Diamondthink27.9as printed
IFEval strict promptno-think54.5as printed
IFEval strict promptthink59.2as printed
MATH-500no-think55.2as printed
MATH-500think77.6as printed
On-disk size
1.50 GBcitedHF file listing — bf16 safetensors, 1,503,300,328 bytes
Intelligence / GB
29.7computed from: quality (MMLU-Redux) (external) ÷ on-disk size (external)
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

Dual-mode model: thinking and non-thinking scores are published separately and shown separately here — never averaged. The headline uses the non-thinking mode (the default on-device chat configuration).

Llama 3.2 1B-Instruct bf16frontier

1.23B params · Llama 3.2 Community License · meta-llama/Llama-3.2-1B-Instruct · added Jul 23, 2026

cited
Benchmark (as printed)ModeShotsScoreMetric
MMLUheadline549.3macro_avg/acc
IFEval59.5as printed
GPQA027.2acc
GSM8K844.4em_maj1@1
ARC-C059.4acc
Hellaswag041.2acc
On-disk size
2.47 GBcitedHF file listing — bf16 safetensors, 2,471,645,608 bytes
Intelligence / GB
20.0computed from: quality (MMLU) (external) ÷ on-disk size (external)
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

Meta states the parameter count as "1B (1.23B)"; 1.23 is used here. Meta reports plain GPQA (not GPQA-Diamond) — the variant difference matters when eyeballing across rows.

Qwen3-1.7B bf16frontier

1.7B params · Apache-2.0 · Qwen/Qwen3-1.7B · added Jul 23, 2026

cited
Benchmark (as printed)ModeShotsScoreMetric
MMLU-Reduxheadlineno-think64.4as printed
MMLU-Reduxthink73.9as printed
GPQA-Diamondno-think28.6as printed
GPQA-Diamondthink40.1as printed
IFEval strict promptno-think68.2as printed
IFEval strict promptthink72.5as printed
MATH-500no-think73.0as printed
MATH-500think93.4as printed
On-disk size
4.06 GBcitedHF file listing — bf16 safetensors, 4,063,515,592 bytes
Intelligence / GB
15.9computed from: quality (MMLU-Redux) (external) ÷ on-disk size (external)
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

Dual-mode model: thinking and non-thinking scores are published separately and shown separately here — never averaged. The headline uses the non-thinking mode (the default on-device chat configuration).

Llama 3.2 3B-Instruct bf16

3.21B params · Llama 3.2 Community License · meta-llama/Llama-3.2-3B-Instruct · added Jul 23, 2026

cited
Benchmark (as printed)ModeShotsScoreMetric
MMLUheadline563.4macro_avg/acc
IFEval077.4avg prompt/instruction acc, loose/strict
GPQA032.8acc
GSM8K877.7em_maj1@1
MATH048.0final_em
ARC-C078.6acc
On-disk size
6.43 GBcitedHF file listing — bf16 safetensors, 6,425,529,048 bytes
Intelligence / GB
9.9computed from: quality (MMLU) (external) ÷ on-disk size (external)
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

Meta states the parameter count as "3B (3.21B)"; 3.21 is used here. Meta reports plain GPQA (not GPQA-Diamond).

Qwen3-4B bf16frontier

4B params · Apache-2.0 · Qwen/Qwen3-4B · added Jul 23, 2026

cited
Benchmark (as printed)ModeShotsScoreMetric
MMLU-Reduxheadlineno-think77.3as printed
MMLU-Reduxthink83.7as printed
GPQA-Diamondno-think41.7as printed
GPQA-Diamondthink55.9as printed
IFEval strict promptno-think81.2as printed
IFEval strict promptthink81.9as printed
MATH-500no-think84.8as printed
MATH-500think97.0as printed
On-disk size
8.04 GBcitedHF file listing — bf16 safetensors, 8,044,982,000 bytes
Intelligence / GB
9.6computed from: quality (MMLU-Redux) (external) ÷ on-disk size (external)
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

Dual-mode model: thinking and non-thinking scores are published separately and shown separately here — never averaged. The headline uses the non-thinking mode (the default on-device chat configuration).

Phi-4-mini-instruct bf16frontier

3.8B params · MIT · microsoft/Phi-4-mini-instruct · added Jul 23, 2026

cited
Benchmark (as printed)ModeShotsScoreMetric
MMLUheadline567.3as printed
MMLU-Pro052.80-shot CoT
GPQA025.20-shot CoT
BigBench Hard070.40-shot CoT
GSM8K888.68-shot CoT
MATH064.00-shot CoT
On-disk size
7.67 GBcitedHF file listing — bf16 safetensors, 7,672,066,216 bytes
Intelligence / GB
8.8computed from: quality (MMLU) (external) ÷ on-disk size (external)
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

Microsoft reports plain GPQA with CoT (not GPQA-Diamond). The card publishes no IFEval or HumanEval rows, so none are listed here.

Gemma 3 1B-it bf16

1B params · Gemma Terms of Use · google/gemma-3-1b-it · added Jul 23, 2026

pending
Benchmark (as printed)ModeShotsScoreMetric
GPQA Diamond019.2as printed
IFEval080.2as printed
BIG-Bench Hard039.1as printed
MMLU-Pro014.7as printed
GSM8K062.8as printed
HumanEval041.5as printed
On-disk size
2.00 GBcitedHF file listing — bf16 safetensors, 1,999,811,208 bytes
Intelligence / GB
pending — a component is not yet available
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

No headline metric: Google publishes no plain-MMLU-family figure for the instruction-tuned model (MMLU-Pro is a different, harder benchmark and is not chart-comparable with other publishers' MMLU/MMLU-Redux figures). This row is excluded from the density ranking and the chart until the unified self-run harness pass lands; its cited per-task scores are all listed.

Gemma 3 4B-it bf16

4B params · Gemma Terms of Use · google/gemma-3-4b-it · added Jul 23, 2026

pending
Benchmark (as printed)ModeShotsScoreMetric
GPQA Diamond030.8as printed
IFEval090.2as printed
BIG-Bench Hard072.2as printed
MMLU-Pro043.6as printed
GSM8K089.2as printed
HumanEval071.3as printed
On-disk size
8.60 GBcitedHF file listing — bf16 safetensors, 8,600,277,880 bytes
Intelligence / GB
pending — a component is not yet available
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

No headline metric: Google publishes no plain-MMLU-family figure for the instruction-tuned model (MMLU-Pro is a different, harder benchmark and is not chart-comparable with other publishers' MMLU/MMLU-Redux figures). This row is excluded from the density ranking and the chart until the unified self-run harness pass lands; its cited per-task scores are all listed. Multimodal model — sizes here are the text weights as shipped.

SmolLM3-3B bf16

3B params · Apache-2.0 · HuggingFaceTB/SmolLM3-3B · added Jul 23, 2026

pending
Benchmark (as printed)ModeShotsScoreMetric
GPQA Diamondno-think35.7as printed
GPQA Diamondthink41.7as printed
IFEvalno-think76.7as printed
IFEvalthink71.2as printed
GSM-Plusno-think72.8as printed
GSM-Plusthink83.4as printed
LiveCodeBench v4no-think15.2as printed
LiveCodeBench v4think30.0as printed
On-disk size
6.15 GBcitedHF file listing — bf16 safetensors, 6,150,235,008 bytes
Intelligence / GB
pending — a component is not yet available
On-device efficiency
pendingdecode · prefill · TTFT · peak RAM · cold start · energy — all await Synthiq Labs device runs
Reproduction config
ships with the first self-run measurement for this artifact

No headline metric: the instruct tables carry no MMLU-family figure (the base-model table reports MMLU-CF, a contamination-free variant not comparable with other publishers' MMLU numbers), so this row is excluded from the density ranking and chart until the self-run harness pass lands. Dual-mode model — think and no-think scores listed separately, never averaged.

Every number here is a git commit. Disagree with one? File the discrepancy.

Follow the research →