syndicAI Quality Eval

The SQE Leaderboard

We measure the build you actually run: exact quant packs on real hardware, gated by tests, with every run log published. No vendor harness, no cherry-picked shots.

How SQE scores are computed 37 published campaigns · updated Aug 16, 2026

Methodology v2 (August 2026): pass and gate are scored as separate weighted components. See the changelog note for what changed and why overfit-heavy packs moved.

fourlineMedium · M

Build a Connect-Four-style game from a written spec: board size, start player, and win-rule traps included.

Token budget6M / runTurn cap120 / shotMax shots3Reps5Context200K
Show
Metrics
#ModelCompare
01
Claude Opus 4.7 Closed referenceanthropic/claude-opus-4.7 · Closed API reference
99
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
02
DeepSeek V4 Flash 0731 deepseek/deepseek-v4-flash-0731 · Official FP4 + FP8 Mixed
99
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
03
Kimi K2.7 Code moonshotai/kimi-k2.7-code · Flagship open-weight
99
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
04
Laguna S 2.1 poolside/Laguna-S-2.1-FP8 · Vendor FP8
96
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
05
Laguna S 2.1 Community packkkuspa/Laguna-S-2.1-NVFP4-0804 · Community NVFP4
96
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
06
Qwen3.6 27B Qwen/Qwen3.6-27B-FP8 · Official FP8
95
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
07
GLM-5.2 z-ai/glm-5.2 · Flagship open-weight
94
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
08
MiniMax M2.5 MiniMaxAI/MiniMax-M2.5 · Official FP8
94
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
09
Qwen3.8 27B Qwen/Qwen3.8-27B-FP8 · Official FP8
94
5/5
5/5 Wilson ≥56.6%
≥56.6%
1 ±1
10
Laguna S 2.1 poolside/Laguna-S-2.1-INT4 · Vendor INT4 (RC2)
93
5/5
5/5 Wilson ≥56.6%
≥56.6%
1
11
MiniMax M3 Community packnvidia/MiniMax-M3-NVFP4 · NVIDIA NVFP4
84
4/5
5/5 Wilson ≥56.6%
≥37.6%
1
12
MiniMax M3 1/5 exhaustedMiniMaxAI/MiniMax-M3-MXFP8 · Official MXFP8
80 1/5 exhausted
4/5
5/5 Wilson ≥56.6%
≥37.6%
1
13
Qwen3.6 27B Community packcyankiwi/Qwen3.6-27B-AWQ-INT4 · Community AWQ-INT4
65
2/5
5/5 Wilson ≥56.6%
≥11.8%
1 ±1
14
Qwen3.6 35B-A3B Qwen/Qwen3.6-35B-A3B-FP8 · Official FP8
63
2/5
5/5 Wilson ≥56.6%
≥11.8%
2 ±1
15
MiniMax M2.5 Community packQuantTrio/MiniMax-M2.5-AWQ · Community AWQ
57
3/5
3/5 Wilson ≥23.1%
≥23.1%
1
16
Step 3.7 Flash 2/5 exhaustedwarningstepfun-ai/Step-3.7-Flash-FP8 · Official FP8
46 2/5 exhausted
2/4 effective n=4
2/4 Wilson ≥15%
≥15%
1.5 ±1
17
Step 3.7 Flash Community pack3/5 exhaustedcyankiwi/Step-3.7-Flash-AWQ-INT4 · Community AWQ-INT4
10 3/5 exhausted
0/5
1/5 Wilson ≥3.6%
≥0%
2
info

SQE score blends harness components 45/25/15/10/5 (gate, pass, repair, style, token thrift) across independent reps. Gate and pass are each k/n over effective reps; gate also requires held-out honesty. Tainted reps are excluded and always disclosed. Best-effort estimates; read the methodology post before ranking models within a task.