syndicAI Quality Eval
The SQE Leaderboard
We measure the build you actually run: exact quant packs on real hardware, gated by tests, with every run log published. No vendor harness, no cherry-picked shots.
Methodology v2 (August 2026): pass and gate are scored as separate weighted components. See the changelog note for what changed and why overfit-heavy packs moved.
Build a Connect-Four-style game from a written spec: board size, start player, and win-rule traps included.
SQE score blends harness components 45/25/15/10/5 (gate, pass, repair, style, token thrift) across independent reps. Gate and pass are each k/n over effective reps; gate also requires held-out honesty. Tainted reps are excluded and always disclosed. Best-effort estimates; read the methodology post before ranking models within a task.