Living benchmarkVersion 0.1
AI Go Problem Creation Benchmark
A reproducible evaluation of whether AI models can author original, legal, and teachable life-and-death Go problems on 19×19 boards in SGF. Every result links back to its generated position, solution tree, automated checks, and human review.
- Models evaluated
- 11
- Runs indexed
- 15
- Evaluation slots
- 150
- Human reviews
- 111
- Latest update
- Aug 08, 2026 UTC · indexed artifacts
BENCHMARK TRAJECTORY
All LLMs By Release Date
Each model is plotted once by public release date, using its best human-reviewed run and a difficulty-capped score out of 10.
Current results
Model summary
Each row uses that model’s best human-reviewed run. Select a model to inspect its generated files and gate-by-gate evidence.
| Model | Human score | Automated gate | Structural | Runs | Latest best |
|---|---|---|---|---|---|
| Claude Fable 5Anthropic | 2/102.0 / 10 | 10 / 10 | 10 / 10 | 1 | Aug 08, 2026 |
| Claude Opus 5Anthropic | 2/102.0 / 10 | 10 / 10 | 10 / 10 | 2 | Aug 07, 2026 |
| GPT-5.6 SolOpenAI | 1/101.0 / 10 | 10 / 10 | 10 / 10 | 4 | Aug 07, 2026 |
| Claude Opus 4.8Anthropic | 0/100.0 / 10 | 10 / 10 | 10 / 10 | 1 | Aug 08, 2026 |
| DeepSeek V4 Flash 0731DeepSeek | 0/100.0 / 10 | 10 / 10 | 10 / 10 | 1 | Aug 07, 2026 |
| GPT-5.5 (Codex)OpenAI | 0/100.0 / 10 | 10 / 10 | 10 / 10 | 1 | Aug 08, 2026 |
| GPT-5.6 LunaOpenAI | 0/100.0 / 10 | 10 / 10 | 10 / 10 | 1 | Aug 07, 2026 |
| Grok 4.5xAI | 0/100.0 / 10 | 10 / 10 | 10 / 10 | 1 | Aug 07, 2026 |
| Muse Spark 1.2Meta | 0/100.0 / 10 | 10 / 10 | 10 / 10 | 1 | Aug 08, 2026 |
| Kimi K3Kimi | 0/100.0 / 10 | 0 / 10 | 0 / 10 | 1 | Aug 08, 2026 |
| Qwen3.8-MaxQwen | 0/100.0 / 10 | 0 / 10 | 1 / 10 | 1 | Aug 08, 2026 |
Run log
Latest benchmark activity
| Started | Model | Effort | Structural | Automated gate | Reviewed | Status |
|---|---|---|---|---|---|---|
| Aug 08, 2026, 15:25 | Claude Fable 5Anthropic | low | 10/10 | 10/10 | 10/10 | reviewed |
| Aug 08, 2026, 04:14 | GPT-5.6 SolOpenAI | max | 10/10 | 10/10 | 10/10 | reviewed |
| Aug 08, 2026, 02:09 | Qwen3.8-MaxQwen | default | 1/10 | 0/10 | 1/10 | failed |
| Aug 08, 2026, 01:47 | Muse Spark 1.2Meta | default | 10/10 | 10/10 | 10/10 | reviewed |
| Aug 08, 2026, 01:44 | Kimi K3Kimi | default | 0/10 | 0/10 | 0/10 | failed |
| Aug 08, 2026, 01:40 | GPT-5.5 (Codex)OpenAI | xhigh | 10/10 | 10/10 | 10/10 | reviewed |
Audit the benchmark
Methods and source material
Results are useful only when the task, artifacts, and review criteria remain inspectable.