Living benchmarkVersion 0.1

AI Go Problem Creation Benchmark

A reproducible evaluation of whether AI models can author original, legal, and teachable life-and-death Go problems on 19×19 boards in SGF. Every result links back to its generated position, solution tree, automated checks, and human review.

Models evaluated
11
Runs indexed
15
Evaluation slots
150
Human reviews
111
Latest update
Aug 08, 2026
UTC · indexed artifacts

BENCHMARK TRAJECTORY

All LLMs By Release Date

Each model is plotted once by public release date, using its best human-reviewed run and a difficulty-capped score out of 10.

  1. Claude Fable 5Anthropic / Jun 9, 20262/10
  2. Claude Opus 5Anthropic / Jul 24, 20262/10
  3. GPT-5.6 SolOpenAI / Jul 9, 20261/10

Model summary

Each row uses that model’s best human-reviewed run. Select a model to inspect its generated files and gate-by-gate evidence.

ModelHuman scoreAutomated gateStructuralRunsLatest best
Claude Fable 5Anthropic2/102.0 / 1010 / 1010 / 101Aug 08, 2026
Claude Opus 5Anthropic2/102.0 / 1010 / 1010 / 102Aug 07, 2026
GPT-5.6 SolOpenAI1/101.0 / 1010 / 1010 / 104Aug 07, 2026
Claude Opus 4.8Anthropic0/100.0 / 1010 / 1010 / 101Aug 08, 2026
DeepSeek V4 Flash 0731DeepSeek0/100.0 / 1010 / 1010 / 101Aug 07, 2026
GPT-5.5 (Codex)OpenAI0/100.0 / 1010 / 1010 / 101Aug 08, 2026
GPT-5.6 LunaOpenAI0/100.0 / 1010 / 1010 / 101Aug 07, 2026
Grok 4.5xAI0/100.0 / 1010 / 1010 / 101Aug 07, 2026
Muse Spark 1.2Meta0/100.0 / 1010 / 1010 / 101Aug 08, 2026
Kimi K3Kimi0/100.0 / 100 / 100 / 101Aug 08, 2026
Qwen3.8-MaxQwen0/100.0 / 100 / 101 / 101Aug 08, 2026
Score definitionHuman-passed problems require a difficulty rating. At most two 20–30 kyu and two 10–19 kyu problems receive credit; 5–9 kyu and harder problems are uncapped.

Latest benchmark activity

View all 15 runs
StartedModelEffortStructuralAutomated gateReviewedStatus
Aug 08, 2026, 15:25Claude Fable 5Anthropiclow10/1010/1010/10reviewed
Aug 08, 2026, 04:14GPT-5.6 SolOpenAImax10/1010/1010/10reviewed
Aug 08, 2026, 02:09Qwen3.8-MaxQwendefault1/100/101/10failed
Aug 08, 2026, 01:47Muse Spark 1.2Metadefault10/1010/1010/10reviewed
Aug 08, 2026, 01:44Kimi K3Kimidefault0/100/100/10failed
Aug 08, 2026, 01:40GPT-5.5 (Codex)OpenAIxhigh10/1010/1010/10reviewed

Methods and source material

Results are useful only when the task, artifacts, and review criteria remain inspectable.