Official MCP RegistryListed
Nonobench
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
First seen 2 Oct 2026. Evidence as of 8 Oct 2026.
11
Tools
From an anonymous probe
1
Source listings
Each with its own history
0
Recorded changes
Since first seen
Tools
| Tool | Description | Behaviour |
|---|---|---|
| check_solution | Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match. | Read-only |
| compare_models | Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names. | Read-only |
| get_leaderboard | Models ranked by accuracy; defaults to all effort levels for compatibility. | Read-only |
| get_model_puzzles | Which puzzles one model solved, missed, timed out on, or has not run. | Read-only |
| get_model_results | Accuracy, cost, latency and token use for one model, broken down by grid size. | Read-only |
| get_puzzle | One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions. | Read-only |
| get_puzzle_results | Per-model outcomes for one puzzle. Answers are omitted unless requested. | Read-only |
| list_families | Model families, available efforts and best variants. | Read-only |
| list_providers | Provider ids, names, families and variant counts. | Read-only |
| list_puzzles | The benchmark puzzles with their ids and row/column clues. | Read-only |
| list_runs | Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output. | Read-only |
Change history
No changes since the first observation. The first snapshot is the baseline.
| Source | Listing | First seen | Last seen | Versions |
|---|---|---|---|---|
| Official MCP Registry | io.github.mauricekleine/nonobench | 2 Oct 2026 | 8 Oct 2026 | 1 |