Skip to content
Official MCP RegistryListed

Nonobench

Part ofNonobenchlisted on 2 directories

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

First seen 2 Oct 2026. Evidence as of 8 Oct 2026.

11
Tools
From an anonymous probe
1
Source listings
Each with its own history
0
Recorded changes
Since first seen

Tools

ToolDescriptionBehaviour
check_solutionCheck a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.Read-only
compare_modelsSide-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.Read-only
get_leaderboardModels ranked by accuracy; defaults to all effort levels for compatibility.Read-only
get_model_puzzlesWhich puzzles one model solved, missed, timed out on, or has not run.Read-only
get_model_resultsAccuracy, cost, latency and token use for one model, broken down by grid size.Read-only
get_puzzleOne puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.Read-only
get_puzzle_resultsPer-model outcomes for one puzzle. Answers are omitted unless requested.Read-only
list_familiesModel families, available efforts and best variants.Read-only
list_providersProvider ids, names, families and variant counts.Read-only
list_puzzlesThe benchmark puzzles with their ids and row/column clues.Read-only
list_runsIndividual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.Read-only

Change history

No changes since the first observation. The first snapshot is the baseline.

Source listings
SourceListingFirst seenLast seenVersions
Official MCP Registryio.github.mauricekleine/nonobench2 Oct 20268 Oct 20261