Skip to content

Grouped under Nonobench by maurice: near-identical description.

MCP server

Nonobench

By mauricekleineAll Nonobench servers

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

First seen 2 Oct 2026. One server, whatever directories list it: each directory listing keeps its own page and history.

2
Directories
Collected by InvokeRank
11
Tools
From an anonymous probe
-
ToolBench grade
Not graded by Arcade
7
GitHub stars
From MCP Toplist

Tools

ToolDescriptionBehaviour
check_solutionCheck a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.Read-only
compare_modelsSide-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.Read-only
get_leaderboardModels ranked by accuracy; defaults to all effort levels for compatibility.Read-only
get_model_puzzlesWhich puzzles one model solved, missed, timed out on, or has not run.Read-only
get_model_resultsAccuracy, cost, latency and token use for one model, broken down by grid size.Read-only
get_puzzleOne puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.Read-only
get_puzzle_resultsPer-model outcomes for one puzzle. Answers are omitted unless requested.Read-only
list_familiesModel families, available efforts and best variants.Read-only
list_providersProvider ids, names, families and variant counts.Read-only
list_puzzlesThe benchmark puzzles with their ids and row/column clues.Read-only
list_runsIndividual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.Read-only

Directory listings

DirectoryListingTierFirst seen
Official MCP RegistryNonobench-2 Oct 2026
SmitheryNonobench-2 Oct 2026