Grouped under Nonobench by maurice: near-identical description.
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
Listed on
First seen 2 Oct 2026. One server, whatever directories list it: each directory listing keeps its own page and history.
2
Directories
Collected by InvokeRank
11
Tools
From an anonymous probe
-
ToolBench grade
Not graded by Arcade
7
GitHub stars
From MCP Toplist
Tools
| Tool | Description | Behaviour |
|---|---|---|
| check_solution | Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match. | Read-only |
| compare_models | Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names. | Read-only |
| get_leaderboard | Models ranked by accuracy; defaults to all effort levels for compatibility. | Read-only |
| get_model_puzzles | Which puzzles one model solved, missed, timed out on, or has not run. | Read-only |
| get_model_results | Accuracy, cost, latency and token use for one model, broken down by grid size. | Read-only |
| get_puzzle | One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions. | Read-only |
| get_puzzle_results | Per-model outcomes for one puzzle. Answers are omitted unless requested. | Read-only |
| list_families | Model families, available efforts and best variants. | Read-only |
| list_providers | Provider ids, names, families and variant counts. | Read-only |
| list_puzzles | The benchmark puzzles with their ids and row/column clues. | Read-only |
| list_runs | Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output. | Read-only |