Skip to content
MCP server

Nodegrove VRAM: can I run it?

By nodegroveAll Nodegrove servers

Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.

Listed on

First seen 3 Oct 2026. One server, whatever directories list it: each directory listing keeps its own page and history.

2
Directories
1 via MCP Toplist
6
Tools
From an anonymous probe
-
ToolBench grade
Not graded by Arcade
0
GitHub stars
From MCP Toplist

Tools

ToolDescriptionBehaviour
can_i_runCan this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.Read-only
estimate_from_hf_repoReads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.Read-only
estimate_vramHow much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).Read-only
list_gpusThe GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.Read-only
list_modelsThe open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.Read-only
what_fitsWhich open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.Read-only

Directory listings

DirectoryListingTierFirst seen
Official MCP RegistryNodegrove VRAM: can I run it?-3 Oct 2026
GlamaListed there according to MCP Toplist’s dataset; not collected by InvokeRank.