Official MCP RegistryListed
FitLLM
Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.
First seen 2 Oct 2026. Evidence as of 6 Oct 2026.
3
Tools
From an anonymous probe
1
Source listings
Each with its own history
0
Recorded changes
Since first seen
Tools
| Tool | Description | Behaviour |
|---|---|---|
| check_llm_fit | Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled). | Read-only |
| list_supported | List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected. | Read-only |
| what_fits_on_hardware | Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model. | Read-only |
Change history
No changes since the first observation. The first snapshot is the baseline.
| Source | Listing | First seen | Last seen | Versions |
|---|---|---|---|---|
| Official MCP Registry | run.fitllm/fitllm | 2 Oct 2026 | 6 Oct 2026 | 1 |