Skip to content
InvokeRank
Index snapshot 1 Oct

Methodology

Score method v1 · published 1 Oct 2026

InvokeRank measures one thing: when a person asks an AI agent for work your product can do, how often does the agent complete that work through your product? Every number on this site says where it came from, how many runs stand behind it, and how sure we are.

The funnel

  • Surfaced: your tool entered the agent’s candidate set (tool search result or deferred load).
  • Selected: the agent chose your tool first among competing tools.
  • Invoked: the call had valid arguments and was not blocked.
  • Executed: the tool returned a non-error, non-empty result.
  • Completed: deterministic checks and a calibrated judge agree the task was done.

The score

Each stage is conditional on the one before, so multiplying the five rates gives the probability that an eligible request is completed through your product. InvokeRank is 100 times that probability, shown with a 90% interval. We do not weight stages: a 10% loss at any stage removes the same 10% of completed tasks. Where two intervals overlap we say the two are statistically tied and do not rank them.

Evidence

Every figure carries its evidence class: R registry data, S controlled runs through the OpenAI and Anthropic APIs, H analytics a host gave you, T your server’s telemetry, O business outcomes, A human-observed runs. Classes are never pooled silently. API runs approximate the consumer apps; they are labelled as such, and we never automate chatgpt.com or claude.ai.

Fair play

We never recommend tool copy that steers a model toward you over other tools. Both OpenAI and Anthropic forbid it, and it raises false invocations, which we measure on negative prompts in every experiment.