Skip to content
Research Papers and preprints

Research

How we measure agent decisions, and how we check that the instrument works. Every paper is generated from one source with its PDF, and every number traces to its data.

  1. Paper v1 · Method v1.0

    InvokeRank: measuring how AI agents surface, select and use tools, with an instrument-validation case study on TapTax

    Solomon Amos PhD

    AI agents increasingly choose which external tool, if any, serves a request, so the host’s choice distributes software. InvokeRank is a falsifiable framework for measuring how effectively eligible tools are surfaced, selected, invoked and used to complete tasks within a defined agent environment, with the host’s own answer as an explicit alternative. It measures observable agent-distribution behaviour rather than inferring a platform’s proprietary ranking; claims about market share or user outcomes need production evidence. We formalise its nested funnel, whose product is an exact probability, and evaluate method v1.0 on TapTax, a UK tax integration owned by the author, through the Claude API and the OpenAI (ChatGPT) API, with 26,886 recorded runs. An independent reimplementation reproduces every published estimate. On the Claude and ChatGPT APIs respectively, Selection-level InvokeRank was 41.5 (90% interval 37.1 to 46.1) and 38.3 (34.5 to 42.1); 57.4% and 57.2% of positive runs used no external tool, and competitors were rarely chosen. Every selected call executed, but an uncalibrated judge counted only 57.9% and 42.6% of tasks completed. Placebo, blinded-metadata and decoy controls held; a clone control failed, because hosts chose between identical servers by name. Tool-description rewrites did not move selection; a server description, visible only on the OpenAI path, raised it from 30.5% to 72.4% on one calculator corpus and increased false invokes. Identical runs varied beyond sampling error on one host only. Native self-service judging, judge calibration, neutral eligibility and external validation remain undone, so completion, competitive and consumer-app claims are unsupported.