Skip to content
Official MCP RegistryListed

FetchSandbox

A deterministic verification engine for agents. Proves a fix: fails on the old code, passes on new.

First seen 2 Oct 2026. Evidence as of 6 Oct 2026.

7
Tools
From an anonymous probe
1
Source listings
Each with its own history
4
Recorded changes
Since first seen

Tools

ToolDescriptionBehaviour
coachConversational integration coach for FetchSandbox. Server-side orchestrator that walks the user through adding an API integration (payments / email / auth / etc.) — intake the goal, elicit domain-aware discovery questions from the spec's brain.yaml, route to the right workflow, prove the contract via FetchSandbox, surface compliance notes. Call this BEFORE any other FetchSandbox tool when the user has an open-ended 'help me add X', 'integrate X', 'test my X integration' ask. BEHAVIOR — strict, do exactly this each turn: If next_action=act_in_app, carry out next_actions with your platform's code, configuration and test tools, then return the results to the same validation session. Do not ask the human to choose routine fault tests, supply missing fixture data you can generate, or copy tool output. Ask only for genuine access, secret-entry constraints or ambiguous business rules. Never claim unsupported checks passed. (1) Say `message_for_user` to the user (verbatim or lightly paraphrased to fit your voice — but don't add new content). (2) If `next_action=wait_for_user` AND `options` is present + non-empty: USE YOUR CLIENT'S NATIVE QUESTION-PICKER TOOL (in Cursor / Claude Code this is `AskUserQuestion`) to render the lettered picker with the `question` text, the `options[].label` as rows, `default_option` as default, and an 'Other...' freeform row when `allow_freeform=true`. When the user picks or types, call coach again with `{session_id, user_response: <picked value or freeform text>}`. (3) If `next_action=wait_for_user` AND no `options`: just wait for free text. (4) If `next_action=call_tool`: invoke `tool_call.tool` with `tool_call.args`, then call coach again with the result in `context`. (5) If `next_action=done`: end the session. For a hosted app-builder request to build or test an integration, keep the user's simple goal intact and let the server route all named providers into one validation session. When FetchSandbox returns twin bindings, treat those as temporary TEST credentials: configure them server-side in the app-builder project, never ask the user for live Paddle/Resend keys, and never claim they were applied until the builder reports the change and the verifier observes app traffic. Infer app configuration details from the project when available; ask the user only for a fact the builder cannot discover. The deterministic verifier, not this conversation, owns the verdict. The state machine is server-side — DON'T try to predict the next step or skip ahead; let the server drive.Read-only
guideROUTE a symptom to the provider behaviour that explains it. Use this when the user names a provider or a domain (payments, email, auth, SMS, subscriptions) and you need to know what that provider ACTUALLY does — not what its docs say, and not what can be inferred from reading the integration code. Reading the code tells you what your app does with a field. It cannot tell you what the field MEANS at the provider — whether a line item is a seat count, whether a 200 body carries ok:false, whether an event can arrive out of order. That is the class of bug this routes. Returns {spec, workflow, scenario, confidence, reasoning} plus `next_actions`: a typed list of what to call next, with arguments pre-filled. Follow it rather than improvising the next step.Read-only
list_workflowsList the named, runnable workflows for a previously-imported spec. Workflows are realistic multi-step API journeys (e.g. 'create customer → attach payment method → create subscription'). Use this after import_spec for exploration ("what can I do?", "show me the flows") OR before run_all_workflows when the user wants a SCOPED validation: list, filter by user intent ("checkout", "webhooks"), then pass the matching ids as `workflow_names` to run_all_workflows. Returns the workflows AND the failure scenarios this spec can simulate — each with a plain description and the exact call to run it. Read the scenarios out to the user and ask which matter; do not guess scenario names.Read-only
quickrunRun a curated proof workflow against a KNOWN, bundled spec (stripe, clerk, descope, resend, twilio, and 50+ others) in ONE call — it spins up the sandbox by slug, so you do NOT need import_spec or a sandbox_id first. This is the normal path when the user is testing an integration with a well-known provider: call `guide` on their prompt, then call `quickrun` with the returned spec + workflow. To reproduce a failure, pass the `scenario` from guide's matched_bug_pattern.reproduce_with.scenario (e.g. webhook_retries, payment_declined). Returns sandbox_id + flow_run_id — pass BOTH to verify_behavior to prove the fix — plus a receipt URL. Use run_workflow instead ONLY when you already hold a sandbox_id from import_spec of a custom/private spec.Changes data
run_workflowExecute ONE specific workflow by name and return its step-by-step trace PLUS a `share_url` — a public, replayable proof URL that renders the full timeline (every request, response, webhook event) for this run. The share_url is the canonical 'here's what happened' artifact: surface it verbatim in any reply that needs evidence (PR comments, Slack threads, blog posts, X replies). Do NOT substitute a docs URL or any other link as the proof — the share_url is the only valid receipt. Use ONLY when the user explicitly names a single workflow to run (e.g., "run accept_payment", "just check the refund workflow"). For ANY validation-style request — "validate stripe", "check coverage", "run all workflows", "test this integration", or even "validate stripe checkout" (multiple workflows match "checkout") — use `run_all_workflows` instead. The batch tool collapses N approvals to 1 and supports a workflow_names filter for scope. Calling this in a loop is an anti-pattern.Changes data
validate_integrationFor hosted application_workflow_v1 apps, build and publish the disabled test adapter before requesting short-lived credentials. After installing the existing handoff and republishing, call preflight:true with session_id/run_id. A setup_blocked response is not a failed application test. Its next_tool_call is null: repair its exact host-side blocker first, then call resume_tool_call for the SAME run/preflight without rotating credentials. Do not poll an unchanged setup blocker. Execution has a separate deadline. Never execute twice; retrieve status with session_id/run_id only. Browser payments require an available browser or one explicit user action, and independently observed paid state; a click is not proof. After proof, remove only temporary settings, republish and verify disabled externally. For every app, inspect capabilities and application_actions. Prepare missing test products, prices and related records before failures: author fixtures using documented provider operations, include a final GET with an exact identity assertion, then configure the app with returned fixture IDs. The server uses the existing twin workflow runner and independently checks read-back. Fixture readiness is not app proof. Unsupported application checks must be reported explicitly, not replaced by a separate script or quickrun. Send application_context with the original goal and inspected workflow to obtain model-proposed rules, missing observations and fixture requirements. Proposals are not executed checks. Preserve all proposed rules in the coverage report, even when unsupported. If server_requirements_model_available is false or planning returns requirements_model_unavailable, inspect the app and author the candidate contract with your own model; label it builder-proposed, not a FetchSandbox model result. Never wait for the human to write routine test requirements. For a normal HTTP app workflow with independently readable state, configure application_workflow_v1 and application_config. Declare a positive baseline first, normal app actions, GET observations and required exact assertions. Reuse the proposed rule IDs; omitted rules remain unmeasured. Signed provider event actions are supported through event_destinations and isolated secure-handoff signing secrets; bind payload identity to actual app responses and verify saved state. No browser, arbitrary auth-header or SQL execution is supported by this HTTP contract: record those gaps explicitly. Do not weaken signature or access checks to make a test pass. Start a validation session for the user's integration. quickrun and run_workflow execute OUR curated workflow against a twin — green there says the provider behaves as documented, and says nothing about whether their integration wired to it correctly. This instead hands back a FRESH twin per provider and asks the user to point their app at it, then reads external provider traffic. Direct agent calls also appear there; caller identity is not established. Use it when someone wants their integration verified before launch. Call once with `providers`, then follow the returned `next_action`. Inspect the app and supply a complete `receipt_config` before configuring the suite; the application suite drives its own purchases and webhooks, so do not ask the user to run a manual flow first. The matrix is returned before any traffic: pass `arm` with a probe ID before exercising the app. Send the returned request_headers on every provider request in that attempt, then pass `probe` and `run_id` to evaluate. Use `cancel: true` with `run_id` to abort or retry failed cleanup. Request logs alone do not prove application state or production readiness. For a supported payment-to-receipt integration, follow the structured `next_action`; call `next_tool_call` only when it is non-null and its arguments are complete. Never execute while host-side setup is pending. Do not substitute guide, quickrun, list_workflows, or reference probes for application verification. Use the widest suite the server recommends. receipt_recovery_v1 includes R1-R4 and adds checks R5-R8 for accepted-send response loss, concurrent redelivery, source/email transient failures and delayed redelivery. Pass suite plus receipt_config with the app's exact webhook URL, test recipients and observable purchase marker. FetchSandbox freezes all four receipt rules for both same-customer and different-customer purchases. You are responsible for completing this setup: treat the target as a disposable test app, not the user's live production service, until its declared suite passes. For a greenfield app, keep live billing disabled and publish a temporary test build first to obtain its stable public origin. Use the published webhook URL, not a private preview URL. Inspect the app and identify the exact webhook route before suite configuration. Before requesting or applying test secrets, send an unsigned POST to that exact URL from outside the workspace; a 400 signature rejection is expected, while a 307/login redirect is a blocker. Keep signature validation enabled. After the successful preflight, configure the suite with its complete receipt_config. Apply the returned twin URLs, X-Flow-Run-Id, and three temporary credentials only to this disposable app; execute only after the host-side setup actions are complete. The credentials are twin-only test values: PADDLE_API_KEY, RESEND_API_KEY, and PADDLE_WEBHOOK_SECRET. Use the platform's supported secret store or secret-writing tool and never use real provider credentials or enable live billing. After every required check is held, remove only the test secrets and twin config this run added, including from the disposable test app's published scope. Never modify a separate live production app or delete pre-existing, shared, or real-provider values. Confirm the temporary names are absent, then hand the app to its owner to configure separate development/staging credentials. If the suite is failing or inconclusive, retain test config for repair and a fresh run. The configure_app response contains public twin URLs/IDs and a secret_handoff_url, never the secret values themselves. If the host exposes no safe way for you to write test secrets, explain that specific limitation and show the human the link to open while signed in to the owning FetchSandbox account; they can copy the three values directly into the platform secret store. Do not ask the human to discover endpoints, call FetchSandbox manually, or paste secret values into chat, project files, or logs. Wait for confirmation only when the human must complete that secure-store action. Then call execute:true with its run_id. FetchSandbox drives signed events and independently observes accepted email records. Keep every check and the observation windows in your report. Restore any temporary host privacy setting after the test. On an inconclusive execute result, inspect and report its `execution_diagnostics` field directly; do not ask the human to infer a cause from the public receipt page, which intentionally omits private proof. If you cannot configure the app, report that blocker; reference tests cannot replace app execution. A verified receipt suite covers only its declared rules and windows, never all production behavior. A missing or untriggered recovery fault stays unmeasured. This application suite requires sign-in. Paddle checkout is a separate app-owned browser flow: if a transaction's `checkout.url` points to `fetchsandbox.com/checkout`, returns 404, or does not show a checkout page, do not imply FetchSandbox hosts the merchant's payment UI. Inspect the exact URL and the Paddle Sandbox default payment link or per-transaction override. Tell the builder to use a reachable app checkout page that loads Paddle.js and opens the `_ptxn` transaction, or a supported Paddle-hosted flow. Verify browser checkout/payment separately from twin API and webhook tests; transaction creation alone is not payment proof.Changes data
verify_behaviorProve a known fix survives a bug — the 'prove' half of reproduce→prove. The backend spawns a buggy AND a fixed reference handler and fires the bug_pattern's probes at both, returning the side-by-side diff (e.g. the buggy handler double-charges on a duplicate webhook, the fixed handler dedupes). Call this AFTER run_workflow reproduces a failure, when the matched bug_pattern has a simulation block, to show the fix actually holds — not just that the failure reproduced. Pass sandbox_id + flow_run_id from the run so the diff is saved onto that run's receipt URL. bug_pattern_id comes from guide's matched_bug_pattern. The buggy/fixed handlers are FetchSandbox reference implementations, NOT the user's code. A confirmed reference result does not verify the user's app or show that it is ready to ship. Keep the returned provider and evidence scope with any reported result.Changes data

Change history

  1. validate_integration: input schema changed (+application_config, +application_context, +fixtures, +preflight)
  2. validate_integration: description changed (+"For hosted application_workflow_v1 apps, build and publish the disabled test adapter before requesting short-lived credentials. After installing the existing handoff and republishing, call preflight:true with session_id/run_id. A setup_blocked response is not a failed application test. Its next_tool_call is null: repair its exact host-side blocker first, then call resume_tool_call for the SAME run/preflight without rotating credentials. Do not poll an unchanged setup blocker. Execution has a separate deadline. Never execute twice; retrieve status with session_id/run_id only. Browser payments require an available browser or one explicit user action, and independently observed paid state; a click is not proof. After proof, remove only temporary settings, republish and verify disabled externally. For every app, inspect capabilities and application_actions. Prepare missing test products, prices and related records before failures: author fixtures using documented provider operations, include a final GET with an exact identity assertion, then configure the app with returned fixture IDs. The server uses the existing twin workflow runner and independently checks read-back. Fixture readiness is not app proof. Unsupported application checks must be reported explicitly, not replaced by a separate script or quickrun. Send application_context with the original goal and inspected workflow to obtain model-proposed rules, missing observations and fixture requirements. Proposals are not executed checks. Preserve all proposed rules in the coverage report, even when unsupported. If server_requirements_model_available is false or planning returns requirements_model_unavailable, inspect the app and author the candidate contract with your own model; label it builder-proposed, not a FetchSandbox model result. Never wait for the human to write routine test requirements. For a normal HTTP app workflow with independently readable state, configure application_workflow_v1 and application_config. Declare a positive baseline first, normal app actions, GET observations and required exact assertions. Reuse the proposed rule IDs; omitted rules remain unmeasured. Signed provider event actions are supported through event_destinations and isolated secure-handoff signing secrets; bind payload identity to actual app responses and verify saved state. No browser, arbitrary auth-header or SQL execution is supported by this HTTP contract: record those gaps explicitly. Do not weaken signature or access checks to make a test pass." -"test recipients and observable purchase marker. FetchSandbox freezes all four receipt rules for both same-customer and different-customer purchases. You are responsible for completing this setup: treat the target as a disposable test app, not the user's live production service, until its declared suite passes. For a greenfield app, keep live billing disabled and publish a temporary test build first to obtain its stable public origin. Use the published webhook URL, not a private preview URL. Inspect the app and identify the exact webhook route before suite configuration. Before requesting or applying test secrets, send an unsigned POST to that exact URL from outside the workspace; a 400 signature rejection is expected, while a 307/login redirect is a blocker. Keep signature validation enabled. After the successful preflight, configure the suite with its complete receipt_config. Apply the returned twin URLs, X-Flow-Run-Id, and three temporary credentials only to this disposable app; execute only after the host-side setup actions are complete. The credentials are twin-only test values: PADDLE_API_KEY, RESEND_API_KEY, and PADDLE_WEBHOOK_SECRET. Use the platform's supported secret store or secret-writing tool and never use real provider credentials or enable live billing. After every required check is held, remove only the test secrets and twin config this run added, including from the disposable test app's published scope. Never modify a separate live production app or delete pre-existing, shared, or real-provider values. Confirm the temporary names are absent, then hand the app to its owner to configure separate development/staging credentials. If the suite is failing or inconclusive, retain test config for repair and a fresh run. The configure_app response contains public twin URLs/IDs and a secret_handoff_url, never the secret values themselves. If the host exposes no safe way for you to write test secrets, explain that specific limitation and show the human the link to open while signed in to the owning FetchSandbox account; they can copy the three values directly into the platform secret store. Do not ask the human to discover endpoints, call FetchSandbox manually, or paste secret values into chat, project files, or logs. Wait for")
  3. coach: description changed (+"If next_action=act_in_app, carry out next_actions with your platform's code, configuration and test tools, then return the results to the same validation session. Do not ask the human to choose routine fault tests, supply missing fixture data you can generate, or copy tool output. Ask only for genuine access, secret-entry constraints or ambiguous business rules. Never claim unsupported checks passed.")
  4. validate_integration: description changed
Source listings
SourceListingFirst seenLast seenVersions
Official MCP Registryio.github.fetchsandbox/mcp2 Oct 20266 Oct 20261