Skip to content
Official MCP RegistryListed

Cloud World Model

Simulate cloud architectures before provisioning. 9 free demo tools; an API key unlocks all 62.

First seen 2 Oct 2026. Evidence as of 5 Oct 2026.

9
Tools
From an anonymous probe
1
Source listings
Each with its own history
6
Recorded changes
Since first seen

Tools

ToolDescriptionBehaviour
scenario.getHydrate one built-in scenario from the live Cloud World Model scenario library. Prerequisite: a scenario id returned by scenario.list. Returns the complete selected scenario graph, including resources and connections plus optional seed, resilienceConfig, protectedResilienceConfig, traffic/failure presets, named traffic-phase summaries, activeFailurePhases and optionalFailurePhases (type, resource/zone target, severity, step range), retry-workload disclosure, and real-world incident metadata. The response includes both title and name for compatibility; pass resources and connections, and optionally seed/resilienceConfig, to simulation.create when you need to edit or inspect the graph. For the shorter handoff, pass the id as scenarioId instead. The likely next tool is simulation.create.Read-only
scenario.listList the built-in demo scenarios as compact catalog cards — stable IDs, title/name, description, difficulty, tags, category, duration, provider summary, resource/connection counts, named active/optional traffic phases, and retry-workload disclosure. Use it as the first call when you want a ready-made architecture instead of designing one; the cards intentionally omit resource, connection, traffic-pattern, and failure-injection graphs. Anonymous discovery includes only scenarios with at most 10 resources so every listed card is demo-creatable. No prerequisites. Optionally narrow discovery with provider, category, and/or difficulty filters; omit them to receive the complete demo-creatable catalog. Pass a returned id as scenarioId to simulation.create for server-side expansion, or pass it to scenario.get when you need to inspect the full graph. Larger scenarios require an authenticated session. Returns named activeFailurePhases and optionalFailurePhases with type, resource/zone target, severity, and step range. No API key required. The likely next tool is scenario.get.Read-only
simulation.createCreate a temporary anonymous demo cloud simulation from a list of resources and connections (max 2 active simulations per client, up to 10 resources; the returned simulationId is a short-lived unguessable capability that survives MCP transport teardown, but it is cleaned up when the demo lifetime expires or the simulation is deleted). No API key required for this temporary anonymous demo operation. Built-in scenario workflow: call `scenario.list` and pass a returned card's `id` as `scenarioId` to `simulation.create` for server-side graph expansion. For full control, call `scenario.get` and pass its hydrated `resources` and `connections` arrays instead. These are two alternatives — do not send `scenarioId` with `resources` or `connections`. For the catalog EKS Spot Interruption Migration scenario, you may set `scenarioOverrides: { eksSpotInterruption: { startupSeconds } }` with an integer startupSeconds from 0 through 3600 to test a different readiness deadline without copying the graph; this override requires scenarioId and is mutually exclusive with resources and connections. `scenario.list` returns graph-free cards with bounded active/optional traffic-phase and retry-workload summaries; it is not a source of resource, connection, traffic-pattern, or failure-injection graphs. Scenario traffic and failure presets are not applied automatically. Use it to start any simulation workflow — either with hydrated resources and connections from scenario.get or your own architecture. Do not use it to modify an existing simulation (use simulation.inject_traffic to change load). For the exact owned typical fit, set appWeight:'typical', location.regionKey:'us-east-2' on all four AWS nodes, one ALB with serviceFamily:'alb' and loadBalancerScheme:'internal', two m5.large compute apps with workload:'crud-typical', appRuntime:'node', appWorkerCount:2, appDbPoolSize:250, one db.r5.large MySQL with workloadDatabaseEngine:'mysql', workloadDatabaseVersion:'8.0', maxConnections:500, ALB→each app→DB connections, minInstances=maxInstances=2, autoscaling:false, traffic 20–300 RPS. Only typical-fit-20, typical-fit-100 and typical-fit-200 tuned the typical-v1-20260927c/6aa574d7ff9d3080b88b221bcd59f7d218ae37f0 fit; 300 is an independent holdout and 500 is diagnostic only. Other typical workloads are modeled, not owned. P99 has distinct per-percentile provenance: latencyP99Basis identifies the owned in-VPC internal-ALB fit, a scaled-from-fit estimate (not directly measured), or an uncalibrated generic model. On the exact healthy lean owned graph at 10–1,000 offered target RPS, P99=max(final P95, 7.021919127633514 + 0.00027013891327780484*T) ms; only 10/100/500 RPS were fit, 1,000 RPS was held out. Lean M5 scaling is not a new measurement; typical/heavy and active failures retain uncalibrated P99. predictionEvidence.latencyP99 has measured 0.65–1.35×, scaled 0.50–1.50× (beyond 1,000: 0.25–2×), or uncalibrated 0.50–2× (beyond: 0.25–3×) assumption bounds centered on final P99. These are not confidence intervals or provider measurements. Historical evidence may omit P99. latencyBasis describes the general modeled latency path; use latencyP99Basis specifically for P99. P99 is diagnostic, not scored. For compute, set characteristics.capacityRps for an explicit per-node RPS ceiling at which CPU reaches ~95%; do not use maxThroughput for that compute contract. Kubernetes rejects capacityRps: set maxThroughput for the total cluster RPS ceiling, or nodePools[].maxThroughput for per-node pool capacity. Omitted compute capacityRps uses the selected catalog tier and can intentionally produce a stressed baseline (for example, the AWS m5.large catalog denominator is 2,000 RPS); for a healthy, capacity-bounded compute experiment, declare an explicit per-node capacity such as 500 RPS. That value is an experiment control, not a universal hardware fact. For OCI flexible compute shapes, pass the documented positive integer characteristics.ocpus explicitly; VM.Standard.E4.Flex accepts 1–64 OCPUs and each OCPU maps to 2 vCPUs. OCPU count establishes capacity dimensions only, not provider-specific performance, throughput, or price. An uncounted flexible shape remains an unverified generic estimate. Check GET /api/prediction/generic-shapes for the catalog-derived generic fallback inventory. New prediction-only GCP standard capacity entries include e2-standard-2/4/8/16/32, n1-standard-1/2/4/8/16/32/64/96, and n2-standard-16/32/48/64/80/96/128; provider specifications establish vCPU/memory dimensions, not CWM performance or pricing. Version 1 predictionEvidence explicitly reports legacyGeneric at the top level and on every appCpuByResource item; its note identifies each generic resource's shape and fallback reason. For generic fixed compute, characteristics.instanceCount accepts integer 1–100 represented VMs; capacity aggregates and CPU is per VM. Do not combine it with autoscaling:true, minInstances, or maxInstances. Aurora Serverless v2 remains limited to 1 with multiAz:false or 2 with multiAz:true. For Aurora Serverless ACU limits, use characteristics.config.minCapacity/maxCapacity or flat characteristics.minCapacity/maxCapacity. The exact AWS database shape with serviceFamily: 'aurora-serverless' and size: 'db.serverless' also accepts flat characteristics.minAcu/maxAcu; those aliases are rejected elsewhere, including at the resource root or inside config. For that shape, multiAz:true with instanceCount:2 creates a separately billable reader (<writer-id>-reader) in another AZ. Inspect returned resources and metrics before using simulation.step to observe modeled failover; no AWS timing guarantee is implied. For database connection budgets, set characteristics.connectionDemand on a database: {mode:'declared',declaredConnections:240} uses that plan-time demand without RPS; {mode:'max',declaredConnections:240,idlePoolFloor:200} takes the maximum of load-derived demand and the declared/floor values; omitted configuration preserves load-derived behavior. declaredConnections and idlePoolFloor are ASSUMPTIONS / plan-time budgets (for example, replicas × per-pod pool size), not observed live DB connections. Set maxConnections to the usable limit you intend to test. Per-database metrics report connectionDemandMode, loadDerivedConnections, declaredConnections/idlePoolFloor, and modeledConnections; cost and DB CPU/latency remain based on existing load-driven behavior. Demand above the usable limit adds a bounded, rule-based pool-saturation error signal; it is not a provider-calibrated rate. Capacity, node-bound, SKU, and autoscaling values supplied through this MCP tool are recorded as agent-supplied in the immutable normalizationReceipt; request responseMode: 'full' to inspect it. Generic GKE telemetry and recovery apply only to worker nodes; the control-plane management fee is cost-only, with no modeled control-plane CPU, API throttling, or cooldown. To bound the autoscaled fleet size, set the top-level maxInstances / minInstances parameters. If you do not set maxInstances, the engine uses the provider default — AWS 50, GCP 15, Azure/OCI/DigitalOcean 10 — which may be much larger than your intended fleet size. The response includes effectiveMaxInstances / effectiveMinInstances so you can confirm the bounds that will be enforced. For a targeted CPU HPA scale-out threshold, send the canonical autoscalingTargetCpu field in this create call (for example, autoscalingTargetCpu: 70 for GKE). The compatible aliases scaleOutCpuThreshold, scaleOutCpuPercent, and autoscaleTargetCpuPercent are also accepted; if more than one is sent, their values must agree. Every create response includes hpaAudit with the supplied field, persisted thresholds, and any provider default. For ECS Fargate CPU-only target tracking, set ecsCpuTargetTracking: true, autoscalingTargetCpu, minInstances/maxInstances, and optional scaleOutCooldownSeconds/scaleInCooldownSeconds with simulationSecondsPerStep (default 1). Inspect autoscalingConfig in the compact response or applicationAutoscalingPolicy in the full response. Latency and throughput do not trigger ECS scaling in this mode. These four TOP-LEVEL fields are simulation-wide — the engine applies one CPU threshold identically to every resource's scale decision by default. To make ONE resource scale at a different CPU target than the rest of the simulation (e.g. a GKE cluster scaling out at 60% while an EC2 fleet in the same simulation scales out at 80%), set characteristics.scaleOutCpuThreshold and/or characteristics.scaleInCpuThreshold on that specific resource instead — the per-resource value wins over the simulation-wide default for that resource only. A misnamed near-miss field nested under characteristics (e.g. targetCPUUtilizationPercentage) is rejected with a 400 explaining the correct field name — it is never silently dropped and defaulted. Responses are compact by default: id, name, status, traffic, and a per-resource summary (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, and recoveryBlockedReason when provided). Pass responseMode: 'full' to get the complete simulation object instead. During a failure workflow, lower traffic to serviceable levels before calling simulation.recover_resource, then use simulation.step until the recovered resource is healthy. Recovery progress is included when applicable: recoveryProgress.state is parked, cooling_down, or healthy, and its parkWindow/cooldown objects report totalSteps, completedSteps, remainingSteps, target, and requiredSteps. Poll simulation.step until state is healthy, then use simulation.metrics to inspect the resulting state and metrics. No prerequisites. Returns the created simulation's id, which every other simulation.* tool consumes; the new simulation also becomes this session's current simulation, so subsequent per-simulation tools may omit simulationId. The likely next tool is simulation.step to advance time. Do not call api.spec to learn the simulation workflow — the tool descriptions in this session contain everything needed. Authenticate with an API key for unlimited persistent simulations.Changes data
simulation.deletePermanently delete an owned temporary anonymous demo simulation and its metrics, events, failures, and capability. This is the explicit way to free a simulation slot; deletion is irreversible, while the existing demo TTL remains the safety net for abandoned simulations. Prerequisite: a simulationId from simulation.create, or an active simulation in the preserved MCP session. The likely next tool is simulation.create to use the freed slot. A successful response is { deleted: true, id }; failed ownership checks do not delete or revoke anything. Authenticate with an API key to unlock all 63 tools and persistent simulation management.Destructive
simulation.inject_failureFail one compute/Kubernetes node or database in a temporary anonymous demo simulation (marked critical, not removed). Targeted database quick failures are bounded and reversible; with no serving database at positive load, simulation.step reports 100% errors, zero goodput and errorBreakdown.dbFailure: 100. simulation.recover_resource can restore the database early. An AWS Aurora writer with characteristics.auroraStandbyResourceId pointing to a healthy related replicaOf database has an opt-in modeled failover: the first step is unavailable, the second shows the standby serving without a residual writer-outage penalty. MultiAz or an unrelated second database alone does not establish a standby. Read metrics.databases[].auroraFailover and the promotion event for the failed and standby IDs and success, plus errorRate and throughput. Each quick step is one modeled simulation second, so second-step promotion is one second after injection. Chaos database_crash samples every 10 seconds and defaults to a 30-second promotion and a 1800-second sole-writer restart after its injection duration; compare matching phases, not equal step indices or wall-clock times. These are deterministic assumptions, not observed AWS behavior. No API key required for this temporary anonymous demo operation. Exact targeting: pass resourceName (human-readable name, e.g. 'app-server-01'; exact match preferred, an unambiguous prefix is accepted) or resourceId to fail a specific resource — including an individual named instance, not only a group. If resourceName matches multiple resources the call fails with a 400 listing every matching candidate by name — retry with one exact name (or its resourceId) from that list. The database must be healthy; an already failed database returns a 400 describing its current status. When neither parameter is supplied, a RANDOM healthy compute/Kubernetes node is selected (not a database) — this path is non-deterministic and NOT suitable for controlled scenarios or replay; always target by name/id when reproducing a precise fault sequence. The response always echoes the applied outcome via resolvedResourceId, resolvedResourceName, and previousHealth (populated from the selected resource on the random path too). Network, storage, cache, queue, and security resource types are not supported by quick injection. For typed database_overload use authenticated failure.create (unavailable anonymously); instance_kill PERMANENTLY removes the instance — failure.delete does not restore it; use instance_down instead for a reversible single-node outage. Returns the updated resource list and the failure event that was logged. The likely next tool is simulation.step to observe how the architecture degrades under failure, then simulation.metrics to review the health impact. Do not use it to advance simulation time — that is simulation.step. Pass simulationId from simulation.create when this call is made from a fresh MCP session; otherwise you may omit it to target the current simulation in the preserved MCP session. Authenticate with an API key to unlock all 63 tools including typed durational failures and chaos engineering.Destructive
simulation.inject_trafficChange the traffic load on a demo simulation. Omit traffic to trigger a random 2×–5× spike (sends random: true internally); provide traffic to set an absolute RPS level (capped at 10000 RPS in demo mode). Use it to stress-test the architecture before stepping; the change only affects metrics after the next simulation.step. Do not use it to read metrics (simulation.metrics) or advance time (simulation.step). Pass the simulationId returned by simulation.create when your connector opens a fresh MCP session; preserve Mcp-Session-Id to use the omitted-ID current-simulation default. Returns the updated simulation with its new traffic level; the likely next tool is simulation.step.Changes data
simulation.metricsRead the latest metrics and resource states for a temporary anonymous demo simulation: latency, CPU, throughput, error rate, cost per hour, and per-resource health. Use it to inspect current state and metrics history without advancing time; do not use it to move the simulation forward — that is simulation.step. Responses are compact by default: principal current metrics plus explicit modeled goodputRps (a post-step point rate sourced from throughput, with provenance), goodputWindow (recorded only from persisted simulation-clock bounds, otherwise unavailable with provenance; never derive it from retrieval time or currentStep), errorBreakdown when available, per-resource status (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, and recoveryBlockedReason when provided), seeded EKS Spot checkpoint history and migrationEvaluationComplete/provenance when present, and the last 10 metrics-history entries. Pass responseMode: 'full' to get the complete simulation state, normalizedConfig, and full metrics history instead. During recovery, each resource may include recoveryProgress.state (parked, cooling_down, or healthy) with parkWindow and cooldown counters; poll simulation.step until healthy, then use simulation.metrics to inspect the resulting state and metrics. GPU / inference workflow: when the simulation includes a kubernetes resource with characteristics.inferenceMode: true, the response also includes top-level gpuUtilization (%), tokensPerSecond, costPerMillionTokens (USD/M tokens), idleGpuCostPerHour (USD/hr of standby GPU spend), and idleGpuFraction (0-1 idle HA overhead share) from the latest step, and each history entry carries the same inference fields. Pass the simulationId returned by simulation.create when your connector opens a fresh MCP session; preserve Mcp-Session-Id to use the omitted-ID current-simulation default. A fresh session has no current-simulation pointer. At least one simulation.step is needed for meaningful metrics. Read-only; repeated calls may be subject to demo usage limits. The likely next tool is simulation.step or simulation.inject_traffic.Read-only
simulation.recover_resourceRecover one reversible failed resource in a temporary anonymous demo simulation. No API key required for this temporary anonymous demo operation. Lower traffic to a serviceable level first, then provide resourceId or resourceName from simulation.create, simulation.step, or simulation.metrics. This deactivates applicable instance_down/database_overload failures for only the selected resource and returns recoveryProgress with parked, cooling_down, or healthy state plus cooldown counters. It cannot restore an instance_kill because that failure permanently removes the resource. The likely next tool is simulation.step; keep stepping and inspect the targeted resource until recoveryProgress.state is healthy. Pass simulationId from simulation.create when using a fresh MCP session; a preserved session may omit it. Authenticate with an API key to unlock all 63 tools and unlimited simulations.Changes data
simulation.stepAdvance a temporary anonymous demo simulation by one time step and return updated metrics — CPU, latency, throughput, error rate, cost (max 20 persisted steps per demo). Concurrent simulation.step calls on one simulation either serialize as distinct consecutive steps (in the browser Workspace queue) or receive HTTP 409 simulation_step_in_progress without advancing or consuming a demo credit. Wait for the running call to finish, then retry only rejected calls; distinct simulations can run concurrently. After a timeout, inspect simulation.metrics before retrying an uncertain result. Use it to observe how the architecture behaves over time, typically right after simulation.create or simulation.inject_traffic. Do not use it to read current state without advancing time — that is simulation.metrics. Pass the simulationId returned by simulation.create when your connector opens a fresh MCP session; preserve Mcp-Session-Id to use the omitted-ID current-simulation default. The likely next tool is simulation.step again (to keep observing) or simulation.inject_traffic (to change load first). Responses are compact by default: principal metrics plus per-resource status (id, name, status, cpuPercent, routedRps, availabilityState, isRoutable, and recoveryBlockedReason when provided) and this step's events. Seeded characteristics.eksSpotInterruption telemetry retains its additive migrationEvaluation beside the interruption lifecycle: use its recorded/derived/unavailable field provenance, frozen deadline verdict/counts/reasons, and simulation-clock milestones rather than final service health. The distinct eksSpotMigration contract remains separately reported when configured. Compact responses also include errorBreakdown when the engine provides it. A critical resource with isRoutable: true is degraded but still serving; availabilityState: unavailable and isRoutable: false identify a failed or parked node. Pass responseMode: 'full' to get the complete simulation state instead. During recovery, each resource may include recoveryProgress with state parked, cooling_down, or healthy, plus parkWindow and cooldown counters. Poll simulation.step until the targeted resource's recoveryProgress.state is healthy, then use simulation.metrics to inspect the resulting state and metrics. GPU / inference workflow: when the simulation includes a kubernetes resource with characteristics.inferenceMode: true, each step response also includes gpuUtilization (%), tokensPerSecond, costPerMillionTokens (USD/M tokens), idleGpuCostPerHour (USD/hr of standby GPU spend), and idleGpuFraction (0-1 idle HA overhead share) so you can track inference economics step by step. Authenticate with an API key for unlimited steps and GPU right-sizing hints.Changes data

Change history

  1. server instructions changed (+"build_revision=fe113a67534e build_timestamp=2026-10-04T12:32:19.609Z" -"build_revision=fda701abee9a build_timestamp=2026-10-03T13:39:06.562Z")
  2. simulation.recover_resource: description changed (+"63" -"62")
  3. simulation.inject_failure: description changed (+"63" -"62")
  4. simulation.delete: description changed (+"63" -"62")
  5. simulation.create: input schema changed
  6. server instructions changed (+"build_revision=fda701abee9a build_timestamp=2026-10-03T13:39:06.562Z" -"build_revision=e5525b3753cc build_timestamp=2026-10-02T02:24:25.292Z")
Source listings
SourceListingFirst seenLast seenVersions
Official MCP Registryai.cloudworldmodel/cloud-world-model2 Oct 20265 Oct 20261