Skip to content
Official MCP RegistryListed

Misata Studio: verified synthetic data

Part ofMisata Studio: verified synthetic dataMCP server

Hosted: plan and build multi-table test and demo data, export 25 formats, with checks shown.

First seen 2 Oct 2026. Evidence as of 8 Oct 2026.

12
Tools
From an anonymous probe
1
Source listings
Each with its own history
0
Recorded changes
Since first seen

Tools

ToolDescriptionBehaviour
blueprint_guide The reference for the engine's full design language (the `blueprint` argument of generate_dataset, start_generation and validate_blueprint): roles, distributions, formulas that read parent columns and draw noise, per-group sequences (seq, lag, cumsum, ar1 drift, random walk), aggregates, event windows, cause-and-effect, lifecycles, with patterns. Read it before writing a blueprint. Read-only
cancel_generationStop a running start_generation job. It stops at the next stage boundary and keeps nothing.Changes data
export_dataset Export a generated dataset as a file. Returns a `download_url` the person can open (or you can fetch, e.g. with curl) for as long as the dataset is held, about 2 hours. Give the person the link rather than pasting file contents into the chat. Args: dataset_id: from a prior generate_dataset call. format: data: csv, parquet, jsonl, json, avro, xlsx, feather, orc, sqlite, duckdb, sql. code and docs: dbt, notebook, dictionary, dbml, mermaid, prisma, sqlalchemy, typescript, jsonschema, expectations, django, openapi, mockapi, demo. `sql` is schema.sql (DDL with keys) + data.sql (COPY/INSERT) — the way to seed a real database: run the returned SQL through your own database connection, since this server never holds a database credential itself. dialect: for `sql` only: postgres, mysql, sqlite, mssql, oracle, bigquery, snowflake. inline: also return the file itself as `base64` (only for files under a few MB). Use it when you must write the file yourself and cannot fetch a URL. Returns: filename, content_type, bytes, download_url, expires_at (unix seconds), and `base64` when `inline` and small enough. Changes data
find_ready_dataset The ready-made datasets Misata publishes: free sample databases (direct download, public domain) and premium datasets with a full answer key (a free preview, then a one-off price). Check this first when someone wants sample, demo, practice or teaching data for a common scenario (retail, e-commerce, SaaS, manufacturing SPC, predictive maintenance, insurance claims, fraud/AML, clinical, network security): handing over a dataset that already exists is instant. If none fits their tables, generate one instead. Args: query: what the person needs, in their words. Only orders the list (closest first); every dataset is still returned, so judge the fit yourself from the tables and summary. Returns: datasets: [{slug, kind (free|premium), title, summary, rows, tables, page_url, and download_url (free) or free_preview_url + buy_url + price_usd (premium)}]. Read-only
generate_dataset Make a verified relational dataset and wait for it: every foreign key checked, dates correctly ordered, declared aggregates and rates exact, before anything is returned. Deterministic for a seed. Use this when you give a `schema` or `ddl` (seconds). For a plain-English `request` the engine designs the tables with a model, which can take several minutes: call start_generation instead and poll get_status, so the call does not sit open and time out. Args: request: Plain-English description (needs an LLM key unless `schema`/`ddl` is also given). schema: A Misata schema dict for exact structural control. No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. seed: Reproducibility seed — the same request/schema and seed always produce the same rows. research: Ground realistic numbers (prices, growth rates) in real published facts via a web search. Off by default (an anonymous caller's request should not trigger external calls unless asked for); needs a key regardless of `schema`/`ddl`. blueprint: The engine's full design language (blueprint_guide has the reference): readings around a parent's nominal, drift, autocorrelation, cause-and-effect, event windows, exact aggregates. Use it whenever the data must behave like the real process. No key needed; run validate_blueprint on it first. Returns: dataset_id: pass this to get_certificate / query_dataset / export_dataset. passed: whether every check held. A dataset that did not pass is still returned, with `certificate.findings` saying what failed — inspect before trusting it. verification: a short, plain-language account of what was checked and what held — show it to the person as the proof, in place of asking them to take the data on trust. tables: {name: {columns, preview_rows (first 20), total_rows}} — the preview only; every row is in the stored dataset, reachable by query_dataset/export_dataset. certificate: the short form (claims, requirements, findings). get_certificate returns the rest (patterns, realism scorecard, every proof chart). Changes data
get_certificate The full certificate for a dataset made by generate_dataset: every claim (stated vs actual), every requirement's status and evidence, realism findings, planted defects/anomalies if any, and the pattern charts. This is the whole answer key, not the trimmed form generate_dataset returns. Read-only
get_status Where a start_generation job is. While running: the stage it is in and how long it has run. When done: the same answer generate_dataset returns (dataset_id, verification, tables, certificate). Read-only
plan_dataset See the tables, sizes and relationships the engine would build, before any rows exist. Free (no rows are made), so it is worth calling before generate_dataset on anything non-trivial: review what it understood and assumed, then adjust your request or schema before spending a real call. Args: request: Plain-English description (needs an LLM key, unless `schema`/`ddl` is also given — then it is still used to ground realism, e.g. locale and what columns mean). schema: A Misata schema dict (see the server instructions for the format). No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. Returns: route: "chat" (not a dataset request — see `reply`), "design" (the model designed the tables) or "pack" (matched a built-in shape). tables: name, estimated rows, columns, foreign keys for each table the engine would build. understanding: what the engine read the request as (business, archetype, assumptions). requirements: every specific thing the request asked for, so you can see what was understood. Read-only
query_dataset Read-only SQL over a dataset's own tables (one SELECT/WITH statement; every table name is a view over that dataset's own files — nothing else on the server is reachable this way). Args: dataset_id: from a prior generate_dataset call. sql: a single SELECT (or WITH ... SELECT) statement. No semicolons, file paths, or statements that write or read outside the dataset (checked before running). limit: rows returned (capped at 5000). Returns: columns, rows, truncated (whether more rows existed than `limit`). Read-only
start_generation Start a generation in the background and return a job_id at once. Use it for any plain-English `request` (a model designs the tables, which takes minutes) — then call get_status(job_id) every 20-30 seconds until it says done. Arguments are the same as generate_dataset. Returns: job_id, status "running". get_status gives the stage and, when finished, the dataset. Changes data
validate_blueprint Check a blueprint before generating it: every error said as what to change, the design traps it falls into (a column that will round away, a copy of a value not made yet, a window that counts nothing), its size against your row cap, and a small preview run (a few thousand rows) with the verifier's findings and sample rows, so you can see the data behave before the real call. Free. Returns: valid, errors (fix these), warnings (read these), estimated_rows, row_cap, preview: {passed, findings, tables: {name: first rows}} when `preview`. Read-only
whoamiWhich account this connection is, and what it may do: signed in with a Studio MCP key or anonymous, the row cap, and whether a model key is available for plain-English requests.Read-only

Change history

No changes since the first observation. The first snapshot is the baseline.

Source listings
SourceListingFirst seenLast seenVersions
Official MCP Registrystudio.misata/misata2 Oct 20268 Oct 20261
Misata Studio: verified synthetic data on Official MCP Registry | InvokeRank