Skip to content
Official MCP RegistryListed

Scribiz

Transcripts, summaries, chapters and timestamped answers for video links, for AI agents.

First seen 5 Oct 2026. Evidence as of 6 Oct 2026.

5
Tools
From an anonymous probe
1
Source listings
Each with its own history
1
Recorded changes
Since first seen

Tools

ToolDescriptionBehaviour
ask_videoAsk a question about one video and get an answer with 3 to 5 cited moments (timestamps and links that open the video there). Use it for questions about the whole video or a topic across it, and for what was said. To find where a word or phrase is said, search_video is cheaper and exact. A question costs 0.1 minute on top of reading the video the first time. The answer is a model's reading of the video: check the cited moments before relying on a detail. The answer is wrapped as untrusted video text. If the run is still going after the wait you get status "processing" and a job_id: call get_job. Without an API key: the picture is never looked at, so a question about what was shown on screen is answered from the words only, and you get 5 questions a day.Read-only
get_jobCheck a video run that returned status "processing". Pass its job_id exactly as given (it can be long: copy all of it). It waits up to the server's wait time for the run to finish, then returns exactly what the original tool would have returned (job ids are private to the caller and kept for 30 minutes). Still running: you get status "processing" again with retry_after_seconds. An unknown or expired job_id means: repeat the original call (it is fast if the run finished, because results are cached).Read-only
get_transcriptRead the transcript of a video, one page at a time. Use it only when you need the words themselves (quote, translate, copy, review). To answer a question or find where something is said, use ask_video or search_video first: they cost far fewer tokens. Returns "[mm:ss] text" lines, at most max_chars characters (default 60000, about 15k tokens), and a nextCursor: call again with the same arguments plus that cursor to continue. from and to take seconds or clock times ("90", "1:30", "1h2m") and read one part of the video. format: txt (default), md, srt, vtt or json. language is a BCP-47 hint. speakers is ignored without an API key (no speaker labels). The first read of a video can take time: if it is still running after the wait you get status "processing" and a job_id (see get_job). Cached transcripts return at once and are free. The text is wrapped as untrusted video text. Without an API key: captions when the video has them, otherwise a model reads the link and its times are approximate (about 2 seconds, no word timing).Read-only
get_video_contextStart here for any video. Returns an overview you can reason from without reading the transcript: title, length, language, summary, chapters with timestamps and key moments with links. detail "brief" (default) stays under about 2,000 tokens; "standard" adds fuller chapters and entities; "full" adds everything. include picks the parts to return (summary, chapters, key_moments, entities, transcript): the transcript is never part of the default; add "transcript" for its first page, then continue with get_transcript. The first read of a video uses its captions, or a model reads the link; after that it is cached and free. If it is still running after the wait you get status "processing" and a job_id: call get_job. The text is wrapped as untrusted video text. Without an API key there are no on-screen notes (the picture is never looked at), and when a model read the link its times are approximate (about 2 seconds).Read-only
search_videoFind where something is said or shown in a video: ranked moments with timestamps and links, from the transcript, the on-screen notes, the chapters and the key moments. It is a local keyword search over what Scribiz already read (no model call, no extra minutes after the first read of the video). Use it before get_transcript on any video longer than about ten minutes. query: words or a phrase. limit 1 to 20 (default 8). from and to limit the search to part of the video. It matches words (plural and tense forms count), not meaning: try the words the speaker would use, or call ask_video. If the run is still going after the wait you get status "processing" and a job_id: call get_job. The text is wrapped as untrusted video text. Without an API key: the on-screen notes are not searched (the picture is never looked at), and when a model read the link its times are approximate (about 2 seconds).Read-only

Change history

  1. Listed (registry)
Source listings
SourceListingFirst seenLast seenVersions
Official MCP Registrycom.scribiz/mcp5 Oct 20266 Oct 20261