This page describes what happens between pasting a link and reading a transcript, as the product works today. It leaves out keys, prompts and security details, and it says where a step is a workaround rather than a feature.
1. The link is checked
The link has to be a standard YouTube video address: a watch?v= link, a youtu.be short link, an embed link or a Shorts link. Anything else is rejected before any work starts. The app then asks the YouTube Data API for the video's real length and its title, because the length that the transcription model reports for itself has been unreliable (it once reported 97 seconds for a 16-minute video).
2. We look for work already done
If a finished transcript for that video exists, you get it straight away. If someone else's job for the same video is still running, you join it instead of starting another. Neither costs anything or counts against your limits. Only a genuinely new video starts a new job, and then the limits apply: one new job per network address every 15 seconds (people sharing one connection share that limit), one per video at a time, a maximum video length of two hours, and a monthly spending cap on the AI service. A video that YouTube reports as over two hours is refused up front with a clear message. If the length cannot be looked up, that check is skipped and the video is attempted as a single request, which can fail for a very long video.
3. The AI model reads the video from YouTube
The job goes into a queue and the transcription model is given the YouTube address. AIOS Scribe does not download the video. The model returns a list of segments, each with a start time, an end time and the words spoken.
Why long videos are sent in pieces
A video of up to 90 seconds is sent as one request. Anything longer is cut into sections of at most 60 seconds, and up to four are processed at once. Two things drove this:
- A hard wait limit. A single request that takes more than about two minutes to answer is cut off by our hosting provider. A 49-minute video failed this way on every attempt until we split it. With our first splitting setup (seven sections of about eight minutes) it finished in 83 seconds. The current setup uses far more, shorter sections, so that figure is from the older design and we have not re-measured a video that long.
- Completeness. On a 3 minute 25 second song, asking for the whole video at once left stretches of 40 to 50 seconds with about ten words and cut choruses off mid-line. The same stretches asked for on their own came back with roughly twice the words.
A hard cut every minute would hand each request half a sentence wherever speech crosses a cut, so each section is requested with 8 extra seconds on each side. Each section keeps only the lines that start in its own minute (give or take 1.5 seconds), so a sentence belongs to the section that heard its start whole. Where two sections both heard the same line at a cut, the more complete copy is kept, but only when the two copies overlap in time, start within 2.5 seconds of each other, and the wording matches in the same order. Short repeated lines, such as a chant, are deliberately never merged, so a line heard twice at a cut can appear twice. We chose that over silently deleting something that was really said.
4. Timestamps are corrected, within limits
The model times a transcript on its own clock, and that clock can be wrong in either direction. On two talks it ran about 1.6 times fast (a 741-second talk was timed to 1,215 seconds); on an 844-second talk it ended at 715 seconds, which is short. When timestamps run past the video's real end, they are compressed to fit, and a clip's times are shifted to its place in the video. We do not stretch timestamps that run short, because a short span could equally be silence at the end, so a clock that runs slow is left as the model gave it. In two songs the model switched from seconds to a minutes.seconds notation partway through, so 1:30 arrived as 1.3; the app detects that and converts it back. Corrections like these are estimates. A line can still be off by a few seconds, so the timestamp is a way to go and listen, not a measurement.
5. The transcript is saved and shown
For every submission the app stores the video link and ID and a thumbnail address, plus your account ID if you are signed in; a finished transcript adds the title, length and timestamped segments. The video itself is never stored. Later visitors who submit the same public video are served the saved copy, and a signed-in visitor also gets it in their history. If a job fails, the reason is recorded (for example too long, declined by the model, or timed out) and shown to you in plain words. Only “too long” is treated as permanent; a refusal can succeed on the next try, so it is not remembered as final.
The title of a video sent in pieces is taken from the one YouTube shows, because no single piece saw the whole video and the first minute alone can be an intro or an advertisement.
6. What you can do with it
Read it, search it, click a timestamp to open the video at that moment, and copy it or save it as .txt or .md. Separately, you can ask for a summary and key points. That is a second, text-only AI call over the stored transcript. It is generated only when you ask, stored once per video, and labelled as AI-generated.
What this does not do
It does not label speakers, translate, or export subtitle files. It cannot read private or unlisted videos. It does not promise an accuracy figure; see where transcripts go wrong and how we plan to measure it.
Frequently asked questions
- Does AIOS Scribe download the video?
- No. The AI model is given the YouTube address and reads the video from YouTube; the video itself is never copied or stored. What is stored is the link and video ID, the thumbnail address, the title and length, and the transcript with its timestamps, plus your account ID if you are signed in. See the privacy policy.
- Why can a long video take several minutes?
- Videos over 90 seconds are processed in sections of up to 60 seconds, four at a time, and then joined. Short videos finish in seconds and 30-minute to 2-hour videos took a median of about four minutes in our records (a small sample of six jobs). Busy periods take longer.
- Why might a line appear twice at a cut?
- Where two sections both heard the same short line, AIOS Scribe keeps both rather than risk deleting a line that was genuinely said twice. Longer duplicates that are the same audio are merged.