We have not measured AIOS Scribe's transcription accuracy, so this page publishes no accuracy percentage. It does two honest things instead: it shows what we can state from real usage (how often jobs finish, and how long they take), and it sets out the method we will use to measure accuracy, so that when numbers appear here you can judge how they were produced.
What we can report: how jobs actually went
From the application's own records, read-only, for the period 14 September to 7 October 2026 (the time since transcription moved into the app):
| Outcome | Jobs |
|---|---|
| Completed | 225 |
| Failed: reply too long for one request (truncated) | 15 |
| Failed: took too long | 11 |
| Failed: busy | 1 |
| Failed: other, or before reasons were recorded | 52 |
| Total | 304 |
Videos over two hours are refused before a job is created, so they do not appear in this table. That is 74 per cent completed. Read it with care: it includes our own testing and the weeks before several fixes, and a finished job says nothing about whether the words are right. It is a record of completion, not of accuracy.
Timing is better recorded only from the last week. For completed jobs with a recorded start: videos under five minutes took a median of 8 seconds (9 jobs, longest 33 seconds), and videos between 30 minutes and two hours took a median of 228 seconds (6 jobs, longest 382 seconds). These are small samples.
The method we will use for accuracy
- A fixed sample. A published list of public videos chosen in advance, not after seeing the results, covering: a single clear speaker, an interview with two speakers, a podcast with overlap, a lecture with technical terms, a video with background noise or music, and at least two languages other than English. Each entry records the length, the speakers, the accent where known, and the audio conditions.
- A reference transcript. A person corrects a transcript while listening, and a second person checks it. Without a trustworthy reference no accuracy figure means anything.
- A standard measure. Word error rate: substitutions, deletions and insertions divided by the number of words in the reference, after the same normalisation of case and punctuation on both sides. We will also report separately on names, numbers and technical terms, because a low overall rate can hide errors in exactly the words that matter.
- Versioned and dated. Every result will carry the date, the AI model and settings used, and the sample size. When the model changes, the test is run again and the old result is kept, not overwritten.
- Failures included. Videos that were declined or truncated count as failures in the results, not as removed samples.
What we cannot claim yet
- That AIOS Scribe is more accurate than YouTube's automatic captions or than human transcription.
- Any accuracy for a particular language, accent, or kind of audio.
- That timestamps are exact. They are estimates; see how they are corrected.
Until then: when to verify manually
Verify before you quote, cite, publish or submit anything. The limitations guide lists what we have seen go wrong and a short checklist.
Frequently asked questions
- How accurate is AIOS Scribe?
- We have not measured it, so we do not give a figure. Treat the transcript as a draft and verify names, numbers and quotes against the video.
- Why publish a page about accuracy with no accuracy number?
- Because an invented or unmeasured number would mislead. The page shows what is known (job outcomes and timing) and the method for measuring the rest.