Skip to content
AIOS Scribe logoAIOS Scribe
Guide

How We Evaluate AI Transcription Quality (With the Tests We Ran)

What we check on every transcript without a reference, the seven test videos we ran on 7 October 2026 with their results, and what those results cannot tell you.

6 min readChecked against the product on

This page explains how we judge whether an AIOS Scribe transcript can be trusted, what we ran on 7 October 2026, and what those runs can and cannot tell you. We have not measured an accuracy percentage, and nothing here should be read as one.

1. Why quality is hard to measure

A word-level accuracy score needs a reference: a transcript that a person has corrected while listening, ideally checked by a second person. Building that for a varied set of videos takes hours per hour of audio, and a score from one kind of video says little about another. Without a reference, we can only check things about the transcript that are visible without listening.

2. What affects quality

Audio clarity, background music or noise, overlapping speakers, accents, how fast people talk, technical terms and names, and the language. The limitations guide lists the failures we have actually seen.

3. What we test

Seven checks that need no reference, run on every video: did it finish, how long did it take, how many words per minute came back, does the last timestamp land near the video's real length, is the longest stretch with no text short, do any segments run out of order, and does the same line repeat back to back. These catch the failures that matter most for trust (missing stretches, broken timing, loops), not wrong words.

4. Method

We ran each video through the same transcription code the site uses (model gemini-3.5-flash, short sections for long videos, low media resolution), from a developer machine, on 2026-10-07, once per video with up to three attempts if a request failed. It was not run through the production queue, so times are indicative. The videos are public and were chosen to cover different conditions; we did not drop any result.

5. Test cases and results

ConditionVideoLengthFinishedTimeWords per minuteLast timestamp / real lengthLongest stretch with no textOut of orderRepeated back to back
One clear speaker, EnglishObservational Learning and the Bobo Doll Experiment12:32Yes52 s2151.002 s01
Interview, two voices, EnglishElon Musk interview on Wired Science20:06Yes80 s2041.0034 s03
Music, non-EnglishPSY, Gangnam Style (music video)4:13Yes27 s600.9825 s01
Non-English talk (Spanish)TEDx talk on body language18:32Yes58 s1341.0012 s00
Technical terms, EnglishTotal Pixel Space9:28Yes40 s1040.9812 s01
Room audio with interpretation (Korean and English)Korean film press conference12:19Yes51 s840.997 s01
Very shortErase Bullying PSA0:31Yes6 s520.978 s00

One more case comes from production rather than this run: a 49-minute Vietnamese video that failed on every attempt as one request finished in 83 seconds once it was sent in sections (651 segments, timestamps in order, ending at 2,959 s of 2,974 s). That used our first splitting setup, which has since changed.

6. What went right

All seven finished. Timestamps never ran backwards, and in every case the last timestamp landed within 4 per cent of the real length. The Spanish and Korean videos finished without us choosing a language (we have not checked the wording of those transcripts). Repeated lines were rare: zero to three per video.

7. What looks wrong, or that we cannot judge

  • Long stretches with no text. 34 seconds in the interview and 25 seconds in the music video. We did not listen to confirm whether those were silence, music or missed speech, so we cannot say they are errors.
  • Words per minute differs a lot (52 to 215). Some of that is the material, and 215 for a lecture is faster than we would expect; we did not investigate whether it is a fast speaker or extra repetition.
  • Nothing here measures whether the words are right. A transcript can pass every check above and still misspell a name.

8. What AIOS Scribe does well, on this evidence

It finished every video, kept timestamps in order, and its timestamps spanned close to the whole length of each video. That shows the timing is plausible, not that every stretch of speech is present: we did not compare the transcripts with the audio, and two videos had long stretches with no text. It also finished on videos in other languages (how correct the words are, we have not checked). For a first draft you can search and navigate, that is useful.

9. What it still struggles with

Music is the clearest weak spot we have seen: some songs are declined, and one we tested before splitting returned 40 to 50 second stretches holding about ten words (fixed on 2026-10-07 by sending short sections). Names, numbers, overlapping speakers and strong accents remain untested by us; see the limitations guide.

10. How to verify an important quote

  1. Search the transcript for the line and note its timestamp.
  2. Open the video at that moment from the timestamp link and listen to a few seconds either side.
  3. Check names, numbers and technical terms word by word.
  4. Quote from the video, not from the transcript, if the wording matters.

11. What we plan next

The next real step is the one in the accuracy plan: a fixed sample, a person-corrected reference and a word error rate, reported with its date and model. Until that exists, this page is evidence about behaviour, not accuracy. Not yet tested: strong accents in a controlled way, overlapping speakers, code read aloud, and a second run of each video to see how much results vary.

Frequently asked questions

Does this page give an accuracy percentage?
No. The checks measure whether a transcript is complete, in order and close to the video length, not whether each word is right. A word error rate needs a person-corrected reference, which we have not built yet.
Can I repeat these tests?
Yes. The videos are public and linked in the table. Paste each into AIOS Scribe and compare the segment count, timestamps and longest gap. Results can vary between runs.

Try it on a video of your own

Paste a public YouTube link to get a transcript you can search and download. No sign-up needed.

Open the transcript tool