Video Transcription Statistics (2026): What 5,832 Real Transcripts Reveal
Original data from 5,832 transcription requests across seven social platforms: the 45-second median Reel vs the 58-minute median YouTube video, 184-wpm Instagram speech, a 68% non-English audio share, and how often videos contain no speech at all.
Between May 2 and August 12, 2026, people used Voqusa to transcribe 5,832 videos and audio files across seven social platforms. This post publishes what that dataset shows. The headline numbers: the median Instagram Reel submitted for transcription runs 45 seconds, while the median YouTube video runs 58 minutes — a 77× gap. English-language Reels average 184 spoken words per minute, roughly 20% faster than conversational speech. 68% of the audio isn't in English. And 3.1% of the videos people try to transcribe turn out to contain no detectable speech at all.
Every figure below comes from our own production telemetry — aggregate statistics only, with no transcript content read or quoted — and we'll refresh the study as the dataset grows.
Key findings#
- The median Instagram Reel submitted for transcription is 45 seconds; the median YouTube video is 57 minutes 47 seconds. People transcribe clips on social platforms — and full podcasts, lectures, and interviews on YouTube.
- English-language speech in short-form video is fast: 184 words per minute on Instagram and 172 wpm on TikTok, against the commonly cited 120–150 wpm conversational baseline.
- 68.3% of transcribed audio is in a language other than English. Korean is the second-largest audio language (16.9%), ahead of Spanish (16.1%) and Hindi (9.2%).
- 71.5% of YouTube requests could be served from the video's existing captions. On every other platform, effectively 100% required speech-to-text — creator-provided captions barely exist outside YouTube.
- 3.1% of submitted videos contain no detectable speech. Pinterest leads at 6.9% — nearly 4× Instagram's 1.8% — reflecting music-only Idea Pins.
About the dataset#
The dataset covers every transcription processed by Voqusa, a free online transcript generator, between May 2 and August 12, 2026: 5,832 stored transcripts of 5,659 distinct videos and audio uploads across Instagram, YouTube, Pinterest, TikTok, Facebook, X/Twitter, LinkedIn, and direct file uploads. All statistics are aggregates computed over database metadata — platform, transcription mode, detected language, duration, and word counts — with no transcript content read or quoted. Speech-rate figures use English-language, speech-to-text transcripts between 10 seconds and 2 hours long, dividing word count by the audio duration actually transcribed. Language shares exclude 1,046 rows where the pipeline did not record a detected language. The set includes one flagged scripted burst of roughly 230 automated YouTube requests on July 24–25, 2026; excluding it moves YouTube's share from 25.2% to 22.1% and no other platform's share by more than 1.2 points.
Which platforms do people transcribe?#
| Platform | Transcripts | Share | Served from existing captions |
|---|---|---|---|
| 1,720 | 29.5% | 0% | |
| YouTube | 1,469 | 25.2% | 71.5% |
| 953 | 16.3% | 0.1% | |
| TikTok | 804 | 13.8% | 0% |
| File uploads | 424 | 7.3% | — |
| 275 | 4.7% | 0% | |
| X (Twitter) | 159 | 2.7% | 0% |
| 27 | 0.5% | 0% |
Instagram is the most-transcribed platform by a clear margin, and Pinterest — rarely mentioned in transcription conversations — quietly takes third place, ahead of TikTok.
The captions column is the structural story. YouTube is the only platform with a functioning caption ecosystem: nearly three-quarters of YouTube requests could be filled from a caption track that already existed. Everywhere else, the caption track effectively never exists, so every transcript had to be generated from the audio with speech-to-text. If you publish short-form video and care about accessibility or search, assume your platform provides nothing by default.
(One Douyin request reached the dataset and is excluded from per-platform tables.)
How long are the videos people transcribe?#
Source duration is recorded for 3,606 transcripts (62% of the dataset — earlier pipeline versions didn't store it).
| Platform | Median | 90th percentile | Mean | n |
|---|---|---|---|---|
| 40s | 1m 30s | 54s | 924 | |
| 45s | 1m 49s | 63s | 1,478 | |
| File uploads | 59s | 12m 22s | 4m 34s | 107 |
| TikTok | 1m 00s | 2m 55s | 1m 32s | 445 |
| 1m 04s | 8m 51s | 3m 38s | 151 | |
| X (Twitter) | 2m 08s | 25m 44s | 7m 19s | 112 |
| 2m 10s | 3m 36s | 2m 32s | 22 | |
| YouTube | 57m 47s | 2h 36m | 1h 09m | 367 |
Two things stand out. First, the TikToks people transcribe are meaningfully longer than the Reels (median 60s vs 45s) — TikTok's long-video push shows up even here. Second, YouTube is a different product category entirely: a 77× longer median than Instagram. Nobody is transcribing Shorts; they're transcribing podcasts, lectures, and interviews to take notes, quote, and repurpose. That's also why YouTube transcription behaves less like a captions tool and more like a research tool.
How fast do people actually talk?#
Measured on English-language, speech-to-text transcripts only (word counts by whitespace splitting don't transfer to languages written without spaces):
| Platform | English speech rate | Sample |
|---|---|---|
| 187 wpm | 43 | |
| 184 wpm | 312 | |
| X (Twitter) | 174 wpm | 65 |
| TikTok | 172 wpm | 121 |
| 153 wpm | 417 | |
| File uploads | 142 wpm | 35 |
Conversational English is commonly measured at 120–150 words per minute. Short-form platforms run 15–25% above that: at 184 wpm, an Instagram Reel delivers about three words every second, which is why the first two seconds of a script hold roughly six words of hook. Pinterest sits closer to a tutorial voiceover pace, and uploaded files — meetings, interviews, lectures — are the slowest speech in the dataset.
Which languages get transcribed?#
Share of the 4,786 transcripts with a recorded language:
| Language | Share |
|---|---|
| English | 31.7% |
| Korean | 16.9% |
| Spanish | 16.1% |
| Hindi | 9.2% |
| Russian | 7.9% |
| Chinese | 6.9% |
| Portuguese | 2.2% |
| French | 2.0% |
| Arabic | 1.1% |
| Polish | 1.1% |
More than two-thirds of the audio people transcribe is not in English. Korean's second place is striking — bigger than Spanish in this dataset — and Hindi, Russian, and Chinese together account for nearly a quarter. Transcription demand is global in a way English-first tooling conversations rarely acknowledge.
How many videos have no speech at all?#
Measured from product analytics over the same period (6,219 completed attempts, automated traffic filtered out): 3.1% of videos submitted for transcription contain no detectable speech.
| Platform | No-speech rate |
|---|---|
| 6.9% | |
| 3.9% | |
| File uploads | 3.1% |
| YouTube | 2.9% |
| TikTok | 1.9% |
| 1.8% | |
| X (Twitter) | 0.6% |
Pinterest is the silent-video champion at nearly 4× Instagram's rate — a large share of Idea Pins are music plus on-screen text, with nothing for speech-to-text to hear. If a transcript comes back empty on Pinterest, the odds are good the pin genuinely contains no speech.
What this means if you make content#
- Front-load harder than feels natural. At ~3 spoken words per second, a two-second hook is six words. Look at how top performers spend them — the TikTok SEO guide breaks down the query-matching version of this.
- Don't assume captions exist anywhere but YouTube. Outside YouTube, effectively every video needed speech-to-text. If you want your speech to be searchable, quotable, or accessible, you have to produce the text yourself — from a video link or an audio file.
- Long-form is a transcription use case, not just a captions one. The 58-minute YouTube median says viewers and creators use transcripts to mine long content — notes, quotes, blog drafts — not to subtitle it.
Cite this data#
Free to cite with attribution as "Voqusa Transcript Dataset, August 2026 edition", linked to this page. Figures are frozen as of August 12, 2026; we plan to publish refreshed editions as the dataset grows.
Frequently asked questions#
How fast do people talk in short-form videos?
In this dataset, English-language speech averages 184 words per minute on Instagram Reels, 172 wpm on TikTok, and 187 wpm on Facebook — all well above the 120–150 wpm conversational baseline. Practically, that is about three words per second, which is why dense, front-loaded hooks dominate short-form scripts.
What percentage of videos have no speech?
3.1% of videos submitted for transcription contain no detectable speech. Pinterest has the highest silent share at 6.9% — largely music-only Idea Pins — followed by Facebook at 3.9%. Instagram Reels (1.8%) and X videos (0.6%) are the most consistently speech-driven formats in the dataset.
Which platform do people transcribe the most?
Instagram leads with 29.5% of requests, followed by YouTube (25.2%), Pinterest (16.3%), and TikTok (13.8%). YouTube is the outlier in kind rather than volume: it is the only platform where most requests (71.5%) could be served from existing captions, and its median video runs 58 minutes versus under a minute everywhere else.
Methodology notes#
- Counts are stored transcripts (5,832) of distinct videos and uploads (5,659); repeat requests for an already-transcribed video are served from cache and don't add rows.
- Language shares merge the labels "en" and "english" and exclude 1,046 rows with no recorded language. Speech-rate figures use English speech-to-text transcripts of 10 seconds to 2 hours, with word count divided by the duration actually transcribed; platforms with fewer than 30 qualifying samples are omitted.
- Duration statistics cover the 62% of rows where source duration was recorded. No-speech rates come from product analytics events over the ~100 days ending August 12 (a slightly different denominator than the transcript table), with known automated traffic excluded.
- One scripted burst of ~230 automated YouTube requests on July 24–25 is included in totals and quantified in "About the dataset"; one Douyin request is excluded from per-platform tables.
- All statistics are aggregates over metadata. No transcript content was read, quoted, or sampled for this study.

Building Voqusa to make video transcription free, fast, and accurate for creators in every language.

