·8 min read·ANALYSIS

Video Transcription Statistics (2026): What 5,832 Real Transcripts Reveal

Original data from 5,832 transcription requests across seven social platforms: the 45-second median Reel vs the 58-minute median YouTube video, 184-wpm Instagram speech, a 68% non-English audio share, and how often videos contain no speech at all.

Michael LiuMichael Liu·
video transcription statisticsspeech ratewords per minuteshort-form video datareels lengthtranscription languages

Between May 2 and August 12, 2026, people used Voqusa to transcribe 5,832 videos and audio files across seven social platforms. This post publishes what that dataset shows. The headline numbers: the median Instagram Reel submitted for transcription runs 45 seconds, while the median YouTube video runs 58 minutes — a 77× gap. English-language Reels average 184 spoken words per minute, roughly 20% faster than conversational speech. 68% of the audio isn't in English. And 3.1% of the videos people try to transcribe turn out to contain no detectable speech at all.

Every figure below comes from our own production telemetry — aggregate statistics only, with no transcript content read or quoted — and we'll refresh the study as the dataset grows.

Key findings#

  • The median Instagram Reel submitted for transcription is 45 seconds; the median YouTube video is 57 minutes 47 seconds. People transcribe clips on social platforms — and full podcasts, lectures, and interviews on YouTube.
  • English-language speech in short-form video is fast: 184 words per minute on Instagram and 172 wpm on TikTok, against the commonly cited 120–150 wpm conversational baseline.
  • 68.3% of transcribed audio is in a language other than English. Korean is the second-largest audio language (16.9%), ahead of Spanish (16.1%) and Hindi (9.2%).
  • 71.5% of YouTube requests could be served from the video's existing captions. On every other platform, effectively 100% required speech-to-text — creator-provided captions barely exist outside YouTube.
  • 3.1% of submitted videos contain no detectable speech. Pinterest leads at 6.9% — nearly 4× Instagram's 1.8% — reflecting music-only Idea Pins.

About the dataset#

The dataset covers every transcription processed by Voqusa, a free online transcript generator, between May 2 and August 12, 2026: 5,832 stored transcripts of 5,659 distinct videos and audio uploads across Instagram, YouTube, Pinterest, TikTok, Facebook, X/Twitter, LinkedIn, and direct file uploads. All statistics are aggregates computed over database metadata — platform, transcription mode, detected language, duration, and word counts — with no transcript content read or quoted. Speech-rate figures use English-language, speech-to-text transcripts between 10 seconds and 2 hours long, dividing word count by the audio duration actually transcribed. Language shares exclude 1,046 rows where the pipeline did not record a detected language. The set includes one flagged scripted burst of roughly 230 automated YouTube requests on July 24–25, 2026; excluding it moves YouTube's share from 25.2% to 22.1% and no other platform's share by more than 1.2 points.

Which platforms do people transcribe?#

PlatformTranscriptsShareServed from existing captions
Instagram1,72029.5%0%
YouTube1,46925.2%71.5%
Pinterest95316.3%0.1%
TikTok80413.8%0%
File uploads4247.3%
Facebook2754.7%0%
X (Twitter)1592.7%0%
LinkedIn270.5%0%

Instagram is the most-transcribed platform by a clear margin, and Pinterest — rarely mentioned in transcription conversations — quietly takes third place, ahead of TikTok.

The captions column is the structural story. YouTube is the only platform with a functioning caption ecosystem: nearly three-quarters of YouTube requests could be filled from a caption track that already existed. Everywhere else, the caption track effectively never exists, so every transcript had to be generated from the audio with speech-to-text. If you publish short-form video and care about accessibility or search, assume your platform provides nothing by default.

(One Douyin request reached the dataset and is excluded from per-platform tables.)

How long are the videos people transcribe?#

Source duration is recorded for 3,606 transcripts (62% of the dataset — earlier pipeline versions didn't store it).

PlatformMedian90th percentileMeann
Pinterest40s1m 30s54s924
Instagram45s1m 49s63s1,478
File uploads59s12m 22s4m 34s107
TikTok1m 00s2m 55s1m 32s445
Facebook1m 04s8m 51s3m 38s151
X (Twitter)2m 08s25m 44s7m 19s112
LinkedIn2m 10s3m 36s2m 32s22
YouTube57m 47s2h 36m1h 09m367

Two things stand out. First, the TikToks people transcribe are meaningfully longer than the Reels (median 60s vs 45s) — TikTok's long-video push shows up even here. Second, YouTube is a different product category entirely: a 77× longer median than Instagram. Nobody is transcribing Shorts; they're transcribing podcasts, lectures, and interviews to take notes, quote, and repurpose. That's also why YouTube transcription behaves less like a captions tool and more like a research tool.

How fast do people actually talk?#

Measured on English-language, speech-to-text transcripts only (word counts by whitespace splitting don't transfer to languages written without spaces):

PlatformEnglish speech rateSample
Facebook187 wpm43
Instagram184 wpm312
X (Twitter)174 wpm65
TikTok172 wpm121
Pinterest153 wpm417
File uploads142 wpm35

Conversational English is commonly measured at 120–150 words per minute. Short-form platforms run 15–25% above that: at 184 wpm, an Instagram Reel delivers about three words every second, which is why the first two seconds of a script hold roughly six words of hook. Pinterest sits closer to a tutorial voiceover pace, and uploaded files — meetings, interviews, lectures — are the slowest speech in the dataset.

Which languages get transcribed?#

Share of the 4,786 transcripts with a recorded language:

LanguageShare
English31.7%
Korean16.9%
Spanish16.1%
Hindi9.2%
Russian7.9%
Chinese6.9%
Portuguese2.2%
French2.0%
Arabic1.1%
Polish1.1%

More than two-thirds of the audio people transcribe is not in English. Korean's second place is striking — bigger than Spanish in this dataset — and Hindi, Russian, and Chinese together account for nearly a quarter. Transcription demand is global in a way English-first tooling conversations rarely acknowledge.

How many videos have no speech at all?#

Measured from product analytics over the same period (6,219 completed attempts, automated traffic filtered out): 3.1% of videos submitted for transcription contain no detectable speech.

PlatformNo-speech rate
Pinterest6.9%
Facebook3.9%
File uploads3.1%
YouTube2.9%
TikTok1.9%
Instagram1.8%
X (Twitter)0.6%

Pinterest is the silent-video champion at nearly 4× Instagram's rate — a large share of Idea Pins are music plus on-screen text, with nothing for speech-to-text to hear. If a transcript comes back empty on Pinterest, the odds are good the pin genuinely contains no speech.

What this means if you make content#

  • Front-load harder than feels natural. At ~3 spoken words per second, a two-second hook is six words. Look at how top performers spend them — the TikTok SEO guide breaks down the query-matching version of this.
  • Don't assume captions exist anywhere but YouTube. Outside YouTube, effectively every video needed speech-to-text. If you want your speech to be searchable, quotable, or accessible, you have to produce the text yourself — from a video link or an audio file.
  • Long-form is a transcription use case, not just a captions one. The 58-minute YouTube median says viewers and creators use transcripts to mine long content — notes, quotes, blog drafts — not to subtitle it.

Cite this data#

Free to cite with attribution as "Voqusa Transcript Dataset, August 2026 edition", linked to this page. Figures are frozen as of August 12, 2026; we plan to publish refreshed editions as the dataset grows.

Frequently asked questions#

How fast do people talk in short-form videos?

In this dataset, English-language speech averages 184 words per minute on Instagram Reels, 172 wpm on TikTok, and 187 wpm on Facebook — all well above the 120–150 wpm conversational baseline. Practically, that is about three words per second, which is why dense, front-loaded hooks dominate short-form scripts.

What percentage of videos have no speech?

3.1% of videos submitted for transcription contain no detectable speech. Pinterest has the highest silent share at 6.9% — largely music-only Idea Pins — followed by Facebook at 3.9%. Instagram Reels (1.8%) and X videos (0.6%) are the most consistently speech-driven formats in the dataset.

Which platform do people transcribe the most?

Instagram leads with 29.5% of requests, followed by YouTube (25.2%), Pinterest (16.3%), and TikTok (13.8%). YouTube is the outlier in kind rather than volume: it is the only platform where most requests (71.5%) could be served from existing captions, and its median video runs 58 minutes versus under a minute everywhere else.

Methodology notes#

  • Counts are stored transcripts (5,832) of distinct videos and uploads (5,659); repeat requests for an already-transcribed video are served from cache and don't add rows.
  • Language shares merge the labels "en" and "english" and exclude 1,046 rows with no recorded language. Speech-rate figures use English speech-to-text transcripts of 10 seconds to 2 hours, with word count divided by the duration actually transcribed; platforms with fewer than 30 qualifying samples are omitted.
  • Duration statistics cover the 62% of rows where source duration was recorded. No-speech rates come from product analytics events over the ~100 days ending August 12 (a slightly different denominator than the transcript table), with known automated traffic excluded.
  • One scripted burst of ~230 automated YouTube requests on July 24–25 is included in totals and quantified in "About the dataset"; one Douyin request is excluded from per-platform tables.
  • All statistics are aggregates over metadata. No transcript content was read, quoted, or sampled for this study.