File Transcription

Transcribe audio and video files

Speakly transcribes audio and video files you already have: recorded meetings, interviews, voice memos, lecture videos.

Transcribe Media
The Transcribe Media page, opened from the sidebar

How to use

  1. Open Transcribe Media in the left sidebar (that is the tab's real name)
  2. Click New Transcription — the Transcribe Media File dialog opens
  3. Drop the file on the dashed area ("Drop audio or video file here") or click Choose File
  4. Set Media Language — "Select the language spoken in the media for better accuracy"
  5. Click Start Transcription

Supported formats

  • Audio: .mp3, .wav, .m4a, .ogg, .webm, .flac, .aac, .wma
  • Video: .mp4, .mov, .avi, .mkv

.webm counts as audio here, not video. Other extensions are ignored when dropped and rejected with an Unsupported media format error if they reach transcription.

Limits

  • Maximum duration: 2 hours. Longer media is refused with Audio too long (… min). Maximum is 120 minutes.
  • Anything longer than 8 minutes is split automatically into 9-minute chunks with a short overlap, transcribed chunk by chunk, then merged into one transcript
  • The dialog states the same limit: "Audio and video files up to 2 hours. Long files are automatically split into chunks."
  • Chunking no longer costs you the timestamps: each chunk's segments are shifted onto the full timeline, so the merged transcript has real times across the whole file instead of a single 00:00

Progress while it runs

A file transcription no longer shows only a spinner. The dialog names the phase it is in and draws a bar with a percentage, so a long video visibly moves instead of leaving you wondering whether it is stuck.

  • Preparing the audio… — Speakly reads the file and, when it is long enough to be split, re-encodes it into parts. On a long video this whole stretch happens before a single word is produced.
  • Transcribing media… — the parts are transcribed one after another. A file long enough to be split also shows Part x of y under the phase.
  • Putting it together… — the parts are merged into one transcript.

With the Local Model (On-Device) engine the percentage also moves *inside* each part, because that engine reports its own decoding progress as it goes. A cloud engine answers a whole part in one request and has nothing to report in between, so with a BYOK provider the bar advances part by part.

Where the file is processed

Local only on the local engine
Files are processed on your machine only when the engine is Local Model (On-Device). With any Your API Key (BYOK) provider — OpenAI, Groq, ElevenLabs, Google, Deepgram, Mistral or a custom endpoint — the audio is uploaded to that provider. Check or change the engine in Settings → Transcription.

A fresh install starts on the local model; Windows onboarding preselects a cloud provider instead. Either way, the finished transcription is stored only on your device.

Timestamps

The transcript is shown as timestamped lines, grouped so a line ends at the end of a sentence rather than every few seconds; where the speech has no full stop, a line is cut after 320 characters or 30 seconds. Engines that return no timestamps at all still produce one line covering the whole file.

The timestamps here are labels, not buttons. Seeking needs a player to drive, and only YouTube transcriptions embed one — see YouTube Transcription. For a local file you read the time and use it in your own player.

Older transcriptions show a single 00:00
Files longer than nine minutes were split into chunks whose timestamps were thrown away, leaving one segment at 00:00. That data was never stored, so an existing transcription cannot gain timestamps — transcribe the file again.
File Transcription — Speakly