Performance Tips

Optimize Speakly for your system

Local Model Performance

  • Use the quantized turbo Q5 (large-v3-turbo-q5_0) for faster results
  • On Apple Silicon macOS, tick Enable next to GPU acceleration (Metal) under Settings → Transcription; on Intel, Windows and Linux the control does not appear, because no acceleration is available
  • Close other heavy applications during transcription
  • Consider BYOK for faster processing on slower hardware

Memory Usage

  • Larger models need more memory and more CPU time — if transcription is slow, drop one size before trying anything else
  • Close unused tabs/applications if memory is limited
Performance Tips — Speakly