Category
Speech-to-Text APIs pricing
Speech-to-Text APIs convert audio into text, spanning batch (async file) and real-time streaming transcription, with add-ons like speaker diarization, translation, and PII redaction. The category splits into focused voice-AI specialists (Deepgram, AssemblyAI, Speechmatics, Gladia, Rev AI) optimized for accuracy, latency, and generous self-serve free tiers, versus hyperscaler platforms (Google, AWS) and the model-API generalist (OpenAI Whisper) that ride massive infrastructure but offer thinner DX and stingier free tiers. Compare on documentation/DX quality, reliability and proven scale, SDK breadth and ecosystem, and how fast a developer or AI agent can self-serve a working key against transparent public pricing.
| API | Billed by | Free tier | Verified |
|---|---|---|---|
Deepgram Speech-to-Text APIs | Per minute (usage-based) | Free tier or trial | Jun 27, 2026 |
AssemblyAI Speech-to-Text APIs | Per minute (usage-based) | Free tier or trial | Jun 27, 2026 |
OpenAI Whisper / GPT-4o Transcribe Speech-to-Text APIs | Per minute (per-second billed) | No free tier | Jun 27, 2026 |
Google Cloud Speech-to-Text Speech-to-Text APIs | Per 15 seconds | Free tier or trial | Jun 27, 2026 |
Amazon Transcribe Speech-to-Text APIs | Per minute (volume-tiered) | Free tier or trial | Jun 27, 2026 |
Speechmatics Speech-to-Text APIs | Per minute (usage-based) | Free tier or trial | Jun 27, 2026 |
Gladia Speech-to-Text APIs | Per hour (usage-based) | Free tier or trial | Jun 27, 2026 |
Rev AI Speech-to-Text APIs | Per minute (usage-based) | Free tier or trial | Jun 27, 2026 |