Speech-to-Text Server
Self-hosted speech-to-text on EC2 and containers, GPU optional
What you get
Transcribe calls, meetings, and media in your VPC - audio never leaves your AWS account
Add users and revocable API tokens from /admin - not one shared password like most AMI listings
HTTPS wizard, automatic GPU or CPU, and a Batch tab for whole S3 folders
About this product
This product provides a fully self-hosted speech-to-text server packaged as an Amazon Machine Image (AMI) for AWS EC2 and as a container image. The server converts audio files into text and exposes both a web interface and a simple HTTP REST API. All processing occurs within the customer's AWS environment.
The AMI launches with a guided browser setup wizard over HTTPS (port 443). You confirm ownership with the EC2 instance ID, create an administrator account, pick a certificate option, and select a Whisper model. No SSH or user-data editing is required for credentials. After setup, the admin account can add users and issue API tokens from /admin - no need to share one password for automation. HTTP Basic auth with any account still works. GPU instances (for example g4dn.xlarge) enable CUDA automatically; CPU instances still work.
The Marketplace container image authenticates with a PASSWORD environment variable over HTTP on port 8080 and auto-detects a CUDA GPU at startup (pass --gpus all to enable it), falling back to CPU otherwise.
A Batch tab handles whole folders rather than one file at a time. Point it at an S3 bucket from /admin, and it becomes a file browser over that bucket: upload audio, queue transcriptions, and get a sidecar transcript written back beside each source file. Nothing leaves the instance and the bucket stays in your account. Without a bucket configured, any writable directory bind-mounted at /data works the same way.
Common use cases include call center transcription, media and meeting transcription, bulk back-catalogue conversion, compliance-sensitive audio processing, and private speech analytics pipelines.
Whisper models base.en is the default and the only model baked into the image, so the server transcribes English out of the box with no download. Sixteen checkpoints are selectable in total - from the Model tab in /admin on the AMI, or by running /app/scripts/download-ggml-model.sh into a mounted /app/models volume and setting WHISPER_MODEL on the container. Models ending in .en are English-only; the rest are multilingual.
- tiny, tiny.en (39M, ~1 GB VRAM) - the smallest checkpoints; suited to drafts and keyword spotting rather than final transcripts - base (74M, ~1 GB VRAM) - small multilingual, fine for short clips - base.en (74M, ~1 GB VRAM) - the default; good English accuracy, small footprint, and the only model that needs no download - small, small.en (244M, ~2 GB VRAM) - noticeably better than base - small.en-tdrz (244M, ~2 GB VRAM) - small.en with tinydiarize speaker turns; required for the Diarize option in the UI - medium, medium.en (769M, ~5 GB VRAM) - strong accuracy, needs a full GPU - large-v1, large-v2, large-v3 (1550M, ~10 GB VRAM) - the full-size checkpoints; the largest and slowest of the set - large-v2-q5_0, large-v3-q5_0 (1550M quantized, ~6 GB VRAM) - most of the quality at roughly half the disk and VRAM - large-v3-turbo (809M, ~6 GB VRAM) - close to large-v3 quality at roughly half the parameters - large-v3-turbo-q5_0 (809M quantized, ~4 GB VRAM) - turbo quality on smaller GPUs
Selecting any model other than base.en downloads it from huggingface.co, so the instance or container needs outbound internet access at that point. Transcription itself never sends audio anywhere.
How it ships
- EC2 AMIEC2 AMI
Speech-to-Text Server AMI
View on AWS Marketplace - ContainerContainer image
Speech-to-Text Server Container
View on AWS Marketplace
Categories and keywords
- Categories
- Speech to TextSpeech RecognitionAudio
- Keywords
- speech-to-textaudiotranscriptionsubtitlecaptioningbatch transcription