Text-to-Speech Server
Self-hosted Chatterbox Turbo TTS on EC2, GPU-first, air-gapped after launch
What you get
Turn text into natural speech for IVR, narration, and voice apps without a third-party TTS API
Chatterbox Turbo with bundled speakers and zero-shot cloning; air-gapped after launch, no weight downloads
About 3 min first boot on g4dn.xlarge; English synthesis faster than realtime on T4 after warmup
About this product
This product provides a fully self-hosted text-to-speech server packaged as an Amazon Machine Image (AMI) for AWS EC2. The server converts written text into spoken audio using Resemble AI's Chatterbox Turbo (MIT license) and exposes both a web interface and an HTTP REST API. All processing occurs within the customer's AWS environment. Model weights, bundled speakers, and CUDA libraries are baked into the image; the instance does not contact Hugging Face, ECR, or any other external service after launch.
GPU instances are the supported path. The recommended type is g4dn.xlarge (NVIDIA T4). Larger g4dn, g5, and g6 types are also supported. CPU-only types (m6i.2xlarge / m7i.2xlarge and above) run the same model more slowly. Small burst instances such as t3 are not supported and are refused at boot.
The AMI launches with a guided browser setup wizard over HTTPS (port 443). You confirm ownership with the EC2 instance ID, choose an administrator password, and pick a certificate option. No SSH or user-data editing is required for credentials.
After setup completes, Chatterbox Turbo loads onto the GPU from weights baked into the AMI. On a recommended g4dn.xlarge instance, plan for about 3 minutes on first launch from a new EBS volume (cold snapshot read). The web UI and GET /health show loading progress (step and progress_pct) through library import, reading weights from disk, GPU load, and a short warmup synthesis. The progress bar can sit at a low percentage while EBS reads the AMI snapshot for the first time. Reboots on the same volume are usually faster. Nothing is downloaded from external services during startup.
Once ready, English synthesis on a g4dn.xlarge (NVIDIA T4) is faster than realtime for typical sentence-length prompts (about 2.5 s wall clock for a one-line request that produces 5 s of audio). Very short phrases are near realtime. The engine runs one GPU synthesis at a time; parallel requests queue.
Common use cases include IVR and telephony prompts, accessibility applications, content narration, automated announcements, and private voice generation pipelines. Optional zero-shot voice cloning accepts a short reference clip (>5 seconds) on the synthesis API.
How it ships
- EC2 AMIEC2 AMI
Text To Speech Server AMI
Listing coming soon
Categories and keywords
- Categories
- Text to Speech
- Keywords
- text-to-speechttsspeech synthesisvoice generationaudionarrationivrchatterboxgpu