Text-to-Speech Model
Synthesize speech on SageMaker with Chatterbox Turbo
What you get
Synthesize natural speech on a SageMaker endpoint from JSON text input
Same CUDA image as the Text-to-Speech Server AMI, with bundled Turbo weights
Real-time endpoints and batch transform on GPU instances (ml.g4dn.xlarge recommended)
About this product
This SageMaker model package converts text to natural speech using Resemble AI's Chatterbox Turbo (MIT license). Deploy a real-time endpoint or run SageMaker batch transform. Send JSON with an input or text field and optional voice, language, and response_format. The response is raw audio (audio/wav or audio/aiff bytes).
GPU is strongly recommended. Deploy on ml.g4dn.xlarge (NVIDIA T4) or larger g4dn/g5 types for faster-than-realtime English synthesis. The same CUDA image falls back to CPU when no GPU is present, but synthesis is much slower than realtime on CPU. Real-time payloads are limited to 6 MB.
This listing is the SageMaker path. For a self-hosted HTTPS web UI, admin console, bundled batch folder processing, and zero-shot cloning uploads, use the Text-to-Speech Server AMI.
Common use cases include IVR and telephony prompts, accessibility narration, content read-aloud, and private voice generation pipelines where text and audio must stay inside your own AWS account.
Model and training data The model is Chatterbox Turbo from Resemble AI, served with weights and bundled reference voices baked into the image. Sigmodata did not train the base checkpoint; it is distributed under the MIT license. English uses the Turbo checkpoint; other languages use the bundled multilingual model when requested via the language field.
Measured performance - On ml.g4dn.xlarge (NVIDIA T4), measured after engine warmup: a one-line English sentence is about 2.5 s wall clock for 5 s of audio (faster than realtime). A short phrase is about 0.8 s; a typical IVR paragraph about 5 s for 11 s of audio. - First endpoint startup includes model load from the container image; plan for several minutes on a new GPU instance before /ping returns 200. - The GPU serializes synthesis; parallel requests queue (MaxConcurrentTransforms=1).
Known limitations - Real-time invocations accept up to 6 MB. Longer inputs should use batch transform (SingleRecord). - One synthesis runs on the GPU at a time per instance. - Zero-shot cloning from an uploaded reference clip is not exposed on this SageMaker endpoint; use the Text-to-Speech Server AMI for that workflow. - CPU-only instances work but are slower than realtime for typical prompts.
We welcome your feedback at [email protected]. Sample notebook: https://www.sigmodata.com/products/text-to-speech-model/getting-started.ipynb
How it ships
- SageMakerSageMaker model package
Text-to-Speech Model
Listing coming soon
Categories and keywords
- Categories
- Text to SpeechMachine LearningAudio
- Keywords
- text-to-speechttsspeech synthesisSageMakerchatterboxnarration