ST

Speech-to-Text Model

Coming SoonSageMaker

Transcribe English audio on SageMaker, GPU optional

What you get

  • Transcribe English audio on a SageMaker endpoint - WAV, MP3, or JSON base64

  • Same CUDA image as the Speech-to-Text Server, with automatic CPU fallback

  • Real-time endpoints and batch transform; GPU if you have one, CPU if you don't

About this product

This SageMaker model package transcribes English speech using whisper.cpp (baked base.en weights). Deploy a real-time endpoint or run SageMaker batch transform. Send JSON with a base64-encoded audio field, or raw audio bytes with Content-Type audio/wav, audio/mpeg, or application/octet-stream. The response is JSON with a text field.

GPU is optional. Deploy on a GPU instance (for example ml.g4dn.xlarge) for higher throughput; the same image uses CUDA automatically when a GPU is present, and falls back to CPU otherwise. Real-time payloads are limited to 6 MB; longer files should use batch transform.

This listing is the SageMaker path. For a self-hosted HTTP API, web UI, and selectable Whisper models, use the Speech-to-Text Server AMI or container.

Common use cases include transcribing call-centre recordings, meeting and interview audio, podcast and media back-catalogues, and voice notes, and building private speech analytics pipelines where audio must stay inside your own AWS account.

Model and training data The model is OpenAI's Whisper base.en, an English-only encoder-decoder speech recognition model, served through whisper.cpp with the weights baked into the image. Sigmodata did not train it; it was trained by OpenAI on a large corpus of web audio and is used here under the MIT licence. base.en is the small end of the Whisper family, chosen so the model runs acceptably on CPU.

Measured performance - Word error rate 4.4 percent on LibriSpeech test-clean, measured over 300 utterances and 6,428 reference words. Scoring lowercases, strips punctuation other than apostrophes, and collapses whitespace. - Measured by sending the audio to the same container image the model package ships, through the same /invocations endpoint a buyer calls. - LibriSpeech test-clean is read speech recorded in good conditions. Expect a higher error rate on spontaneous conversation, telephony audio, background noise, or strong accents.

Known limitations - English only. Other languages are not supported; use a multilingual Whisper model if you need them. - base.en is the smallest English Whisper model. Larger models are more accurate; the Speech-to-Text Server AMI and container let you pick one. - The response is plain text. Word-level timestamps, speaker diarisation, and segment metadata are not returned by this endpoint. - Real-time invocations accept up to 6 MB. Use batch transform for longer audio; JSON base64 inflates a file by roughly one third. - Accuracy degrades on overlapping speech and on very short clips with little context.

We welcome your feedback at [email protected]. Sample notebook: https://www.sigmodata.com/products/speech-to-text-model/getting-started.ipynb

How it ships

  • SageMakerSageMaker model package

    Speech-to-Text Model

    Listing coming soon

Categories and keywords

Categories
Speech RecognitionTextMachine Learning
Keywords
speech-to-texttranscriptionWhisperSageMakerASRaudio