Decision Model
Typed decisions from text or JSON, with calibrated probabilities, inside your VPC
What you get
Typed answers, not generated text: pick one of your options, a yes/no probability or a rubric level, each with calibrated confidence
Runs inside your own AWS account and VPC - your text is never sent to a third-party API
94 percent accuracy on our validation set; about 0.4 s per question on a GPU instance
About this product
Runs inside your own AWS account: your data is processed on a SageMaker endpoint you control, inside your VPC, and is never sent to a third-party API.
Ask typed questions about unstructured input and get answers your code can act on, instead of free text to parse. Send one "state" (text, or any JSON object or array) and up to 64 questions. Every answer is one of the options you supplied:
- choice: one of your named options, with a probability for each - noul: the probability that a yes/no question is true - score: a level on a rubric you define, with the most likely level and the expected level
Typical uses are routing support tickets, flagging churn, fraud or abuse signals, classifying log lines by severity, and scoring reviews. Every answer carries a confidence, so you can act automatically on confident answers and send the rest to a person. Questions about one state are evaluated together, so adding a question is cheap.
The request and response shapes follow the public Jev System One API, so existing Jev clients work unchanged. Real-time endpoints and SageMaker batch transform (JSON Lines, one request per line) are supported. Use a GPU instance (ml.g4dn.xlarge or ml.g5.xlarge) for production speed; CPU instances (ml.m5.xlarge) suit low volume. Batch transform is validated on ml.m5.xlarge and ml.g4dn.xlarge.
We welcome your feedback at [email protected].
Known limitations
- English-centric. Other languages are not evaluated. - The context window is 4096 tokens. Your questions and options are kept intact first, and a long state is truncated from the end. - One evaluation runs at a time per instance. Add instances for more throughput. - Score questions are the least accurate type. Act on "level" and check "confidence".
Model and training data
Strands Decider 2B from AWS (Apache-2.0), a LoRA fine-tune of Qwen3.5-2B, served as published. The authors' training data is described in the model card: https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19. The container runs in network isolation, so no data leaves your account.
Measured performance
- Accuracy of 94 percent on our validation set, the same on GPU and CPU. - One question, median and 95th percentile: ml.g4dn.xlarge 451 ms and 959 ms; ml.g5.xlarge 389 ms and 994 ms; ml.m5.xlarge 2.2 seconds. - Twelve questions in one request: ml.g5.xlarge 0.6 seconds (about 18 questions per second); ml.g4dn.xlarge 1.6 seconds (about 7 per second); ml.m5.xlarge 14.1 seconds.
How it ships
- SageMakerSageMaker model package
Decision Model
View on AWS Marketplace
Categories and keywords
- Categories
- Natural Language ProcessingText
- Keywords
- decisionclassificationroutingtriagestructured outputSageMaker