SageMaker Savings Plans, the discount and the stranding
SageMaker Savings Plans discount ML compute up to 64 percent for a committed hourly spend, and the commitment prices certainty in a workload class that has none: across our engagements, three year all upfront commitments stranded 20 to 40 percent of their value when the ML platform changed within two quarters. The discount is real, and so is the discipline it demands.
Prepared by Redress Compliance · August 6, 2026 · AWS advisory. Based on 20 to 30 AWS commitment engagements benchmarked 2024 to 2026.
Executive summary
The plan is its own family. SageMaker Savings Plans and Compute Savings Plans are separate commitments that do not cover each other: EC2, Fargate, and Lambda spend draws down a Compute plan, and SageMaker usage draws down only a SageMaker plan.
Buyers who assumed one commitment covered the estate ran uncovered SageMaker at on demand rates beside an underused Compute plan, paying twice for one mistake.
The discount schedule rewards the longest lock. The plan discounts run up to 64 percent at three year all upfront, stepping down through one year and no upfront variants, and cover the SageMaker family broadly: Studio notebooks, training jobs, and inference endpoints all draw the committed rate.
The percentages are genuine, and every point of them is priced in flexibility surrendered.
ML workloads move faster than ML commitments.
Across our engagements, three year all upfront commitments stranded 20 to 40 percent of their value when the workload changed within two quarters: training migrated to newer instance families the moment they launched, inference re-platformed to Bedrock or self hosted serving.
And the committed hourly spend kept billing against usage that no longer existed.
Sizing is the whole game. Right sizing before committing, notebook shutdown schedules, training spot usage, endpoint autoscaling, cut the required commitment 15 to 30 percent before any plan signed, and the plan then priced the optimized floor rather than the observed waste.
The commitment stacks with the EDP, plan discounts applying after the private rate, which makes the sequencing a negotiation decision, not an afterthought.
The two plan families, and what each one covers
| Plan | What draws it down | What never does |
|---|---|---|
| Compute Savings Plans | EC2, Fargate, Lambda, across families, regions, and OS | Any SageMaker usage, notebooks, training, or inference |
| SageMaker Savings Plans | Studio notebooks, training jobs, inference endpoints, the SageMaker family | EC2 hosted ML you run yourself, Bedrock consumption, third party serving |
| The consequence | Two meters, sized separately, each against its own workload floor | One blended commitment covering the ML estate, which does not exist |
The boundary is the first audit. Self managed ML on EC2 draws the Compute plan; the same model moved into SageMaker stops drawing it and starts needing the other family; the same workload re-platformed to Bedrock draws neither.
Every architecture decision moves spend between meters, which is why the commitment map has to be redrawn whenever the platform roadmap is, and why blended forecasts strand value on both sides of the line.
The discount schedule, and what the lock actually buys
The schedule prices flexibility surrendered: three year all upfront reaches the headline 64 percent, one year no upfront sits far below it, and the variants in between trade cash timing against rate.
The committed unit is an hourly spend floor, billed whether consumed or not, which is precisely the construction that punishes ML's actual behavior: training is spiky and migrates to each new instance generation for the price performance, inference is steady until the architecture changes.
And the two together produce the stranding pattern.
The commitment discipline the CUD sizing analysis works for Google applies unchanged here: commit to the floor you would consume in the pessimistic scenario, never the forecast.
The AWS EDP negotiation playbook
The commitment stack in full: the private rate, the Savings Plan layering, the growth curve, and the concessions that survive a replatform.
Get the white paper →The EDP stack, discounts in the right order
For estates on an Enterprise Discount Program the two discounts stack: the EDP private rate applies first, and the Savings Plan discount applies to the discounted rate, which makes the effective ML rate a product of two negotiations conducted on different calendars.
The sequencing matters at the EDP table, committed SageMaker spend is real commit dollars toward the EDP's growth requirement, and the EDP benchmark data prices what the private rate should concede at each tier.
The practical order: negotiate the EDP rate on the full ML forecast, then commit Savings Plans only to the floor, because the EDP discounts the ambition safely while the plan punishes it.
- Percentile standing for your exact deal size and industry, from real closed transactions
- Scenario simulation before the call: test alternative terms and see the financial impact of each
- A negotiation playbook, talking points, and a two page executive brief on day one
What we saw across commitment engagements, 2024 to 2026
Across roughly 20 to 30 AWS commitment engagements benchmarked between 2024 and 2026, the ML commitments were the ones that aged worst:
Three year all upfront commitments against ML platforms that changed within two quarters.
What right sizing cut from the required commitment before any plan signed.
The failure pattern was consistent: committing training and inference together, when only inference had a steady floor, and the training share stranded the moment a new instance family or a Bedrock migration moved the workload.
The successes were equally consistent, inference floors committed on one year terms, training left on spot and on demand, and the commitment map reviewed against the platform roadmap quarterly, the cadence the workload actually changes on.
Your first five moves
- Map spend to the two families first: what draws Compute, what draws SageMaker, what draws neither, because the boundary decides both commitments.
- Right size before committing: notebook shutdown schedules, spot for training, endpoint autoscaling, and let the plan price the optimized floor.
- Commit inference floors, not training peaks: steady serving on shorter terms, spiky training on spot and on demand, never one blended number.
- Sequence against the EDP: the private rate negotiated on the full forecast first, the Savings Plan committed to the floor after.
- Review the commitment map quarterly against the platform roadmap, because ML re-platforms faster than any three year paper. The cost optimization practice runs the sizing with you.
Frequently asked questions
What are SageMaker Savings Plans?
AWS commitment plans that discount SageMaker ML compute, Studio notebooks, training jobs, and inference endpoints, up to 64 percent at three year all upfront in exchange for a committed hourly spend.
They are a separate family from Compute Savings Plans: neither plan's commitment covers the other's usage.
Do Compute Savings Plans cover SageMaker?
No. Compute Savings Plans cover EC2, Fargate, and Lambda; SageMaker usage draws down only a SageMaker Savings Plan.
Estates that assumed one commitment covered both ran uncovered SageMaker at on demand rates beside an underused Compute plan, and every replatform between EC2, SageMaker, and Bedrock moves spend between the meters.
How much do SageMaker Savings Plans save?
Up to 64 percent at three year all upfront, stepping down through one year and no upfront variants.
The realized saving depends on utilization: the committed hourly spend bills whether consumed or not, and the 20 to 40 percent stranding we measured on changed workloads came straight out of the headline discount.
Why do ML commitments get stranded?
Because ML platforms change faster than commitment paper: training migrates to each new instance generation, inference re-platforms to Bedrock or self hosted serving, and the committed spend keeps billing against usage that moved.
Across our engagements, three year all upfront commitments stranded 20 to 40 percent of their value when the workload changed within two quarters.
Should we commit training and inference together?
Usually not: inference carries the steady floor that suits a commitment, while training is spiky and migration prone, better served by spot and on demand.
Committing both as one blended number was the most consistent failure pattern in our benchmarks, stranding the training share at the first platform change.
How do SageMaker Savings Plans interact with an AWS EDP?
They stack: the EDP private rate applies first and the plan discount applies to the discounted rate, while committed SageMaker spend counts toward the EDP's commit requirement.
Negotiate the EDP on the full ML forecast first, then commit the plan only to the pessimistic floor, because the EDP discounts ambition safely and the plan punishes it.