Contents
Key takeawaysWhat the plans coverDiscount and commitment costSizing the planStacking with an EDPWhat we have seenWhat to do nextFAQSageMaker Savings Plans take up to 64 percent off SageMaker instance usage for a fixed hourly commitment, but ML workloads change faster than the term. Commit only to the inference floor left after right sizing, on one year terms.
- Two separate families. Compute Savings Plans never cover SageMaker usage, and a SageMaker plan never covers EC2, Fargate or Lambda.
- The deepest rate needs the longest lock. The top discount comes only with three year all upfront, and the hourly commitment bills whether you use it or not.
- ML commitments strand. Across our engagements, three year all upfront plans lost 20 to 40 percent of their value when the ML platform changed within two quarters.
- Right size first. Notebook shutdown schedules, spot training and endpoint autoscaling cut the required commitment by 15 to 30 percent before any plan was signed.
- Commit inference and leave training out. Steady serving suits a one year plan, while spiky, migration prone training belongs on spot and on demand.
- Sequence against the EDP. Negotiate the private rate on the full ML forecast, then commit the plan only to the pessimistic floor.
SageMaker Savings Plans lower the rate you pay for SageMaker ML instance usage in return for a fixed dollar amount per hour, committed for one or three years. The discount is real. The difficulty is that ML workloads change platform, instance generation and hosting model faster than any three year term allows.
This guide covers what the plans include, what stranded commitment does to the headline rate, how to size a plan, and how it stacks with an AWS Enterprise Discount Program. The worked examples use hypothetical numbers you can rerun on your own bill. For the wider picture, see our AWS Savings Plans hub.
What do SageMaker Savings Plans cover, and what do they leave out?
They cover SageMaker ML instance usage and nothing else. AWS lists Studio notebooks, notebook instances, Processing, Data Wrangler, Training, Real-Time Inference and Batch Transform. The plan rate applies across instance family, size, Region and component, so shifting inference from an ml.c5 instance in one Region to an ml.inf1 instance in another keeps the discount.
AWS now labels them SageMaker AI Savings Plans in the console. Compute Savings Plans are a separate family, and the two never cover each other: EC2, Fargate and Lambda spend draws down a Compute plan, while SageMaker usage draws down only a SageMaker plan.
| Plan | What draws it down | What never does |
|---|---|---|
| Compute Savings Plans | EC2, Fargate and Lambda, across instance families, Regions and operating systems | Any SageMaker usage: notebooks, training or inference |
| SageMaker Savings Plans | Studio notebooks, notebook instances, Processing, Data Wrangler, training jobs, inference endpoints and Batch Transform | ML you host yourself on EC2, Bedrock consumption, third party serving |
| What it means for you | Two meters, sized separately, each against its own workload floor | One blended commitment covering all ML spend does not exist |
Buyers who assumed one commitment covered all their AWS compute ran uncovered SageMaker at on demand rates next to an underused Compute plan, paying twice for one mistake. EC2 Instance and Database Savings Plans do not reach SageMaker either.
Where does the boundary between the two plans shift?
Architecture decisions shift spend between meters. A model you train and serve yourself on EC2 draws the Compute plan. Put the same model into SageMaker and it stops drawing that plan and needs the other family. Rebuild it on Bedrock and it draws neither.
So the commitment map has to be redrawn whenever the platform roadmap is. A blended forecast that ignores the line strands value on both sides of it: unused Compute commitment where ML left EC2, and on demand SageMaker where no plan covers it.
Which ML charges does the SageMaker plan not cover?
- Bedrock. Model calls bill by tokens or through Bedrock's own Provisioned Throughput, outside both Savings Plans families.
- Non instance SageMaker charges. AWS states the plan applies only to SageMaker ML instance usage, so storage volumes, S3 data and data transfer bill at their normal rates.
- Self hosted serving on EC2. This draws a Compute or EC2 Instance plan, never the SageMaker plan.
- Third party platforms. Serving platforms you pay outside AWS, and the software fees on Marketplace listings, sit outside both plan families.
Negotiating AWS 3: The Discount Stack
How much do SageMaker Savings Plans save, and what does the commitment cost?
AWS advertises up to 64 percent off on demand rates, reached only at three year all upfront. Terms are one or three years, paid all upfront, partial upfront or no upfront. The rate steps down as you shorten the term or defer payment, and it varies by instance type, so check your own mix on the purchase page.
What you commit to is an hourly spend at plan rates, billed every hour of the term whether you use it or not. Usage above the commitment bills at on demand. Once the short return window described in the FAQ below has passed, the commitment cannot be changed or cancelled.
Why do ML workloads fit a fixed hourly commitment badly?
Training is spiky and migrates to each new instance generation for the price performance. Inference is steady until the architecture changes. The plan rate carries over to a new instance family, yet the newer instance often does the same work for fewer billable dollars, which leaves less usage to absorb the commitment.
Say a training job ran 10 hours on the old generation and takes 4 on the new one. That instance may cost more per hour, but the job costs less overall, so average hourly spend falls while the commitment bills at the old level. Inference moving to Bedrock or self hosted serving completes the stranding pattern.
What does stranding do to the headline discount?
Say your SageMaker usage runs at $100 an hour at on demand prices, and you buy a three year all upfront plan at 64 percent. The commitment is $36 an hour, or $946,080 paid on day one for 26,280 hours. The table shows what happens when part of that usage leaves SageMaker within two quarters.
| Usage that leaves | Remaining usage at on demand value | Plan dollars still used per hour | Stranded per hour | Effective discount on what remains |
|---|---|---|---|---|
| None | $100 | $36.00 | $0 | 64 percent |
| 20 percent | $80 | $28.80 | $7.20 | 55 percent |
| 40 percent | $60 | $21.60 | $14.40 | 40 percent |
| 64 percent | $36 | $12.96 | $23.04 | 0 percent |
Two things follow from the arithmetic. The share of commitment stranded equals the share of usage that left. Once usage falls by the discount rate itself, the plan costs exactly what on demand would have. In the 40 percent row, this plan wastes $14.40 an hour, which is $126,144 a year.
Why we advise against buying the longest term for the deepest rate
Account teams and many FinOps guides push for maximum coverage on three year all upfront, since it carries the best rate. For a stable EC2 fleet that is reasonable. For ML we disagree. As the table shows, a 64 percent plan that loses 40 percent of its usage does no better than a 40 percent plan with nothing stranded.
Take one year terms on inference floors instead, and keep training out of the commitment. You accept a lower headline rate in exchange for the chance to reset the plan every 12 months.
AWS EDP Negotiation Guide
The full commitment stack, from private rate to Savings Plans, with the terms worth asking for.
Get the white paper →How should you size a SageMaker Savings Plan?
Size it after you remove waste, and only to the usage you would still run in a pessimistic year. In our engagements, right sizing before committing cut the required commitment by 15 to 30 percent, so the plan priced only the usage that survived cleanup.
What should you fix before you buy?
- Idle notebooks. SageMaker Studio can shut down idle JupyterLab and Code Editor applications automatically, set by administrators at domain level. SageMaker notebook instances need a lifecycle configuration script to do the same.
- Training on Managed Spot Training. AWS puts the saving at up to 90 percent against on demand. Jobs need checkpointing to survive interruption, and you set EnableManagedSpotTraining and MaxWaitTimeInSeconds on the job.
- Endpoint autoscaling. Real-Time Inference endpoints scale through Application Auto Scaling, and endpoints built on inference components can scale down to zero instances when idle.
- Instance choice. SageMaker Inference Recommender load tests candidate instance types for a model, which shows whether a smaller or newer instance can serve the same traffic.
- Batch where latency allows. Batch Transform bills only while the job runs, so scoring that tolerates a delay does not need a live endpoint.
Worked example: sizing after the cleanup
Take a hypothetical ML platform averaging $100 an hour of SageMaker usage at on demand prices. The table splits that spend by workload and shows what belongs in a commitment after cleanup.
| Workload | Observed average | After right sizing | Goes into the plan |
|---|---|---|---|
| Studio notebooks | $18 | $8 with idle shutdown | $0, usage follows working hours |
| Training jobs | $32 | $32 of demand, run on Managed Spot | $0, spot and on demand only |
| Real-Time Inference | $50 | $41 with autoscaling | $30, the lowest sustained hourly level over 90 days |
| Total | $100 | $81 before spot savings | $30 |
The observed average invites a commitment sized to $100 an hour. After cleanup the on demand equivalent is $81, 19 percent lower, and training drops out of the plan entirely. The commitment covers $30 an hour of inference, about 73 percent of right sized endpoint spend, converted to plan dollars at your instance mix's rate.
Commit to the pessimistic floor
The rule our CUD sizing analysis applies to Google Cloud works unchanged here. Commit to what you would consume in the pessimistic scenario, and treat the forecast as an upper bound. Anything above that floor is cheaper to buy later, in a second plan, once the usage has proved itself.
Larger ML platforms can buy one year plans in tranches a quarter apart. Each plan then expires on its own date, and you get four chances a year to cut commitment that a roadmap change has made surplus. A small team with a few endpoints is often better off with one modest one year plan, or none.
How can you check your own position?
- Cost Explorer recommendations. Select SageMaker Savings Plans as the plan type. The figure is built from past usage, so run it after cleanup.
- Purchase analysis. The Savings Plans purchase analyzer in the Billing and Cost Management console models a custom commitment amount against your historical usage.
- Savings Plans reports in Cost Explorer. The report on unused commitment shows commitment you pay for and do not use. The coverage report shows SageMaker usage still billing at on demand.
- Cost and Usage Report. With hourly granularity turned on, plan charges appear line by line per hour, which is where you trace stranded commitment to the endpoint or job that left.
How do SageMaker Savings Plans stack with an AWS EDP?
They stack, and the two discounts multiply rather than add. The effective ML rate is the public on demand price reduced by both the plan percentage and the EDP private rate, which makes it the product of two decisions made on different calendars.
Say your instance mix earns 50 percent on the plan and your EDP gives 10 percent. $100 of on demand usage then costs $100 times 0.5 times 0.9, or $45. That is a 55 percent effective discount, short of the 60 percent a simple sum suggests.
Committed SageMaker spend also counts as real commit dollars toward the EDP's growth requirement. That turns the order of the two into a negotiation decision. Our EDP benchmark data shows what the private rate should concede at each tier.
Which should you negotiate first?
Negotiate the EDP rate on the full ML forecast first, then commit Savings Plans only to the floor. ML spend that shifts from SageMaker to Bedrock still counts toward the EDP, while the plan penalizes every dollar that leaves SageMaker. The EDP has its own risk if total AWS spend falls short, which our note on EDP shortfall risk covers.
What will the AWS account team say, and how should you answer?
- "Cost Explorer recommends this hourly commitment." Reply that the recommendation reflects past usage, including idle notebooks and training that will move to spot. You will size after cleanup and against the platform roadmap.
- "Three year all upfront gets you the full discount." Ask for the plan rate on your actual instance types, then show the stranding table. A lower rate on a one year term limits the damage to 12 months if part of the usage leaves.
- "Your Compute Savings Plan already covers ML." It covers only ML you host yourself on EC2. Ask for a coverage report that shows SageMaker usage separately.
- "A plan purchase this quarter helps you hit the EDP commitment." Agree that it counts, then decline to buy commitment on workloads that may leave SageMaker. Close an EDP gap with spend you are sure of.
What have we seen in recent SageMaker commitment negotiations?
Across roughly 20 to 30 AWS commitment engagements we benchmarked between 2024 and 2026, the ML commitments aged worst. Three year all upfront plans stranded 20 to 40 percent of their value when the ML platform changed within two quarters.
The failure pattern was consistent. Buyers committed training and inference together, when only inference had a steady floor, and the training share stranded the moment a new instance family or a Bedrock migration shifted the workload.
Commit the inference floor a year at a time, keep training on spot and on demand, and redraw the commitment map whenever the platform roadmap changes.
The successes were just as consistent. Inference floors were committed on one year terms, training stayed on spot and on demand, and the commitment map was reviewed against the platform roadmap every quarter, the cadence at which ML workloads actually change.
What to do next
- Map spend to the two families. Tag what draws the Compute plan, what draws the SageMaker plan and what draws neither, because that boundary decides both commitments.
- Right size before committing. Turn on idle shutdown for notebooks, move training to Managed Spot Training, add endpoint autoscaling, and let the plan price the optimized floor.
- Commit only the inference floor. Put steady serving on one year terms and keep spiky training on spot and on demand, never one blended number.
- Sequence against the EDP. Negotiate the private rate on the full forecast first, then commit the Savings Plan to the floor after.
- Check every purchase within 7 days. Confirm the plan type, term and hourly amount while the return window is still open.
- Review the commitment map quarterly. Compare it with the platform roadmap, because ML changes platform faster than any three year term. Our cost optimization practice runs the sizing with you.
Is an AWS commitment coming up? Our AWS EDP negotiation team sizes it to your usage and works only for buyers.
Frequently asked questions
What are SageMaker Savings Plans?
They are AWS commitment plans that lower the rate on SageMaker ML instance usage in exchange for a fixed dollar amount per hour over one or three years. The rate applies to notebooks, Processing, Data Wrangler, training, real time inference and Batch Transform. They form a separate family from Compute Savings Plans, and neither plan's commitment covers the other's usage.
Do Compute Savings Plans cover SageMaker?
No. Compute Savings Plans cover EC2, Fargate and Lambda, and SageMaker usage draws down only a SageMaker Savings Plan. Every change of platform between EC2, SageMaker and Bedrock shifts spend from one meter to another, so recheck coverage after each migration as well as at purchase.
How much do SageMaker Savings Plans save?
AWS advertises up to 64 percent at three year all upfront, with lower rates on one year terms and on partial or no upfront payment. What you actually save depends on how much of the commitment you consume, because every committed hour you do not use comes straight out of the headline discount. Before buying, get the plan rate for your own instance types under each term and payment option.
Why do ML commitments get stranded?
The plan commits dollars per hour, while ML work keeps getting cheaper per job and keeps changing platform. Newer instances finish the same training in fewer billable hours, inference shifts to Bedrock or self hosted serving, and the commitment keeps billing against usage that has gone elsewhere.
Should we commit training and inference together?
Usually not. Inference has the steady hourly floor a commitment needs, while training runs in bursts and jumps to each new instance generation, so spot and on demand suit it better. A single blended number was the most consistent failure in our benchmarks, and the training share was the part that stranded first.
How do SageMaker Savings Plans interact with an AWS EDP?
They stack. Both the plan rate and your EDP discount apply to the same usage, so the combined discount beats either one alone but falls short of their sum. Committed SageMaker spend also counts toward the EDP's commit requirement, so settle the EDP on the full ML forecast before you buy, then size the plan to the pessimistic floor.
Can you cancel or return a SageMaker Savings Plan?
Only in a narrow window. AWS accepts returns of a plan with an hourly commitment of $100 or less within 7 days of purchase and in the same calendar month, refunds any upfront charges, and caps how many returns each billing family can make. Larger plans cannot be returned, and no plan can be cancelled once the window closes.
Do SageMaker Savings Plans cover Amazon Bedrock?
No. Bedrock bills by tokens or through its own Provisioned Throughput, which sits outside both Savings Plans families. A model moved from a SageMaker endpoint to Bedrock stops drawing down your SageMaker plan the day it goes live, so plan the commitment around any Bedrock migration on the roadmap.