How an ML startup cut its $22,000 AWS bill by fixing scheduling
An ML startup reduced its monthly GPU costs by $4,800 not by optimizing code, but by scaling down instances during predictable low-traffic hours.
A machine learning startup with eight engineers recently slashed its monthly cloud infrastructure bill by nearly five thousand dollars. The team achieved these savings not by rewriting their inference models or changing instance types, but by aligning their GPU usage schedule with actual user demand patterns. The case highlights how operational timing, rather than just architectural efficiency, drives significant cost differences in cloud environments.
What happened
The startup was running a production inference API that served real traffic, resulting in an AWS bill of $22,000 per month. The majority of this cost came from GPU compute resources used for model inference. When a consultant reviewed the account, they expected to find over-provisioned hardware or inefficient code. Instead, they found that the instance types were well-matched to the workload and utilization rates during active hours were reasonable. There were no obvious technical wastes in the architecture itself.
The root cause of the high cost was temporal, not technical. Analysis of the traffic logs revealed that ninety percent of all user requests occurred between 9am and 11pm US Eastern time. The overnight period was nearly silent, consisting only of automated health checks and minor background jobs. Despite this lack of demand, the GPU instances ran at full capacity and full price for twenty-four hours a day. The team had kept the cluster at peak size overnight because it felt safer, leading to eight hours of near-zero utilization every single night.
To address this, the team implemented a scale-down schedule. At 11pm Eastern, the inference cluster scales down to a minimal warm state. This state is sufficient to handle background tasks and respond to health checks but consumes far fewer resources. At 8am Eastern, before the morning traffic surge, the cluster scales back up to full capacity. The entire change took only two days to implement and test safely.
Key details
- The startup’s monthly AWS bill was $22,000, primarily driven by GPU compute costs for inference.
- Traffic analysis showed 90% of requests occurred between 9am and 11pm US Eastern time.
- GPU instances were running at full capacity 24/7, despite eight hours of near-zero user traffic nightly.
- The solution involved scaling down to a minimal warm state at 11pm and scaling up at 8am Eastern.
- The implementation took two days to test and deploy without changing the underlying architecture or models.
- The adjustment resulted in monthly savings of $4,800 by matching resource allocation to actual demand.
Background
In cloud computing, particularly with specialized hardware like GPUs, costs are often tied to uptime rather than just active processing. Many teams default to keeping clusters at maximum capacity to ensure low latency and high availability, fearing that scaling down might introduce risks or delays. This approach, often called "over-provisioning for safety," can lead to significant waste during predictable lulls in traffic.
FinOps, or financial operations, is the practice of managing cloud costs through collaboration between engineering and finance teams. A core principle of FinOps is matching spend to business value. In this context, it means ensuring that expensive resources are only running when they are actively serving users. Automated scaling policies allow teams to dynamically adjust resources based on time or load, bridging the gap between technical reliability and financial efficiency.
Why it matters
For teams that manage their own software infrastructure, this case illustrates that cost optimization does not always require complex refactoring. Engineers often focus on code efficiency or database indexing to reduce costs, overlooking operational schedules. When traffic patterns are predictable, such as in B2B applications with distinct business hours, static scaling policies can leave money on the table. Recognizing that "safety" via constant peak capacity has a direct financial penalty is a crucial shift in mindset for DevOps leads.
Furthermore, this story underscores the importance of data-driven decision-making in infrastructure management. Without analyzing the timestamp distribution of requests, the team assumed their high bill was justified by constant demand. By simply observing when users were actually active, they identified a clear opportunity for savings. This approach is applicable to many self-hosted services where load fluctuates predictably, from internal tools to customer-facing APIs.
What you can do
- Analyze your service logs to identify peak and off-peak traffic windows, looking for predictable daily or weekly patterns.
- Review your current auto-scaling policies to see if they allow for deep scale-downs during known low-traffic periods.
- Implement a "warm state" configuration that keeps essential health checks and background workers running while reducing expensive compute nodes.
- Test scaling events in a staging environment to ensure that ramp-up times do not impact user experience during traffic surges.
- Set up alerts for unusual traffic spikes outside of expected windows to catch anomalies without keeping resources high permanently.
- Regularly revisit your scaling schedule as user behavior changes, ensuring your cost structure remains aligned with actual demand.



