Lessons from Real-World Case Studies on Cloud Waste
It started with a Slack message in January 2022. Our CEO pinged me about gross margins. They were hovering around 50%.
Fifty percent is not a catastrophe, but it is nowhere near the 80%-plus that elite software companies boast. When markets tighten and capital stops flowing like water, gross margin is the metric that separates surviving businesses from distressed ones. We decided to run a yearlong internal experiment: use our own platform, CloudZero, to overhaul our cloud architecture and lift our gross margin.
Most companies facing this problem immediately throw automated tooling at the bill or pressure engineers to cut features. We took a different path. We wanted to see what happens when engineers get actual visibility into unit economics. The result? We uncovered $1.7M in annualized cloud savings in just 54 hours of engineering time. Even better, 88% of those savings came from engineering architecture decisions, not generic automated rightsizing or vendor discounting.
Step 1: Allocating 100% of Spend Without Tagging
Spend allocation is usually where cloud cost optimization projects go to die. Traditional allocation relies entirely on tagging resources. Anyone who has managed enterprise cloud infrastructure knows that tagging is a fool's errand. It requires continuous human discipline, breaks whenever teams reorganize or merge, and fails entirely when handling shared or untaggable resources.
To bypass tagging hell, we used CostFormation—our code-driven cost allocation method. It ingested 100% of our cloud bill and mapped it directly to our business structure without requiring a single new tag. We broke our spending into three core buckets: POCs for sales trials, internal R&D, and Customer Direct COGS.
When we looked at the breakdown, the picture became crystal clear. Direct COGS was the only segment eating into our gross margins. If we wanted to move the needle, we had to look at what specific features, services, and customer behaviors were driving that direct infrastructure bill. Without measuring that baseline first, any cost-cutting initiative would have been shooting in the dark.
Step 2: Quick Wins That Took Under 60 Hours
Once we understood our COGS drivers, we hunted for low-hanging fruit that required minimal friction. Our billing ingest pipeline was our largest single cost driver, consuming immense amounts of cloud compute.
Instead of undertaking a massive, disruptive refactor, we looked at caching mechanisms and data retrieval patterns. By identifying redundant queries and inefficient data parsing, we realized we could eliminate massive blocks of redundant compute. In less than 60 hours of total engineering investigation, prototyping, and deployment, we cut annualized costs by $1,098,000.
The best part wasn't just the dollar figure. It was how little disruption it caused to our product roadmap. Our engineers didn't have to stop shipping features; they just needed to see where their code was burning cash. Automation tools had previously flagged none of these architectural nuances because automated scripts lack architectural context.
Step 3: Negotiating Vendor Contracts Like Snowflake
A $1.1M reduction was incredible, but we weren't done. The next logical step in any mature FinOps strategy is examining vendor contracts.
Most teams start blindly by locking into generic multi-year Reserved Instances or AWS Enterprise Discount Programs. Because we had granular cost visibility via AnyCost, we didn't have to guess where our money was going. We discovered that Snowflake accounted for a staggering 63% of our total cloud spend.
Armed with accurate forecasts of our future data ingestion and query growth, we approached Snowflake to negotiate a committed use agreement. Because we could prove exactly how much capacity our multi-tenant architecture required, we secured a significantly better rate. That single negotiation yielded $145,000 in annual savings with zero engineering hours invested.
Step 4: Engineering Culture and How Much Does an LLM Cost
To this point, we had protected our core engineering teams from deep cost reviews so they could ship major product updates like our AI Hub and Budgets features. But true, sustainable savings require engineering engagement.
As teams increasingly adopt generative AI features, engineering leaders frequently ask: how much does an llm cost at scale? The answer depends entirely on token volume, prompt sizing, model selection (such as Anthropic Claude or OpenAI GPT models), and whether inference is cached or hitting raw APIs. Unmonitored LLM queries and chat endpoints can quietly bleed margins just as fast as unoptimized cloud databases or serverless functions.
In our own infrastructure, serving data for a surge of new engineering users was getting increasingly expensive. As a serverless shop, our default instinct was to spin up resources on demand without restriction. One of our engineers realized we could predict the exact resource size a user request actually needed rather than applying a blanket serverless configuration.
By prototyping and rolling out intelligent right-sizing for serverless requests across our ecosystem, costs dropped by another $144,000 annually. More importantly, it fostered a permanent culture shift. When engineers see the direct cost consequences of their architecture, they optimize instinctively. Automated scripts could never replicate that human judgment.
The Final Tally: 54 Hours and a 24x ROI
When we tallied up the results at the end of the year, the numbers spoke for themselves:
- Total Annualized Savings: $1.7 million.
- Engineering-Led Savings: $1.498 million (88% of the total).
- Vendor Negotiation Savings: $145,000.
- Total Staff Effort: Approximately 54 hours of engineering time.
- Gross Margin Impact: A 10% overall improvement.
If CloudZero were an external customer paying for our platform and investing $17,000 in engineering hours, the net project profit would still sit around $1.6 million—representing a staggering 24x return on investment.
Engineering-led cost optimization isn't about slapping a temporary band-aid on your cloud bill. It is about connecting technical decisions to business reality, empowering developers with clear context, and building a sustainable engineering culture where cost efficiency is just another mark of clean code.