Skip to main content
VeloraX
All insights
  • Cloud
  • Cost optimisation
  • FinOps

What actually drives your cloud bill, and what to fix first

Cloud cost problems are rarely caused by the thing teams assume. A practical order of investigation, from the changes that take an afternoon to the ones that need architecture.

· 7 min read · VeloraX Consulting

Densely cabled server racks in a data centre, with rows of green status indicators.

When a cloud bill becomes a board-level topic, the first reaction is usually to look for something expensive and switch it off. That instinct is understandable and it is occasionally correct, but it tends to produce a week of disruption and a saving that quietly reverses within a quarter. Cost work rewards the opposite approach: find out where the money is actually going, then fix things in the order of effort required.

You cannot reduce what you cannot attribute

The most common blocker is not technical. It is that nobody can say which team, product or customer is responsible for which portion of the bill. A consolidated invoice arrives, it is larger than last month, and the conversation stalls there because there is no way to connect the number to a decision anyone made.

Before optimising anything, get to a position where every significant resource carries tags for owner, environment and workload, and where those tags are enforced rather than requested. This is unglamorous work. It is also what converts an argument about the bill into a series of specific, answerable questions.

One caution: resist the urge to build an elaborate internal chargeback model at this stage. You need attribution good enough to direct attention. Precise cross-charging can wait until the obvious waste is gone.

The order that usually works

In our experience the investigation is most productive in roughly this sequence, because it moves from cheap and reversible to expensive and structural.

  1. 01Resources nobody owns. Detached storage volumes, idle load balancers, old snapshots, environments spun up for a demo two years ago. This costs an afternoon and often removes a surprising share of the bill.
  2. 02Non-production running out of hours. Development and staging environments frequently run every hour of the week to serve a team that works forty of them. Scheduled shutdown is a small change with an immediate effect.
  3. 03Over-provisioned compute. Instances sized during a launch and never revisited. Compare provisioned capacity against observed utilisation over a representative period, not against a peak everyone remembers.
  4. 04Storage tiering and retention. Data written once and never read again sitting on the fastest, most expensive tier, often with no retention policy at all.
  5. 05Data transfer. Cross-zone and cross-region traffic that comes from where components were placed rather than from anything the business needs. This one hides well and is worth looking for explicitly.
  6. 06Commitment coverage. Once usage is genuinely stable, reserved capacity and savings plans convert a predictable baseline into a lower rate. Do this last — committing to a workload you have not yet optimised locks in the waste.
  7. 07Architecture. Managed service choices, database sizing, caching strategy, and whether a workload needs to be running continuously at all. Real savings live here, and so does real risk, which is why it comes after the cheaper work.

The part teams skip

Almost every cost programme finds savings. Rather fewer keep them. Spend drifts back because the conditions that produced it are still in place: no environment expiry, no cost visible at the point where an engineer chooses an instance size, no review when a workload changes shape.

The durable fix is to make cost a property of the system rather than a periodic clean-up. Budget alerts routed to the team that owns the workload, cost impact noted in infrastructure change reviews, automatic expiry on non-production environments, and a short monthly review of the largest movements. None of this is difficult. It simply has to exist before the next deadline arrives and attention moves elsewhere.

A cost reduction that nobody owns is a cost deferral.

When cutting cost is the wrong goal

It is worth saying plainly: not every large cloud bill is a problem. If spend is growing in proportion to revenue, and the alternative is slower delivery or worse reliability, then the bill is doing its job. The question worth asking is not whether the number is large but whether it is understood and deliberate.

We have also seen optimisation programmes that succeeded on their own terms and hurt the business — capacity trimmed so close that a normal seasonal peak caused an outage, or a caching layer removed to save a modest monthly sum and paid for many times over in support load. Set a floor before you start: the reliability and performance you are not willing to trade, whatever the saving.

A reasonable first week

If you want somewhere concrete to start: enable detailed billing export and get a month of data into a tool you can query. Tag your top twenty resources by spend. List everything with no owner. Shut down non-production out of hours. Then look at the largest remaining line and ask the team that owns it what it does.

That week will not finish the job. It will, however, replace a vague sense that the bill is too high with a specific list of decisions — and that is the point at which cost work stops being anxious and starts being ordinary engineering.

Tell us what is not working

A first conversation costs nothing and commits you to nothing. Describe the problem in your own words and we will tell you honestly whether we are the right people for it.