Cost optimisation that starts with a measurement.
Inference cost per million tokens, model routing, caching, right-sizing and reservations, applied to AI workloads and to the cloud estate around them. Fixed-fee assessment; savings reported against a baseline you agreed.
The problem
AI spend is unusual: it scales with usage, it hides inside a few line items, and the levers are technical. Switching a model, routing simple requests to a smaller one, caching repeated prompts or moving inference to different silicon can each move the bill by double-digit percentages. Nobody in finance can pull those levers, and most engineering teams have not been asked to.
We measure first, so that every change is reported against a baseline rather than a feeling. Then we work through the levers in order of return, with your team, inside your accounts.
What we deliver
- Cost baseline: per workload, per model, per million tokens, with the top ten drivers
- Model routing and fallback design: the right model per request class
- Prompt and response caching, batching, context trimming
- Inference placement: Bedrock, Inferentia, Trainium, GPU, by workload
- Cloud estate review: right-sizing, reservations and savings plans, idle resources
- Cost dashboard and monthly governance cadence handed to your team
How we work
Baseline
Two weeks of measurement. Where the money goes, by workload and by driver, in a report your CFO can read.
Levers
Each lever quantified: expected saving, effort, risk. You choose the order.
Apply
Changes made in your tenancy with quality checked against your evaluation suite, so savings never come at the cost of answers.
Govern
Dashboard, budget alerts and a monthly review handed to your team.
What backs it
Common questions
How do you charge?
The assessment is a fixed fee. Implementation is fixed scope per lever. We do not take a percentage of savings, because that rewards the wrong behaviour.
Will quality suffer?
Not without you knowing. Every change is checked against your evaluation suite before it is kept, and the report shows the quality numbers next to the cost numbers.
Is this only for AI workloads?
No. Because we also migrate general workloads to AWS, the estate review covers compute, storage and data services around the AI systems.
Related services
Talk to an engineer about this
Thirty minutes, no slides. Bring the workload and we will tell you what we would do and what it would cost.