
Sponsored By:
Thursday, October 1st
9:00 am ET
Cloud cost management matured around a predictable set of levers: rightsizing, reserved capacity, autoscaling, storage tiering. AI workloads break most of those assumptions. Spend is now driven by tokens, inference patterns, model choice, GPU utilization, and context window size, variables that don't map cleanly to the VM and container based cost models most engineering and FinOps teams already know how to manage.
Attendees should join this session because AI spend is becoming one of the largest, least understood line items in the cloud bill, and most teams don't yet have a way to explain where it's going, let alone forecast it. The problem this webinar solves is the gap between traditional cost visibility tools and the actual cost drivers of AI: teams can tell you their EC2 spend by team and service, but not their cost per inference by feature or customer.
Five reasons to attend: (1) understand exactly what makes AI cost tracking fundamentally different from cloud cost tracking, (2) learn the new cost drivers, tokens, model selection, retries, context length, that traditional tagging and showback models miss, (3) get practical engineering efficiency tactics like model routing, caching, and prompt optimization that reduce spend without hurting product quality, (4) walk away with an approach to attributing AI costs to teams and features when usage is bursty and uneven, and (5) build a shared language between engineering and finance before AI spend becomes a budget crisis instead of a planning input.
This session is built on real patterns emerging across engineering organizations right now, not a product pitch, and is aimed at engineering leaders, platform teams, and FinOps practitioners trying to get ahead of a cost curve that's moving faster than most budgeting cycles can keep up with.
Key Takeaways:
Why traditional cloud cost allocation (tagging, showback, chargeback by team or service) breaks down when applied to AI workloads
The new cost drivers unique to AI spend: tokens, inference vs. training cost, model selection, context length, and retry/fallback logic
How to attribute AI costs to teams, features, or customers even when usage is bursty and inconsistent
Concrete engineering efficiency levers (caching, prompt optimization, model routing, right sizing model choice) that cut cost without cutting quality
How to build forecasting and budgeting practices that hold up when AI usage spikes overnight
Register Below:
We'll send you an email confirmation
.png?width=189&height=189&name=broganpatrick-modified%20(1).png)