Token Management, Spend, and How to Actually Bring AI Costs Down
AnswerRocket
0:00 / 0:00
Token Management, Spend, and How to Actually Bring AI Costs Down
79 просмотров · 11 дней назад
AnswerRocket
374 подписчика
79 просмотров · 11 дней назад
Managing AI costs has become increasingly difficult. Uber’s CTO made headlines when he shared the company had burned through its entire 2026 AI budget in just 4 months. That's the problem this episode tackles head on. Token prices are falling for a given level of intelligence, but most companies' actual AI spend keeps climbing, because the bar for "good enough" intelligence keeps moving too. Nobody has a clean way to forecast it yet.
In this episode of AI, Actually, Jim Johnson sits down with Shanti Greene, Stew Chisam, and first-time guest Jake Barger to break down what's actually driving enterprise AI costs and what to do about it. The conversation moves from the economics of model pricing into a real case study: a 33x cost reduction on a production agent, achieved in about two weeks.
You'll get insight into:
• Why token prices dropping doesn't mean your AI bill is dropping
• How model routing, caching, and context management quietly drive cost
• The step-by-step process behind a 33x cost reduction on a live enterprise agent
• When open weight and self-hosted models actually make sense, and when they don't
• Practical takeaways for building a repeatable cost optimization process
Follow the Gang:
Jim Johnson, Managing Partner, AnswerRocket - / jim-johnson-bb82451
Shanti Greene, Head of Data Science and AI Innovation, AnswerRocket - linkedin.com/in/shantigreene/
Stew Chisam, Operating Partner, StellarIQ - / stewart-chisam-7242543
Jake Barger, Director of AI & Machine Learning, AnswerRocket | linkedin.com/in/jacob-barger-30b1ab1b9/
Chapters
00:00 Introduction and panel overview
02:44 Token Economics: Why Costs Stay High
07:34 Model Routing, Caching & Task Fit
12:22 Continuous Optimization for AI Agents
16:51 The 33x Cost-Reduction Case Study
23:36 Context Management & Smart Escalation
29:28 Open-Weight and Self-Hosted Models
37:22 Enterprise Reliability and Deployment Tradeoffs
44:26 Closing Takeaways: Managing AI Spend
#TokenManagement #AICostOptimization #LLMTokenSpend #ModelRouting #AgenticAICosts #OpenWeightModels #SelfHostedLLM #KVCacheOptimization #EnterpriseAIROI #AIInfrastructure
____________________________________________________________________
SUBSCRIBE TO OUR PODCAST
AI, Actually
Tired of the AI hype? So are we. Welcome to AI, Actually: the podcast that cuts through the noise and gets real about how artificial intelligence can work for your business. In each episode, our resident AI and business transformation experts–along with occasional industry guests–hold a candid, jargon-free conversation on what it takes to get actual value from AI.
Join us as we tackle topics like: the real difference between the latest LLM models, why generic AI can't make sense of your messy company data, how to get your GenAI use case off the ground, and what the rise of AI agents means for your business. This is your practical playbook for putting AI to work. No PhD required.
AI, Actually is produced by AnswerRocket. Since 2013, our enterprise AI solutions have helped Fortune 500 companies achieve measurable results through their AI transformations. This podcast is where we share what we’ve learned.
____________________________________________________________________
LEARN MORE ABOUT ANSWERROCKET
https://answerrocket.com/
FOLLOW US ON SOCIAL MEDIA
Facebook: / answerrocket
Instagram: / answerrocket
LinkedIn: / answerrocket
Twitter: / answerrocket