Перейти к содержимому

Token Management, Spend, and How to Actually Bring AI Costs Down

AnswerRocket

0:00 / 0:00

Token Management, Spend, and How to Actually Bring AI Costs Down

79 просмотров · 11 дней назад
AnswerRocket
374 подписчика
79 просмотров · 11 дней назад
Managing AI costs has become increasingly difficult. Uber’s CTO made headlines when he shared the company had burned through its entire 2026 AI budget in just 4 months. That's the problem this episode tackles head on. Token prices are falling for a given level of intelligence, but most companies' actual AI spend keeps climbing, because the bar for "good enough" intelligence keeps moving too. Nobody has a clean way to forecast it yet. In this episode of AI, Actually, Jim Johnson sits down with Shanti Greene, Stew Chisam, and first-time guest Jake Barger to break down what's actually driving enterprise AI costs and what to do about it. The conversation moves from the economics of model pricing into a real case study: a 33x cost reduction on a production agent, achieved in about two weeks. You'll get insight into: • Why token prices dropping doesn't mean your AI bill is dropping • How model routing, caching, and context management quietly drive cost • The step-by-step process behind a 33x cost reduction on a live enterprise agent • When open weight and self-hosted models actually make sense, and when they don't • Practical takeaways for building a repeatable cost optimization process Follow the Gang: Jim Johnson, Managing Partner, AnswerRocket -   / jim-johnson-bb82451   Shanti Greene, Head of Data Science and AI Innovation, AnswerRocket - linkedin.com/in/shantigreene/ Stew Chisam, Operating Partner, StellarIQ -   / stewart-chisam-7242543   Jake Barger, Director of AI & Machine Learning, AnswerRocket | linkedin.com/in/jacob-barger-30b1ab1b9/ Chapters 00:00 Introduction and panel overview 02:44 Token Economics: Why Costs Stay High 07:34 Model Routing, Caching & Task Fit 12:22 Continuous Optimization for AI Agents 16:51 The 33x Cost-Reduction Case Study 23:36 Context Management & Smart Escalation 29:28 Open-Weight and Self-Hosted Models 37:22 Enterprise Reliability and Deployment Tradeoffs 44:26 Closing Takeaways: Managing AI Spend #TokenManagement #AICostOptimization #LLMTokenSpend #ModelRouting #AgenticAICosts #OpenWeightModels #SelfHostedLLM #KVCacheOptimization #EnterpriseAIROI #AIInfrastructure ____________________________________________________________________ SUBSCRIBE TO OUR PODCAST AI, Actually Tired of the AI hype? So are we. Welcome to AI, Actually: the podcast that cuts through the noise and gets real about how artificial intelligence can work for your business. In each episode, our resident AI and business transformation experts–along with occasional industry guests–hold a candid, jargon-free conversation on what it takes to get actual value from AI. Join us as we tackle topics like: the real difference between the latest LLM models, why generic AI can't make sense of your messy company data, how to get your GenAI use case off the ground, and what the rise of AI agents means for your business. This is your practical playbook for putting AI to work. No PhD required. AI, Actually is produced by AnswerRocket. Since 2013, our enterprise AI solutions have helped Fortune 500 companies achieve measurable results through their AI transformations. This podcast is where we share what we’ve learned. ____________________________________________________________________ LEARN MORE ABOUT ANSWERROCKET https://answerrocket.com/ FOLLOW US ON SOCIAL MEDIA Facebook:   / answerrocket   Instagram:   / answerrocket   LinkedIn:   / answerrocket   Twitter:   / answerrocket