Перейти к содержимому

The New AI Stack: Models Are Just One Layer

AI deepdive

0:00 / 0:00

The New AI Stack: Models Are Just One Layer

19 просмотров · 1 месяц назад
AI deepdive
10 подписчиков
19 просмотров · 1 месяц назад
The New AI Stack: Models Are Just One Layer A modern AI product is not just a prompt connected to a model. The model is the engine — but production reliability comes from the stack around it: identity, routing, RAG, tools, orchestration, memory, guardrails, evals, observability, deployment, and human oversight. This is the bigger map behind the recent AI Deep agent/security arc. Your agent needs a computer. Your coding agent needs a codebase map. Prompt injection needs architecture, not prompt magic. VLM supply-chain risk means downloaded models are dependencies. All of those are symptoms of the same shift: production AI needs architecture, not vibes. In this video, we break down the modern AI stack and why bigger models do not remove the need for system design. Models reason, but systems decide what data they see, what tools they can call, what permissions they have, how failures are caught, and how behavior improves over time. Topics covered: Why the simple prompt-to-model demo mental model is incomplete Model routing and LLM gateways Identity, permissions, and enterprise access control RAG as a data pipeline, not magic memory Tool calling and validated execution Agent orchestration with loops, graphs, checkpoints, and runtime state Memory and state management Guardrails, prompt injection, sandboxing, and approval gates Evals, traces, observability, cost, latency, and quality metrics Chapters: 0:00 The Demo Model Is Wrong 1:01 The Model Is Still the Engine 2:09 Identity Before Intelligence 3:11 RAG Grounds the Model 4:18 Tools Turn Answers Into Actions 5:29 Orchestration Is the Runtime 6:29 Memory Creates Continuity 7:29 Guardrails Are the Brakes 8:29 Evals Make It Improve 9:26 The Product Is the Composition 10:26 What Layer Does the Work? Sources and further reading: Patrick Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” arXiv:2005.11401, submitted May 22, 2020; revised April 12, 2021: https://arxiv.org/abs/2005.11401 Shunyu Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models,” arXiv:2210.03629, submitted October 6, 2022; revised March 10, 2023: https://arxiv.org/abs/2210.03629 Timo Schick et al., “Toolformer: Language Models Can Teach Themselves to Use Tools,” arXiv:2302.04761, submitted February 9, 2023: https://arxiv.org/abs/2302.04761 Anthropic, “Introducing the Model Context Protocol,” November 25, 2024: https://www.anthropic.com/news/model-conte... Model Context Protocol documentation, “What is the Model Context Protocol (MCP)?”: https://modelcontextprotocol.io/docs/getti... Google, “Introducing Gemini 2.0: our new AI model for the agentic era,” December 11, 2024: https://blog.google/innovation-and-ai/mode... Google, “Gemini 2.5: Our most intelligent AI model,” March 25, 2025: https://blog.google/innovation-and-ai/mode... Anthropic, “Introducing Claude 4,” May 22, 2025: https://www.anthropic.com/news/claude-4 Anthropic, “Introducing Claude Sonnet 4.5,” September 29, 2025: https://www.anthropic.com/news/claude-sonn... Google, “A new era of intelligence with Gemini 3,” November 18, 2025: https://blog.google/products-and-platforms... Google Cloud, “Bringing Gemini 3 to Enterprise,” November 18, 2025: https://cloud.google.com/blog/products/ai-... LangChain documentation, “LangGraph overview.”: https://docs.langchain.com/oss/python/lang... LangChain, “LangGraph: Agent Orchestration Framework for Reliable AI Agents.”: https://www.langchain.com/langgraph LlamaIndex, “Agentic RAG With LlamaIndex: Architecture Guide,” January 30, 2024: https://www.llamaindex.ai/blog/agentic-rag... Pinecone, “Retrieval-Augmented Generation (RAG),” June 12, 2025: https://www.pinecone.io/learn/retrieval-au... AWS, “Amazon Bedrock Agents.”: https://aws.amazon.com/bedrock/agents/ Microsoft Learn, “What is Microsoft Foundry Agent Service?”: https://learn.microsoft.com/en-us/azure/ai... Snowflake documentation, “Cortex Agents.”: https://docs.snowflake.com/en/user-guide/s... LangSmith documentation, “LangSmith Observability.”: https://docs.smith.langchain.com/ LiteLLM documentation, “Getting Started.”: https://docs.litellm.ai/docs/ vLLM documentation: https://docs.vllm.ai/en/latest/ Weaviate documentation: https://weaviate.io/developers/weaviate Chroma documentation: https://docs.trychroma.com/ Subscribe for technical AI explainers on LLM infrastructure, agents, open models, and production AI systems. #AIDeepDive #LLM #AIAgents #RAG #LLMOps #ProductionAI