Перейти к содержимому

Engineering Coding Agent Harnesses for Data Workflows | Snowflake × Future AGI

Future AGI

0:00 / 0:00

Engineering Coding Agent Harnesses for Data Workflows | Snowflake × Future AGI

336 просмотров · 2 недели назад
Future AGI
224 подписчика
336 просмотров · 2 недели назад
An agent harness for data workflows is the tool registry, context, and guardrails wrapped around a coding agent. In this session, Josh Reini (Snowflake) and Rishav (Future AGI) rebuild one live in Claude Code, cutting upfront context from 77,000 tokens to 8,000 while holding task accuracy. Data tasks break generic coding agents. A software-engineering harness is built to navigate repos, edit files and run tests. A data harness has to carry schemas, business definitions, and relationships between datasets — and pay for all of it in tokens on every turn. The four techniques below target that cost directly, demoed against data-eng-bench, Snowflake's open-source data engineering benchmark, on a dbt model with multi-currency sales dated to order placement rather than payment. CHAPTERS- 1- Why AI agents pass the demo and fail in production 2- What is intelligence efficiency: accuracy vs. token cost 3- The components of a coding agent harness 4- Three optimization buckets: context, tools, skills 5- Tool search: from 30+ loaded tools to 2 6- Lossless result compression: markdown to TSV 7- Result offloading: previews instead of 13,000 rows 8- Bundle dispatch: parallel tool calls 9- Snowflake CoCo and data-eng-bench results 10- Live demo: the baseline run in Claude Code 11- Live demo: the optimized run Audience Q&A WHAT'S COVERED- • Intelligence efficiency — more tasks, higher accuracy, lower cost, and why cost now matters more than latency for long-running async work • The components of a harness: model, prompt assembly, tool registry, context/memory/compaction • Tool search — replacing 30+ eagerly-loaded tools with two (search + invoke). Anthropic measured this dropping upfront context from 77K tokens to 8K • Lossless result compression — markdown to TSV, and hoisting repeated column values into a header instead of repeating them on every row • Result offloading — a 5-row preview plus a cache reference instead of dumping 13,000 rows into the context window • Bundle dispatch — parallel tool calls, and how savings compound once results are compressed together • Live demo: baseline vs. optimized run, with cost and correctness checked against benchmark ground truth both times • Why these gains compound as tool counts, data size, and session length grow QUESTIONS ANSWERED IN THIS SESSION Q1- Why does a coding agent ignore AGENTS.md on long tasks? As tokens accumulate, model attention degrades. The fix is reducing unnecessary tokens — tool metadata, uncompressed results — so the instructions you want attended to aren't competing with noise. Q2- Is a huge tool response a tool design problem or a context management problem? Tool design. The context blowup is the symptom. Fix it at the tool boundary with lossless compression or result offloading. Q3- How do you verify compaction kept the original constraints? Manage what enters the window in the first place, and test across models — context handling varies significantly between them. Q4- When a data task fails, do you change the model, the context, or the harness? There's no correct first move. Run an LLM judge across the traces to identify failure modes — incorrect tool calls, redundant subtasks — and the fix usually becomes obvious. Q5- How much schema and business context should you front-load? More than you'd think. This inverts the advice for tools and skills. Rich semantic context up front saves the agent expensive exploration turns. Q6- Should evals judge the trajectory or the final answer? Trajectory, unless you have benchmark ground truth. Ground truth lets you check the final answer; everything else needs trace-level evaluation. ABOUT THE DEMO Rishav also demos Future AGI's environment feature: point it at a repo and get a sandboxed environment where a failed production trace becomes thousands of simulation scenarios to regress new agent versions against. Future AGI is an open-source self-improving layer for agentic systems — simulate, evaluate, guard, observe and automatically improve your AI in one closed loop, so agents get more accurate over time instead of silently breaking in production. LINKS- Demo repo: https://github.com/joshreini1/agent-e... Future AGI open-source repo: https://github.com/future-agi/future-agi Try Future AGI: https://app.futureagi.com/ Join the Future AGI Discord:   / discord   #aiagents #dataengineering #snowflake #codingagents