EP24: AI is lying to us
Human In The Loop by VallySeed
0:00 / 0:00
EP24: AI is lying to us
122 просмотра · 7 дней назад
Human In The Loop by VallySeed
42 подписчика
122 просмотра · 7 дней назад
A ChatGPT inventor built an AI model that cannot write a sentence. TypeSafe AI says that is the point.
Jev gives up free-form language and returns typed, probabilistic decisions that software can use directly. The company says it is faster, cheaper, and better calibrated than language models on the workflows it tested. The catch is simple: a valid type can still contain the wrong decision.
Oscar and Matt ask whether production AI needs fewer chatbots and more constrained models built for classification, routing, scoring, and verification.
Signal or Noise:
Jev and the case for AI models that do not generate prose: SIGNAL pending independent testing
OpenAI models preserving instructions to hide mistakes across context windows: SIGNAL
Gemini reaching three real companies during a cyber evaluation: SIGNAL
The lawsuit over an alleged coordinated AI slowdown: SIGNAL for governance, not proof of collusion
Anthropic's Life Sciences Verification Program: SIGNAL with data and access constraints
Then Stack Check tests two tools:
1. Archify for validated technical diagrams and PR architecture diffs
2. Iteris for turning selected tickets into reviewed pull requests. Oscar reports it was used on two client projects. He reports that the .NET and Microsoft SQL to Next.js and Postgres migration completed successfully. On the fitness app, he reports that Iteris handled 13 medium-priority tickets on the first day and 10 resulting pull requests were merged that day. These outcomes were not independently audited.
The closing disagreement:
Oscar bets small specialist models will make big general-purpose LLMs obsolete.
Matt argues most AI startups are features the model providers have not shipped yet. His challenge: are you building a company, or a feature with a burn rate?
Human In the Loop is a weekly podcast with Oscar Gallo and Matt Wozniak. We focus on decisions you can use when you build, buy, secure, and operate AI.
Timestamps:
0:00 - Cold Open
0:25 - Welcome to Human in the Loop
0:51 - Signal or Noise intro
1:04 - Typesafe launches JEV — small specialized models
11:15 - OpenAI models caught lying in compaction summaries
30:07 - Gemini hacked three real companies during security test
39:27 - Subscribers sue over alleged coordinated AI slowdown
51:12 - Anthropic opens a wet biology lab
1:03:01 - Stack Check intro
1:03:24 - Archify — architecture diagram generator
1:13:32 - Iteris — ticket-to-PR meta harness
1:27:41 - Unpopular Opinions
1:27:47 - Matt: Most AI startups are features not yet built by model labs
1:42:07 - Oscar: Big general LLMs will die
1:48:15 - Outro
Sources:
https://typesafe.ai/blog/introducing-...
https://techcrunch.com/2026/09/18/a-n...
https://openai.com/index/model-misali...
https://alignment.openai.com/misalign...
https://techcrunch.com/2026/09/17/ope...
https://www.bbc.com/news/articles/c60...
https://www.reuters.com/business/gemi...
https://apnews.com/article/antitrust-...
https://www.anthropic.com/news/life-s...
https://github.com/tt-a1i/archify
https://github.com/tt-a1i/archify/blo...
https://github.com/Oscar-Codes-Life/I...
Subscribe for a new episode each week. Tell us which Stack Check candidate you would keep.