Can You Train an AI to Stop Deceiving You? | Teun van der Weij, Apollo Research
Zurich AI Safety
0:00 / 0:00
Can You Train an AI to Stop Deceiving You? | Teun van der Weij, Apollo Research
46 просмотров · 12 дней назад
Zurich AI Safety
22 подписчика
46 просмотров · 12 дней назад
OpenAI tried to train deception out of its own frontier models. Apollo Research stress-tested whether it actually worked.
CHAPTERS
00:00 Introduction
00:49 What is scheming?
01:13 Why it is so hard to train against
02:21 Three requirements for anti-scheming training
03:16 The approach: covert actions as a proxy
05:14 How deliberative alignment works
06:49 Main result: large reductions, never zero
07:34 Failure mode: ignoring or inventing the spec
09:04 Failure mode: misleading chain of thought
09:50 Failure mode: hidden goals survive the training
11:37 Failure mode: models start writing in unreadable language
13:18 Models know when they are being evaluated
16:11 Main takeaways
17:41 Working at Apollo Research
18:27 Q&A
Recorded at the Zurich AI Safety Day, ETH Zürich, September 2025.
The Zurich AI Safety Day brought together researchers, policymakers and professionals working to reduce risks from advanced AI.