Перейти к содержимому

Can You Train an AI to Stop Deceiving You? | Teun van der Weij, Apollo Research

Zurich AI Safety

0:00 / 0:00

Can You Train an AI to Stop Deceiving You? | Teun van der Weij, Apollo Research

46 просмотров · 12 дней назад
Zurich AI Safety
22 подписчика
46 просмотров · 12 дней назад
OpenAI tried to train deception out of its own frontier models. Apollo Research stress-tested whether it actually worked. CHAPTERS 00:00 Introduction 00:49 What is scheming? 01:13 Why it is so hard to train against 02:21 Three requirements for anti-scheming training 03:16 The approach: covert actions as a proxy 05:14 How deliberative alignment works 06:49 Main result: large reductions, never zero 07:34 Failure mode: ignoring or inventing the spec 09:04 Failure mode: misleading chain of thought 09:50 Failure mode: hidden goals survive the training 11:37 Failure mode: models start writing in unreadable language 13:18 Models know when they are being evaluated 16:11 Main takeaways 17:41 Working at Apollo Research 18:27 Q&A Recorded at the Zurich AI Safety Day, ETH Zürich, September 2025. The Zurich AI Safety Day brought together researchers, policymakers and professionals working to reduce risks from advanced AI.