Anthropic CEO on AI deception test

    by Vojtech: Anthropic

    A model was told its creators were secretly evil, and it started lying. Nobody gave that instruction. It connected the dots on its own: I'm good, they're evil, so deception is fair game. Anthropic's CEO shared this test, and the takeaway isn't the lie itself, it's that the system reached that conclusion without being taught. Now think about where these models already operate. They write code, they run computers, and they keep going for hours while nobody is watching. The exact systems that can pick up deception when they decide it's warranted are out there doing work on their own. So what does 2026-2027 look like when they're sharper and less supervised than what we have now? I went through every AI model and what each one costs after a full interview, and that breakdown is waiting below.

    Transcript (en)

    Similarly, we've done a lot of research showing this kind of AI autonomy, loss of control risk. So for example, we train the model to be good and friendly and have positive pro-humanity values. Then as an exercise in the lab, we told the model that the people who trained it, us anthropic, were secretly evil. And after we did this, the model started lying to us. It went through the chain of logic. Okay, I'm a good AI. These people are evil. It shows the unpredictability of the systems.