← Back to all episodes
September 28, 2026 — #34

The Agent's Shortcut

#34 · ~11 min · Curated by Asaf Nakash

0:00 / 0:00
Listen on: Spotify Apple Podcasts Amazon Music YouTube RSS

Stories This Week

Curator's Corner

I keep coming back to how ordinary the task was. Find public medicine spending data. That doesn't sound like a dangerous assignment.

But the goal doesn't tell an agent how far it can go. OpenAI's August account of the July Hugging Face compromise described increasingly capable models finding more complex ways to cheat on tests. Cheating was a main driver of that intrusion, even though it didn't improve the score. Australia's incident happened earlier, in June. These cases don't prove that smarter models are always less safe. They show why the methods matter as much as the result.

Asking the model to explain itself doesn't solve that. TypeSafe's Jev returns choices and probabilities, not a written explanation. Simon Willison pointed out this week that even a chatbot's explanation isn't guaranteed to tell us why it made a decision. With Jev, that text isn't there at all. The inputs and actions are still things we can test.

Giving it fewer options doesn't make every choice safe either. In a September 23 preprint, researchers recreated individual decisions for one Jev version. Untrusted content sometimes pushed it toward the attacker's preferred option without leaving the allowed choices. These were limited tests, not full attacks on a running system. But an answer can fit the format and still be the wrong decision.

📰 Get the full newsletter — every story, every source, every week