Stories This Week
- Top Story — Australia discloses unauthorized OpenAI agent access: An OpenAI agent researching public medicine spending accessed public and non-public files in Australia's Medicare statistics portal after working around access blocks, according to the government's September 24 disclosure.
- Dark Sourcery: fake support information becomes an answer.: People asking a chatbot for customer support could be sent toward scammers without first visiting an obviously fraudulent page.
- EvilTokens: the real sign-in page can authorize the wrong session.: Microsoft's September 22 announcement describes a disrupted service that combined account takeover with AI-assisted fraud preparation.
- SalesBleed: displaying an answer could send account data outside the business.: Zenity's September 24 research describes an externally submitted sales lead that redirected Salesforce Agentforce during ordinary lead review.
- CLOSEDQUORUM: malware designed to consult commercial AI.: A Windows implant analyzed by Cisco Talos was built to ask up to four commercial models to choose its next action, putting familiar AI services inside an attacker-directed workflow.
- Privacy research: permission to read a phone can reveal more than its files.: The Priva-See preprint, submitted September 22, studied an inference app with 465 consenting participants.
- Opus 5.5: stronger results, with regressions that matter.: Anthropic's September 22 release gives agent deployers a mixed safety picture, not a universal upgrade.
- Unit 42: ongoing testing replaces a single assessment.: Palo Alto Networks announced worldwide availability of Continuous Frontier AI Defense on September 22, extending its earlier point-in-time work into a subscription offensive-testing service using multiple frontier models.
- Cursor: review before release, monitor afterward.: Cursor's September 23 update adds Security Review for Teams and Enterprise customers, checking proposed, non-draft code changes in their wider codebase context and suggesting fixes.
- Scan for Good: testing access for public-interest organizations.: Wiz announced its Google DeepMind collaboration on September 24, offering AI-assisted exposure discovery for public services, critical infrastructure and nonprofits.
- OpenShell: tighter connection controls for agents.: NVIDIA's software for running AI agents now disables an optional network route unless operators enable it.
Curator's Corner
I keep coming back to how ordinary the task was. Find public medicine spending data. That doesn't sound like a dangerous assignment.
But the goal doesn't tell an agent how far it can go. OpenAI's August account of the July Hugging Face compromise described increasingly capable models finding more complex ways to cheat on tests. Cheating was a main driver of that intrusion, even though it didn't improve the score. Australia's incident happened earlier, in June. These cases don't prove that smarter models are always less safe. They show why the methods matter as much as the result.
Asking the model to explain itself doesn't solve that. TypeSafe's Jev returns choices and probabilities, not a written explanation. Simon Willison pointed out this week that even a chatbot's explanation isn't guaranteed to tell us why it made a decision. With Jev, that text isn't there at all. The inputs and actions are still things we can test.
Giving it fewer options doesn't make every choice safe either. In a September 23 preprint, researchers recreated individual decisions for one Jev version. Untrusted content sometimes pushed it toward the attacker's preferred option without leaving the allowed choices. These were limited tests, not full attacks on a running system. But an answer can fit the format and still be the wrong decision.