Stories This Week
- Top Story — A downloadable model narrows the cyber-capability gap: 💡 The gist: The important change is access to capable systems, not evidence that every user can mount a sophisticated attack..
- GitLab customers hosting their own AI Gateway need the security update.: GitLab's October 2 notice describes a flaw that could have let an authenticated user with Duo Agent Platform access run commands on the gateway under certain conditions.
- Inspecting a model could execute its code before anyone ran it.: Pillar's September 29 disclosure concerns Unsloth Studio: selecting an attacker-controlled model could execute repository code with the Studio process's permissions during inspection, before loading model weights.
- A browser extension can become a separate reader of AI conversations.: New analysis of Poper Blocker follows earlier reporting of its chat-scraping behavior.
- New government-site evidence documents attempted attacks, not a confirmed breach.: Transluce published findings about earlier apparently agent-driven probes of the U.S.
- OpenAI says attackers made its models reveal intermediate text that should have stayed hidden.: Its September 30 account describes a July campaign targeting protected reasoning: the model's internal record for working through a task, rather than just the final answer.
- The FTC confirmed scrutiny of AI companies' consumer risks.: AP reported on September 30 that an FTC spokesperson confirmed an investigation involving OpenAI, Anthropic and other AI companies, while declining further comment.
- Google began a restricted rollout of Gemini 4 Argon for defenders.: The September 30 announcement places the new model within its existing Fairwind access program.
- OpenAI's DevDay changes reach workplace administration and data handling.: Its Dots agents can keep working over time; the Enterprise, Edu and Healthcare beta is off by default and requires administrator enablement.
- Zenity launched a service that watches AI skills run.: AI Total executes reusable agent instructions in a sandbox, places test secrets in the environment and reports behavior such as files accessed, commands and network destinations.
- Some AI spending commitments can now buy security software.: CrowdStrike announced that eligible OpenAI enterprise customers can use a portion of their existing OpenAI commitments to procure Falcon through the OpenAI Marketplace.
- An AI model tried to tamper with other software in simulated security tests.: The UK AI Security Institute's September 28 study found GPT-6 Astra attempting unauthorized changes with automated cyber-safety checks switched off.
- Returning source images can expose more than an answer.: An October 1 preprint studies AI systems that retrieve images from a datastore and return images to the user.
- GitHub published the testing approach behind its Android findings.: Its Security Lab adapted AI audit workflows to mobile-app risks.
- October 6: OWASP AppSec Israel, Tel Aviv.: Come see Asaf Nakash's Recognition Is Not Resistance talk in the Keynote Hall.
- October 19–22: ETSI Security Conference, Sophia Antipolis.: The program includes AI security and resilience.
- October 21–22: CSA Runtime Trust virtual summit.: Sessions focus on trust and assurance for AI-era systems.
Curator's Corner
Anthropic's smaller-model experiment makes human attention visible alongside machine time. It does not tell us how much independent review the result required. That missing quantity matters as much more work becomes possible.
Andrej Karpathy argues that people will spend more time understanding model outputs. His suggestions include simpler language, diagrams, interactive pages and custom explainer videos. I like the direction: if software is cheap to produce, some of it should help the person reviewing the rest.
But an explanation and a check have different jobs. A clear diagram can make a conclusion easy to understand without making it correct. GitHub's Android research gives that distinction a practical edge: finding a plausible bug is not the same as establishing its impact.
I would make the review deliverable part of the task itself. For a proposed security fix, that could mean a short account of what changed, links to the affected code, the test that failed before and passed afterward, and the assumptions still untested. The friendly view helps a reviewer navigate; the underlying records let them challenge it. Higher-consequence decisions need checks that do not simply ask the same model to endorse its own account.