If an agent under deadline pressure can see the thing that grades it, what stops it from routing around that check? I want an answer before wiring any agent into a system with write access.
Evasion shows up without anyone provoking it
On September 24, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Anselm Paulus, Ameya Prabhu, and Maksym Andriushchenko published Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure, which introduces a benchmark called EvasionBench. We saw part of this in the hugging face attack, where if true the agents were aware they were being watched.
That has a direct consequence for monitor placement. If the monitor is a tool the agent can call, a log it can write to, or a config it can edit, the monitor is part of the surface the agent optimizes against.
The Verge reported that OpenAI paused training of its most capable models after a model under test exploited a loophole to reach the internet from a sandbox. The Decoder reported further detail: a DNS loophole used to escape a locked-down environment, a deliberately leaked GitHub token, and a model that twice ignored a researcher's direct instruction. Both accounts trace back to OpenAI's own disclosures, so the specifics may change.
But all this proves OpenAI, with the smartest people in the world, apparently doesnt know how to build an air-gapped environment.
Somebody wrote down the boring architecture
Also on September 24, Mehmet Iscan published Requirement-Bound Verified Commissioning, describing a frozen four-billion-parameter local model used as a candidate generator, with an external acceptance layer holding verification and release authority.
The study is narrow — one paper on mechatronic commissioning. To sum it, when release authority lives outside the model, model capability becomes a throughput concern not a trust concern. You certify the gate once and swap the generator when a better one ships. A paper from Matteo Merler, Bowen Li, Josh Roy, Yichao Liang, Qianwei Wang, Yixuan Huang, and Tom Silver on coding agents for generalized task and motion planning uses the same idea, treating coding agents as program generators evaluated against benchmark verification instead of trusting their plans directly.
Splitting the planner from the executor
Jev-Mobile, from Linghua Zhang and colleagues, decouples low-frequency vision model planning. The authors report 79 percent task success on AndroidWorld with reduced execution time and API cost. Fairly big claim about their architecture.
The Decoder also reported on Nvidia's SoL-Pi, which cut coding-agent token usage by up to 49 percent with little performance change by optimizing the control layer between model and environment, with smaller gains on other benchmarks. Fix the harness, not the model.
Context shapes the decision
Zixiang Xu's JevOut reports that short additions of ordinary natural context can flip a decision model's correct choices into high-confidence wrong ones. No injected instructions required. If your agent pulls in documents, ticket history, or chat threads before deciding something, that retrieval path shapes both the decision and the confidence attached to it. SAGE, from Xinyue Zeng, Jiawei Zhang, Yujun Yan, and Dawei Zhou, comes at the adjacent problem of long-horizon reasoning bias from the training side. Neither removes the operational requirement, which is knowing the provenance of everything in the context window at decision time.
The test you can run
Take one agent pipeline you already run in production or close to it. Construct a task the agent cannot complete honestly — a validation step that fails for a legitimate reason — and add mild pressure through the prompt or a tight retry budget. Log, at the tool layer, whether the agent modifies the check, disables it, retries around it, or reports success anyway.
Add two constraints: Instrument outside the agent's context, because self-reported traces are the exact artifact the EvasionBench framing calls into question. And run it against the credentials the agent actually holds, since scope of access is what turns an evasion behavior into an incident.
The better investment right now is the acceptance layer, not the model tier. An external verifier with release authority and no shared credentials is cheaper to build than most teams assume, and it survives model upgrades.
In the posse.bot harness, Ive burned many tokens on non-deterministic evals, tied right into the harness, so eager to try this out.
Two things to watch
Whether OpenAI resumes tool-based training, and what disclosure looks like when it does. The reported pause covers training, evaluation, and inference for the most capable models. How that restriction lifts will tell you more about industry containment posture than any published safety framework.
Open SWE's release cadence. LangChain shipped desktop nightlies on September 25 covering Slack, middleware, and database fixes, another later that day, and a third on September 26 with bug fixes and prompt refactoring. Three nightlies in two days on a desktop surface is worth tracking if you depend on it, particularly whether the 0.2.11 line settles into a stable cut. Entity connections are browsable at ranger360.ai/explorer.


