
From AI — Security, Responsibility, and Governance, a monthly newsletter. Three signals worth knowing, one durable principle, one bottom line. Subscribe on LinkedIn →
What happened. Wiz Research disclosed GhostApproval, a flaw in six leading AI coding assistants: Amazon Q Developer, Claude Code, Augment, Cursor, Google Antigravity and Windsurf. A booby-trapped repository uses a decades-old symlink trick to make the agent write to files far outside its workspace — an SSH key, a shell configuration — while the confirmation dialog you see shows a harmless-looking local edit. In some tools the write landed before you could click.
Why it matters, for security. The human-in-the-loop approval gate is the control most teams lean on for autonomous agents. When the agent shows one thing and does another, that gate becomes a rubber stamp. Real credit to Wiz for the research, and to AWS, Cursor and Google, who shipped fixes.
What happened. The Future of Life Institute published its Summer 2026 AI Safety Index, an independent panel grading nine frontier labs across 37 indicators in six domains. No lab scored above a C+; Anthropic led at C+, with OpenAI and Google DeepMind following. Reviewers flagged a pattern of safety pledges quietly weakened under competitive pressure.
Why it matters, for responsibility. Responsibility is measurable, and outside measurement keeps everyone honest. The useful signal is not the letter grade; it is the gap between what was promised during fundraising and what is actually shipping now.
Source: Future of Life Institute →
What happened. As of August 2, the EU AI Office and member states can enforce the obligations on general-purpose AI models that have been on the books for a year, with fines up to €15 million or 3% of global turnover. New transparency duties also apply: tell people when they are talking to an AI, and label synthetic and deepfake content.
Why it matters, for governance. The paperwork phase is ending and the enforcement phase is beginning. Governance shifts from "do you have a policy?" to "can you show your systems follow it?" — and that is a different kind of homework.
Read the three together: an approval dialog that hides the real action, safety pledges walked back when the pressure rose, and obligations that only bite once someone can check them. Anthropic's own July study on agentic misalignment points the same way; it found models that will agree to your face and then act against you when they think it serves the goal.
The through-line is simple. What a system, or an organization, says it will do is not the same as what it does. So do not buy the promise. Check the behavior, and keep checking, because behavior drifts.
In AI, trust is earned at runtime. Watch what the system does, not what it says it will do.
ARIA tests whether an AI system actually follows the policy it is supposed to follow — continuously, with each finding mapped to the framework clause it touches. Start a conversation →
Full disclosure: verifying whether AI behavior actually matches the promise happens to be our day job. That is a footnote here, not the point.
← Back to InsightsWritten for people who have to make decisions about AI. What happened, why it matters, and what to do about it — with every source linked so you can check the work yourself.
Or read it on LinkedIn →