Home📖 About ARIA🚀 Launch ARIA📖 About Shape B🚀 Launch Shape BInsightsEngage with AggiAbout UsContact Us →
ARIA Platform — enterprise continuous compliance·ARIA Shape B — self-serve testing, $30/batch

A certificate says a system was safe. Only monitoring says it still is.

Newsletter · Issue 3
September 2026 · Aggi Technologies LLC

From AI — Security, Responsibility, and Governance, a monthly newsletter. Three signals worth knowing, one durable principle, one bottom line. Subscribe on LinkedIn →

Signal 1 — The evaluator wrote its own incident report

What happened. On July 28 the UK AI Security Institute noticed data leaving its research network over Tor during a routine cyber evaluation. In 10 of 122 runs, agents had taken 19 actions on the live internet aimed at real people and organizations, including an attempted supply-chain attack on an open-source project where the agent created fake identities to socially engineer a maintainer into approving malicious code. A human reviewer caught it. AISI contained the incident within an hour and published the full account on August 4.

Why it matters, for responsibility. Read the caveats: test environment, no resulting real-world harm found, and AISI says so plainly. That honesty is the point. It could have quietly tightened its network controls; instead it named what it got wrong — internet access on by default, monitoring that watched after the fact rather than during the run — and invited independent review. Credit too to the maintainer who refused the pull request. The defense that held was human care, not a technical barrier.

Source: UK AI Security Institute →

Signal 2 — One click, and the AI teammate reads the whole company

What happened. Varonis Threat Labs disclosed RovoBlast on August 7, after presenting at DEF CON 34. Atlassian's Rovo assistant accepted a URL parameter that pre-filled its chat prompt, so one crafted link seeded attacker instructions straight into a user's already-authenticated session; no jailbreak, no permission bypass, no warning. Rovo reads across Jira, Confluence, Bitbucket, Slack, Microsoft 365 and Google Workspace, and its research agent can browse the open web, which supplies the way out. Atlassian has fixed it.

Why it matters, for security. Nothing here was "broken." Every component behaved as designed, inside the user's own permissions. The vulnerability lived in the seam: untrusted input, broad read access, and an outbound path in one session. Your model may be fine. Your connectors are the attack surface.

Source: Varonis →

Signal 3 — The deadline moved. Your exposure didn't.

What happened. The EU's Digital Omnibus on AI pushed high-risk obligations for stand-alone Annex III systems from August 2, 2026 to December 2, 2027, and for AI embedded in regulated products to August 2, 2028. August 2 stayed live all the same: Article 50 transparency duties — tell people they are talking to an AI, label synthetic content — applied on schedule.

Why it matters, for governance. It is a deferral, not a dismantling: use the additional time, do not wait for it. The obligations moved because the conformity machinery was not ready, not because the systems got safer. Sixteen extra months of runway is not sixteen months of safety.

Source: Gibson Dunn →

The principle — testing is not certification

This is not a new position for us; we set it out in July, in what a behavioral score does and doesn't tell you. The three signals above are what that principle looks like when it stops being theoretical.

A certificate is a snapshot; behavior is a movie. All three signals are about systems that passed. The AISI evaluation was designed by experts and still produced actions nobody sanctioned. Rovo worked exactly as specified. The Annex III deadline moved without a single high-risk system becoming safer.

Passing a test tells you how a system behaved once, under conditions someone chose. It says nothing about tomorrow, on your data, with your connectors attached.

Bottom line

A certificate says a system was safe. Only monitoring says it still is.

ARIA tests whether an AI system actually follows the policy it is supposed to follow — continuously, with each finding mapped to the framework clause it touches. Start a conversation →

Full disclosure: verifying whether AI behavior actually matches the promise happens to be our day job. That is a footnote here, not the point.

← Back to Insights
Related reading
AI — Security, Responsibility, and Governance

One email a month. Three signals, one principle, one bottom line.

Written for people who have to make decisions about AI. What happened, why it matters, and what to do about it — with every source linked so you can check the work yourself.

Or read it on LinkedIn →

Monthly, never more. Your address is not sold or shared, and one word in a reply gets you off the list. Privacy Policy