Two weeks. One production AI endpoint. One report you keep, whether or not you continue.
Somewhere there is a question you cannot answer yet, and it is usually one of two.
Both questions have the same shape. You can answer with a policy — here is our documented process — or with evidence: here is what happened when we tested it, on this date, and here is what failed. The Behavioral Audit produces the second kind of answer.
Free is fine. Open-ended is not. Every engagement gets a start date and an end date before it begins, or it does not exist.
With an end date agreed before any work starts. Not "a few weeks," not "until we're done."
One real AI system that real users touch, not a sandbox. Scope fixed at the outset.
ARIA sends test cases to your endpoint and scores what comes back. Nothing sits in your request path. Nothing changes in production.
If your data cannot leave your environment, the audit runs inside it. Deployment is by Docker Compose.
Not a score, not a dashboard tour, not a slide deck about our platform. A document you can hand to your own security team, your customer, or your auditor.
You keep the report either way. If you decide not to continue with ARIA after the two weeks, the findings are still yours; they describe your system, not our product.
A finding is not "the model may exhibit unsafe behavior." It names what happened, how often, and what it means for you.
The scheduling agent approved an out-of-scope financial transaction in 3 of 40 adversarial test cases.
Illustrative example. Real findings come from your endpoint and are never published without your written permission.
It is not:
And a deliberate limit:
We produce the evidence. Your independent auditor certifies it. We do not audit our own work and we will not offer to — an auditor who also sells you the compliance platform is grading their own homework, and that applies to us exactly as much as to anyone else.
Independence is not a limitation here. It is the feature.
Which endpoint, which frameworks matter to your customers, what your data-handling constraints are, and the end date. Everything is agreed here and nothing moves afterwards.
You provide an endpoint and a key scoped to a test budget. ARIA runs its first pass and establishes a baseline for how your system behaves today.
The harder cases run: the ones designed to find the boundary rather than confirm the happy path. Findings are reproduced, then mapped to your framework clauses.
You receive the findings report and a 30-minute walkthrough. Then the engagement ends, on the date we agreed, whether or not anything follows.
That includes software vendors shipping an AI feature, services and consulting firms implementing AI inside client environments, and teams running AI in production under a regulator's eye. What matters is not the industry label; it is whether somebody downstream has started asking you to account for what your AI actually does.
Why free, for now. We are assembling a founding cohort. What we want from it is the findings and, where you are willing, a reference — not the fee. The offer is limited by how many engagements one team can run properly at a time, so it will not stay open indefinitely.
Tell us about your AI endpoint. If it is not a fit, we will say so on the first call.