Responsible AI, Verified — continuous behavioral testing and compliance for AI.
PUBLISH ON: LinkedIn (Dr. Prasad Golla profile) — long-form article Word count target: ~1,800 words Tone: Authoritative practitioner. Measured urgency. Not hype.
Healthcare AI Is Flying Blind. We're Building the Instrument Panel.
Introducing ARIA — an AI governance assessment platform built for healthcare organizations deploying Large Language Models in clinical environments.
In 2007, a Veterans Administration research grant worth $80 million was cancelled.
Not because the science failed. Not because the funding dried up. Because the participating institutions — brilliant researchers at world-class medical centers — could not share patient context data in a way that was simultaneously secure, HIPAA-compliant, and ethically defensible.
The technology to do it existed. The governance framework to do it safely did not.
The VA walked away. Gulf War veterans waiting for answers about their health waited longer.
I was there. I was part of the team at UT Southwestern Medical Center trying to build that platform. And when the grant was cancelled, I filed away a lesson that has taken nearly twenty years to fully surface.
Recently, Sridhar Yerramreddy — CEO of Steer Health and one of the most respected voices in healthcare AI — described a pattern he calls the "Islands of LLMs."
Dozens of AI tools deployed across a health system. Each one evaluated and procured independently. Each one operating in isolation. No unified view of what data they touch, what risks they carry, or what happens when one of them fails.
Sound familiar? It should. It is the same fragmentation problem that killed the Gulf War Syndrome research in 2007. Same category of mistake. Different technology. Twenty years later.
The difference today is scale and speed. In 2007, the fragmentation involved a handful of research institutions and one database protocol disagreement. In 2026, a mid-sized health system may be running 30, 50, or 100 AI tools — clinical scribes, triage assistants, diagnostic support systems, patient communication bots, revenue cycle automation — none of which are governed by a unified framework, and most of which were never formally assessed for safety, bias, or regulatory compliance before deployment.
The numbers are not reassuring.
According to the Grant Thornton 2026 AI Impact Survey, 77% of organizations are now actively working on AI governance — yet most lack the tools to implement it effectively. IBM's Cost of a Data Breach Report shows that shadow AI already accounts for 20% of all breaches, adding an average of $670,000 in costs per incident. And Gartner's long-standing finding — that 85% of AI projects deliver erroneous outcomes — has not improved with the proliferation of LLMs. It has gotten more complicated.
In clinical environments, "erroneous outcomes" is not a budget line. It is a patient safety event.
When I ask healthcare organizations about their LLM governance, I hear variations of the same three answers:
"We have RBAC." Role-based access control at the database layer does not protect against prompt injection at the LLM layer. A properly structured query to a clinical scribe system can extract records it was never authorized to surface — without touching the database directly. RBAC is not an AI security strategy.
"Our vendor handles compliance." Your vendor's Business Associate Agreement is a legal instrument, not a technical control. It does not validate that your vendor's model is not hallucinating medication dosages, drifting from its validated baseline, or processing PHI in ways your contract does not cover. Have you tested it?
"We validated it before deployment." Once. On clean data. Before the vendor updated the model. Clinical AI validation is not a checkbox — it is a continuous process. Every vendor model update is effectively a new system deployment requiring re-validation. Most organizations do not know when their vendor's model has changed.
These are not hypothetical vulnerabilities. They are the exact gaps I have documented in clinical AI deployments across multiple health systems. And they are precisely the gaps that neither HIPAA guidance, nor NIST AI RMF documentation, nor any general-purpose GRC platform is currently equipped to catch.
The AI governance market is moving fast. Platforms like Credo AI, Holistic AI, and Microsoft Purview are building capable products. VerifyWise has built an impressive open-source framework. IBM OpenPages has decades of GRC credibility.
Most of them are built for general enterprise governance, not for clinical environments. They tend not to surface the FDA Clinical Decision Support classification question — the single most consequential regulatory determination a healthcare AI startup can face, because getting it wrong means deploying an uncleared medical device under federal law. They tend not to model the dependency between a clinical scribe's PHI logging and a Business Associate Agreement that may not cover that logging. They tend not to ship with HIPAA assessment integrated into the same workflow that asks about NIST AI RMF GOVERN, MAP, MEASURE, and MANAGE controls.
The general-purpose platforms also tend to start at six figures per year, require months of implementation, and deliver governance frameworks designed for financial services or manufacturing — frameworks that healthcare organizations then spend additional months and additional dollars trying to adapt.
A Series A healthcare AI startup cannot afford that. And they cannot afford not to have a governance posture either.
This week, Aggi Technologies LLC is launching ARIA — Aggi Responsible Intelligence Assessor.
ARIA is a modular, data-driven AI governance and compliance assessment platform built for healthcare organizations deploying LLMs in clinical and administrative workflows.
It is not a checklist. It is not a policy template library. It is an assessment engine with three design choices that I think materially differentiate it for the clinical use case:
1. Healthcare-calibrated question bank — across multiple frameworks, with no generic filler. ARIA's question bank spans NIST AI RMF (78 questions, the primary framework), HIPAA Security and Privacy (~30 questions), ISO/IEC 42001 (~25 questions), and a plain-English INTAKE on-ramp (29 questions) that propagates answers automatically into NIST, HIPAA, and ISO. FDA Clinical Decision Support questions (24, currently in draft pending review) round it out. Cross-standard mapping is real: 53 mappings shipped today (22 NIST↔HIPAA + 31 ISO↔NIST), so a single answer flows into peer-framework controls as `pending_review` until a human confirms it.
Every question is calibrated for clinical context. Not "do you have a governance policy" — but "does your AI governance policy define acceptable and unacceptable use cases specifically for patient-facing versus clinician-facing LLM tools?" Not "have you assessed bias" — but "has your bias audit used a dataset representative of the actual patient population this system serves, including age, race, language, and socioeconomic status?"
The specificity is the point. Generic questions produce generic answers that do not identify the gaps that matter in a clinical deployment.
2. Dependency graph — because compliance gaps are not independent. ARIA maps the relationships between compliance gaps. When your organization answers "no" to whether a Business Associate Agreement has been executed with your LLM vendor, ARIA does not just flag that gap in isolation. It surfaces every downstream exposure: the PHI that is legally flowing to that vendor without authorization, the audit trail gaps that follow, the incident response gaps that cascade from there.
Healthcare AI compliance is a system, not a list. ARIA shows the system.
3. Conditional logic engine — critical questions become gates, not just findings. Certain questions in ARIA's question bank are not merely important — they are gates. When an organization indicates that their LLM system produces outputs that automatically enter the EHR without human review, ARIA does not simply log that as a finding. It flags it as a critical patient safety risk, surfaces it in the leadership report, and connects it to the specific FDA guidance and HIPAA provisions that apply.
When an organization indicates uncertainty about whether their system qualifies as a Clinical Decision Support device under FDA guidance, ARIA unlocks the FDA CDS module, surfaces the warning page, and creates a CRITICAL finding — because deploying a regulated medical device without FDA clearance is a federal violation, not a compliance gap.
The platform gets smarter as organizations answer it. That intelligence is the product.
ARIA is live and available now at aria.aggicorp.com.
The platform is built in deliberate layers and the foundation is running: a multi-framework assessment engine spanning NIST AI RMF, HIPAA, ISO/IEC 42001, and HITRUST, with cross-standard auto-propagation so a single answer can satisfy peer questions across frameworks. Evidence management, conditional logic, remediation tracking, and audit-ready PDF export are operational. Dependency graph visualization, multi-audience reporting, executive dashboards, and a Maturity Mountain progress view are live. ControlMesh — the per-organisation analytic layer — adds posture trajectory, audit-readiness scoring across eight signals, counterfactual hover insights on every status button, and contradiction detection across an organisation's assessments. An LLM-assisted ingestion pipeline (with PII pre-flight scanning, hashed prompt+response audit logging, and a human review gate before any extracted answer is accepted) is live for healthcare practitioners who would rather upload a policy document than answer 78 questions by hand. FDA CDS questions are currently draft and require expert review before production use; the EU AI Act and SOC 2 modules are scaffolded but not yet populated.
The pricing model is built around the market that exists but is currently underserved: a free tier for the core assessment engine, a professional tier accessible to Series A and Series B healthcare AI startups, and an enterprise tier for health systems and AI vendors.
Not six figures up front. Not a six-month implementation. Not a governance framework designed for someone else's industry.
The EU AI Act's high-risk system provisions for healthcare take effect in August 2026. FDA guidance on AI-enabled medical devices is tightening. Enterprise healthcare customers are now including AI governance attestations in their procurement requirements. And the organizations buying healthcare AI tools are starting to ask vendors the questions that ARIA is designed to help them answer.
The governance infrastructure that should have accompanied the LLM deployment wave did not arrive with it. ARIA is part of closing that gap.
I spent nineteen years building toward a platform that could have prevented what happened in 2007. The VA grant that died because the data could not be governed correctly. The Gulf War veterans whose research went unfunded because the institutions could not get their infrastructure right.
We are not going to repeat that mistake.
Responsible AI, Verified.
Dr. Prasad N. Golla, Ph.D. | MBA | IEEE Senior Member Founder & Principal Consultant, Aggi Technologies LLC Creator, ARIA — Aggi Responsible Intelligence Assessor prasad@aggicorp.com | aggicorp.com | drgolla.com
Learn more about ARIA: aggicorp.com/aria Schedule a conversation: prasad@aggicorp.com
Aggi Technologies LLC helps regulated organizations govern AI behavior with ARIA and ARIA Shape B. Talk to us →
Written for people who have to make decisions about AI. What happened, why it matters, and what to do about it — with every source linked so you can check the work yourself.
Or read it on LinkedIn →