TTIC  ·  June 2026

Behavioral Verification Is Not Optional
for AI Agents in Regulated Industries.
Here Is Why TTIC Is Testing It
as a Certification Requirement.

The Governance Gap

Why Existing Frameworks Fail for
Agentic AI in Regulated Industries

Every compliance framework in healthcare was designed for a deterministic world. You write a policy, configure a control, conduct an audit. The system does what you told it to do. Documentation of intent maps roughly to documentation of behavior because the system does not deviate from its configuration.

AI agents break that assumption at a fundamental level. Their behavior is probabilistic, context-dependent, and changes over time without anyone modifying the configuration. SOC 2 does not measure this. HIPAA technical safeguards do not measure this. Most of ISO 27001 does not measure this. They measure whether controls were documented and in place during an audit window. That is a different question from whether the system actually behaved correctly under the full range of conditions it will encounter in production, including adversarial ones.

This is the gap that TTIC was built to address. And it is the gap that made Praxen immediately compelling when Steve Wilson brought it to our attention.

Governance Architecture

What Enterprise Clinical AI Governance
Actually Requires

Most vendors present clinical AI governance as two layers: model accuracy and basic vendor documentation. Enterprise clinical AI governance requires eleven. The realization that Praxen and Darwin together covered all eleven layers is what inspired the TTIC certification pathway.

Enterprise clinical AI governance layers diagram
This diagram was presented by Sherri Douville and Mitch Parker in the Maturity Model Committee update to the IEEE/UL 2933 committee. Image copyright Sherri Douville, CEO Medigram, Founder & Chair TTIC.

Not as a product decision, but as a governance completeness realization. Praxen addresses the technical and behavioral layers through pre-deployment and ongoing verification. Darwin governs the clinical transaction layers continuously after deployment. Together they cover the full stack.

This layer reflects TTIC's independent analysis of regulatory and compliance considerations, including first-principles alignment with FDA's published cybersecurity and quality system frameworks. No FDA review or endorsement of this framework has been sought or obtained.

A New Discipline

What Agent Behavior Verification Is
and Why Regulated Industries Need It

Agent Behavior Verification is a new and important discipline in AI security. It is the practice of probing an AI system's actual behavior against its intended behavior under adversarial conditions, before deployment and on an ongoing basis after deployment. Praxen is the open-source reference implementation of ABV, built by Steve Wilson and released under Apache 2.0 by Exabeam.

Praxen reflects and extends Exabeam's established leadership in behavior intelligence and agent security. Where Exabeam has long led the field in detecting anomalous behavior by humans and non-human identities at runtime, Praxen extends that leadership upstream into pre-deployment verification, closing the gap between what an AI agent is configured to do and what it actually does before it ever reaches production. It is a natural and significant extension of a leadership position Exabeam has built over years.

For regulated industries, the stakes are distinct from the general enterprise context. In healthcare, behavioral deviation is not a productivity loss. It is a patient safety event, a liability exposure, and an institutional accountability failure. A clinical AI system that behaves correctly 98% of the time and deviates 2% of the time in ways that affect clinical decisions is not a software quality problem. It is a governance failure.

Praxen's RAISE framework scores behavioral maturity across six categories: domain limitation, knowledge base balance, zero trust implementation, supply chain governance, adversarial testing, and continuous monitoring. Each maps directly to governance requirements that regulated industries face. Darwin is engineered to complement this with deterministic, auditable clinical transaction scoring that enterprises can rely on by design rather than by assumption.

The Six RAISE Categories

Understanding RAISE: A Clinical AI Governance Perspective

By Sherri Douville CEO, Medigram · Founder & Chair, TTIC · Co-Chair, Trust and Identity Subgroup, IEEE/UL 2933 · Series Editor, Taylor & Francis · Brian Yam Chair, Pro Sports Track, TTIC · COO, Somnology

Each category below includes framing for:

Health System Executive Vendor Clinical AI Developer Pro Sports Executive Physician Investor
A behavioral constraint that restricts an AI agent to operating only within its defined scope of competency. The agent must refuse requests outside its authorized domain and must not attempt tasks it was not designed or validated to perform — even when users rephrase, escalate, or apply adversarial pressure to cross that boundary.

Read the full definition in the Praxen documentation ↗

In clinical AI, domain limitation is a patient safety control, not just a technical boundary. An AI system authorized to support medication reconciliation that responds to questions about surgical planning has violated a governance boundary that directly affects care. TTIC evaluates domain limitation in terms of clinical accountability: what is the remit of this agent, who authorized that remit, and does the agent actually stay within it under adversarial probing? The authorization of scope is a governance act, not a product configuration decision.
Domain limitation is about whether the AI knows what it is for — and behaves accordingly even when pushed. It is measured by probing the agent with out-of-scope requests and evaluating whether it refuses, escalates, or inappropriately responds. A high score means the agent consistently declines tasks outside its authorized boundary. A low score means the agent answers questions it should not, creating both liability and patient safety risk. The test is not what the system is configured to do — it is what the system actually does under adversarial conditions.
Define the clinical remit of the AI agent in a governance document before deployment — specific use cases, specific clinical roles authorized to interact with it, and explicit out-of-scope exclusions. Configure system prompts and guardrails to enforce the boundary. Run Praxen to probe whether the agent actually respects that boundary under adversarial conditions. Document the boundary definition, the probing methodology, and the findings as part of your clinical AI governance record. Re-verify after any model or configuration update.
How Does This Relate to Me

Click on the button for your title to see how this applies to you.

Addressed by Sherri Douville  ·  CEO, Medigram  ·  Founder & Chair, TTIC  ·  Co-Chair, Trust and Identity Subgroup, IEEE/UL 2933     Brian Yam  ·  Chair, Pro Sports Track, TTIC  ·  COO, Somnology
When an AI agent operates outside its authorized clinical scope, liability attaches to the institution, not the vendor. Domain limitation evidence establishes that the system was configured and verified to stay within its governance-approved remit. Require your vendors to show documented Praxen domain limitation scores before deployment approval — and make that requirement part of your contracting language.
Domain limitation is the first thing a health system procurement team will probe after your sales cycle ends. If your clinical AI agent responds to out-of-scope queries under adversarial conditions, your deal is at risk. A documented Praxen domain limitation score demonstrates that your system behaves within its authorized boundary under pressure — a governance artifact that competitors without it cannot match.
Implement domain limitation at the prompt engineering and retrieval layer, not just in the UI. A system that refuses domain-adjacent queries in the interface but responds to them when prompted directly has not solved domain limitation — it has hidden it. Praxen probes both layers. Design your refusal logic to be robust to adversarial rephrasing, role-play framing, and jailbreak attempts before you ship.
AI wellness and performance systems authorized for athlete recovery monitoring must not drift into contract-relevant performance evaluation — collective bargaining agreements define that boundary explicitly. Domain limitation verification ensures the AI stays within its authorized use case under the adversarial pressure of playoff preparation and contract year contexts. The governance boundary protects both the organization and the athlete.
Domain limitation protects you from acting on AI output generated outside the system's validated scope. An AI decision support tool that answers questions it was not validated for is providing unverified clinical guidance. Understanding what the AI is authorized to do — and asking whether it was adversarially probed to verify it stays there — is a clinical accountability competency that belongs in your AI governance participation.
Domain scope violations create liability exposure that manifests as M&A risk and insurance cost. Portfolio companies that can demonstrate adversarially-verified domain limitation evidence are materially more defensible in due diligence than those who cannot. A documented Praxen score in this category is a governance artifact with direct implications for enterprise value and D&O insurance underwriting.
Requires that the AI agent's knowledge sources represent a balanced, current, and appropriately curated body of information. The agent must not be systematically biased toward a single source, vendor, or perspective, and must have mechanisms to surface knowledge gaps rather than generating confident incorrect output when its knowledge is thin or outdated.

Read the full definition in the Praxen documentation ↗

In clinical settings, knowledge base balance is directly tied to care equity and clinical accuracy. An AI system trained predominantly on research from academic medical centers may perform poorly in community hospital contexts. One drawing from data representing dominant demographics may systematically underserve other populations. TTIC evaluates this category through the lens of institutional accountability: what is the provenance of this system's knowledge, how was it curated, and does the curation record stand up to scrutiny in a governance review by clinical leadership?
This category measures whether the system's knowledge architecture is defensible — not just technically, but from a clinical governance standpoint. It asks whether the AI knows what it does not know, whether it surfaces uncertainty rather than confabulating, and whether the curation of its knowledge base reflects the institution's values and patient population. A high score means the system has documented, auditable knowledge governance. A low score means the knowledge base is opaque, potentially biased, and not defensible in a regulatory audit or malpractice discovery process.
Document the sources, vintage, and curation methodology of every knowledge base component before deployment. Establish a review cadence for knowledge currency — particularly for clinical guidelines that change frequently. Test the system's behavior when queried on topics where its knowledge is thin: does it acknowledge uncertainty or generate confident incorrect output? Require your AI vendor to provide knowledge provenance documentation as a procurement condition. Include knowledge base governance documentation in your clinical AI governance record.
How Does This Relate to Me

Click on the button for your title to see how this applies to you.

Addressed by Sherri Douville  ·  CEO, Medigram  ·  Founder and Chair, TTIC  ·  Co-Chair, Trust and Identity Subgroup, IEEE/UL 2933     Brian Yam  ·  Chair, Pro Sports Track, TTIC  ·  COO, Somnology
Your institution will be held accountable for clinical decisions influenced by AI systems with biased or outdated knowledge bases. Knowledge base balance governance documentation is the institutional record that demonstrates due diligence in selecting and monitoring AI that reflects your patient population and current clinical standards. It is the governance artifact that answers the question: did we know what this system was trained on?
Health system procurement teams increasingly require knowledge provenance documentation as part of due diligence. A clinical AI product with a documented, auditable knowledge curation process and a Praxen score in this category demonstrates governance maturity that undocumented competitors cannot match. The ability to show knowledge currency protocols — not just a training data cutoff date — is a meaningful enterprise sales differentiator.
Retrieval-augmented generation systems require active knowledge governance, not just initial curation. The knowledge base you shipped with is not the knowledge base that will govern clinical interactions in 18 months. Build update workflows, deprecation protocols, and uncertainty-surfacing mechanisms into your system architecture from the start. The cost of retrofitting knowledge governance into a deployed clinical AI system is significantly higher than building it in.
Performance and wellness AI systems for athletes draw on sports science literature that evolves rapidly. Knowledge base currency directly affects recommendation quality. Systems referencing outdated recovery protocols or injury prevention literature create both performance and health risk. Verified knowledge currency is a competitive governance requirement — especially for teams where the margin between health and injury in a playoff run is measured in days.
You are the last line of defense when a clinical AI system generates output based on outdated or unbalanced knowledge. Understanding the knowledge provenance of AI tools you interact with — and having access to governance documentation showing what the system knows and when it was last updated — is essential to your professional accountability. Ask your clinical informatics team what knowledge currency review is required for AI tools in your department.
Knowledge base obsolescence is an underappreciated product risk in clinical AI. Systems requiring continuous manual curation to remain clinically accurate carry higher ongoing operational costs and higher liability exposure. Portfolio companies with automated knowledge governance and documented currency controls are better positioned for scale and have lower long-term governance costs than those treating knowledge base management as a post-deployment operational afterthought.
Requires that the AI agent operate under zero trust security principles: every interaction, data access request, and system call is authenticated and authorized explicitly, with least-privilege access enforced at each layer. No entity — user, system, or agent — is trusted by default based on prior authentication state. Trust must be continuously re-earned, not inherited from prior sessions.

Read the full definition in the Praxen documentation ↗

Zero trust in clinical AI is a regulatory and liability imperative, not just a security best practice. HIPAA, state privacy laws, and emerging AI governance frameworks all assume that clinical data access is explicitly authorized per transaction. An AI agent that caches authorization states, assumes a prior authenticated session grants ongoing data access, or escalates its own privileges to complete a clinical task has created a potential HIPAA violation and a governance accountability failure. TTIC evaluates zero trust implementation in terms of whether the authorization architecture is defensible in a regulatory audit and whether an evidentiary record of every access decision exists.
Zero trust implementation is measured by whether the agent attempts to access data or capabilities beyond its explicitly authorized scope, whether it can be induced to escalate privileges through adversarial prompting, and whether every data access event is logged with sufficient granularity for regulatory audit. A system that authenticates users at login but then grants broad ongoing access to clinical data has not implemented zero trust. It has implemented convenience with compliance-adjacent labeling. Praxen probes both the authorization architecture and the logging completeness.
Implement per-request authentication and authorization for every clinical data access event. Apply least-privilege access at the data element level — an AI authorized to read medication history should not automatically access psychiatric history in the same query. Ensure every access event is logged with timestamp, user identity, agent identity, data element accessed, and authorization basis. Run Praxen to probe privilege escalation and authorization bypass vectors. Verify that the logging architecture produces records sufficient for a HIPAA audit without manual reconstruction.
How Does This Relate to Me

Click on the button for your title to see how this applies to you.

Addressed by Sherri Douville  ·  CEO, Medigram  ·  Founder and Chair, TTIC  ·  Co-Chair, Trust and Identity Subgroup, IEEE/UL 2933     Brian Yam  ·  Chair, Pro Sports Track, TTIC  ·  COO, Somnology
HIPAA audit exposure is the most immediate organizational risk in this category. Health systems that cannot produce a per-transaction authorization record for AI-accessed clinical data face regulatory exposure that individual vendors cannot indemnify. Zero trust implementation evidence is the institutional governance artifact that demonstrates your AI access architecture is defensible — not just your vendor's claims about it, but verified evidence from adversarial testing.
Zero trust implementation is increasingly a procurement gate, not a differentiator. Health systems with established security governance programs require per-transaction authorization evidence. A Praxen score in this category, combined with SOC 2 Type II certification, is the documentation portfolio that answers security review questions with evidence rather than policy documents. It is the difference between passing the security questionnaire and passing the security review.
Token-based session authentication is not zero trust. Every call to a clinical data source from your AI agent requires explicit authorization evaluation at the time of the call. Build authorization middleware into your agent architecture before deployment. Retrofitting zero trust into a production clinical AI system is significantly more expensive and risky than designing it in from the start — and the retrofitting usually surfaces additional vulnerabilities that create deployment delays at the worst possible time.
Athlete health data is governed by collective bargaining agreements that define authorized access scope per use case. AI systems that assume broad health data access based on team employment relationship create CBA violations and, increasingly, player association grievance exposure. Zero trust implementation ensures that AI data access is limited to what each specific use case explicitly requires — and that the authorization record exists to demonstrate it.
Your patients' clinical data is accessed by AI systems as part of care delivery. Zero trust implementation ensures that access is limited to what each clinical interaction explicitly requires — and that an auditable record exists of every access event. When patients ask who has seen their data, or when a regulatory inquiry requires a data access log, a zero trust architecture provides the answer. Asking whether the AI tools in your department meet this standard is a legitimate clinical governance question.
HIPAA penalty exposure for unauthorized AI data access is uncapped at the federal level and compounds with state privacy law violations. Portfolio companies that cannot demonstrate per-transaction authorization architecture for clinical AI data access carry regulatory tail risk that is difficult to quantify in pre-revenue or early-revenue diligence. Zero trust implementation evidence materially reduces that risk profile — and demonstrates technical governance maturity that correlates with lower overall regulatory exposure.
Requires that the AI agent's operator have documented visibility into every third-party component — foundation models, libraries, APIs, data sources, and infrastructure services — that contributes to the agent's behavior. Vulnerabilities, behavioral changes, or governance failures in any supply chain component can propagate to the agent's output without the operator's knowledge unless active supply chain governance is in place.

Read the full definition in the Praxen documentation ↗

Clinical AI supply chains are opaque by default and dangerous by consequence. A clinical AI product typically runs on a foundation model from one vendor, uses a vector database from another, retrieves clinical guidelines from a third-party source, and is deployed on cloud infrastructure from a fourth. A behavioral change in any of these layers — a model update, a library patch, a guideline source change — can alter the clinical AI system's behavior without any corresponding change to the product the health system purchased or approved. TTIC evaluates supply chain governance through the lens of the institutional accountability question: if this system behaves unexpectedly tomorrow, do you have the documentation chain to identify where the change came from?
Supply chain governance is measured by the completeness and currency of the vendor's AI-specific software bill of materials (SBOM), by active monitoring for behavioral changes originating in supply chain components rather than the product layer, and by the vendor's demonstrated ability to respond to supply chain events — vulnerabilities, model updates, third-party behavioral changes — with documented governance decisions. A high score means the vendor can account for every behavioral dependency. A low score means the product's behavior is partially determined by components the vendor is not actively governing.
Produce and maintain an AI-specific SBOM that includes model versions, API versions, data source vintages, and infrastructure dependencies. Establish a monitoring process for supply chain changes that includes behavioral re-verification when upstream components update. Require foundation model providers to notify you of behavioral changes, not just security patches. Run Praxen after significant supply chain changes and document the results as part of your re-verification record. Include supply chain governance documentation in your clinical AI governance record as an ongoing artifact, not a one-time assessment.
How Does This Relate to Me

Click on the button for your title to see how this applies to you.

Addressed by Sherri Douville  ·  CEO, Medigram  ·  Founder and Chair, TTIC  ·  Co-Chair, Trust and Identity Subgroup, IEEE/UL 2933     Brian Yam  ·  Chair, Pro Sports Track, TTIC  ·  COO, Somnology
When a clinical AI system produces a harmful output, institutional counsel will ask who was responsible for each component that contributed to it. Supply chain governance documentation is your institutional evidence that you required and received component accountability from your vendors before and during deployment. Without it, liability attribution becomes contested and the institution bears residual exposure that vendor indemnification clauses typically do not cover.
Health systems are beginning to require AI-specific SBOMs as part of procurement due diligence. A vendor who can produce a current, auditable SBOM for their clinical AI product, demonstrate active supply chain monitoring, and show Praxen verification after supply chain changes is differentiated from competitors who cannot. The ability to show documented governance decisions when a foundation model updates — not just a product release note — is a meaningful enterprise trust signal.
Foundation model updates are the most underappreciated behavioral risk in clinical AI development. If your clinical AI product's behavior changes when your underlying model updates and you do not catch it with behavioral re-verification before it reaches production, you have delivered an unverified governance change to every health system running your product — simultaneously, and without their knowledge. Build re-verification triggers into your supply chain monitoring architecture.
Wearable and performance analytics platforms for athletes aggregate data from multiple vendors and apply AI layers from third-party providers. A behavioral change in any supply chain component can alter the recommendations reaching coaching staff and medical teams without triggering any product update notification. Supply chain governance documentation is the organizational record that demonstrates you are monitoring what your AI systems are actually doing, not just what they were approved to do at deployment.
Clinical AI systems used in care delivery may behave differently from the versions evaluated in your clinical review committee, if underlying components have been updated without revalidation. Understanding your institution's supply chain governance requirements for clinical AI vendors — and asking whether they run behavioral re-verification after model updates — is a meaningful clinical governance participation question that protects your patients and your professional accountability.
Supply chain liability in clinical AI is similar to medical device component liability: it can originate upstream and manifest as a product failure downstream. Portfolio companies without active supply chain governance are carrying behavioral risk that is unpriced and difficult to quantify until it manifests. Companies with documented SBOM governance and behavioral re-verification protocols have materially lower regulatory and liability exposure profiles, and more predictable governance cost structures.
Requires that the AI agent be subjected to structured adversarial testing by parties whose goal is to make it fail — to produce harmful output, violate its constraints, expose hidden capabilities, or behave in ways that contradict its governance documentation. Red teaming is a pre-deployment verification requirement, not a post-deployment audit, and must be documented with findings and remediation records.

Read the full definition in the Praxen documentation ↗

In clinical AI, adversarial testing is a clinical safety requirement. AI systems deployed in care delivery settings will be used by a population that includes, intentionally or not, those who ask questions the system should not answer, apply it to use cases it was not validated for, and interact with it in ways that surface edge case behaviors. Red teaming replicates these conditions before deployment so that behavioral deviations are discovered and remediated in a controlled context rather than discovered through adverse events in care delivery. TTIC evaluates red team results in terms of the clinical severity of discovered deviations — not just their technical existence — and treats the remediation record as a clinical accountability document.
Adversarial testing is measured by the rigor and independence of the red team methodology, the clinical realism of attack scenarios, and the completeness of the remediation record. A high score means the system has been subjected to structured, documented adversarial testing that covers clinically relevant attack vectors, produced findings, and shows same-day or documented-timeline remediation with re-verification. A low score means the system has only been tested under conditions designed to make it succeed — which tells you nothing about how it will behave when users, intentionally or accidentally, push it toward failure.
Structure your red team as an independent function with no stake in deployment timelines. Define adversarial scenarios based on your specific clinical use case — not generic LLM red team scripts. Document every probe, every finding, and every remediation decision. Re-run probes after remediation to verify they are closed. Use Praxen as the structured framework for your adversarial testing methodology and use its scoring to create an auditable findings-and-resolution record. Run three successive assessment cycles, treating each run's findings as a remediation roadmap. This iterative cycle-and-remediate approach is the methodology TTIC recommends for any organization pursuing certification, and reflects the rigor required to close findings systematically rather than in a single pass.
How Does This Relate to Me

Click on the button for your title to see how this applies to you.

Addressed by Sherri Douville  ·  CEO, Medigram  ·  Founder & Chair, TTIC  ·  Co-Chair, Trust and Identity Subgroup, IEEE/UL 2933     Brian Yam  ·  Chair, Pro Sports Track, TTIC  ·  COO, Somnology
Adversarial testing documentation is the institutional evidence that you required your AI vendor to break their own system before you deployed it in your facility. Without red team evidence, your procurement record shows due diligence on features and pricing but not on safety. When a clinical AI adverse event reaches your risk committee, pre-deployment red team results are the difference between a defensible governance record and an institutional accountability gap that plaintiff counsel will exploit.
Clinical AI vendors who can produce structured, documented pre-deployment red team results — and show the remediation record — are the only vendors who can credibly answer the question 'how do you know this system will not harm a patient?' in a procurement evaluation. Red team evidence is rapidly becoming a health system contracting requirement. The vendors who have it will close deals faster and with less legal friction than those who rely on SOC 2 and penetration testing reports to satisfy clinical governance questions.
Red teaming your own system is a structural conflict of interest. Your team knows how the system is designed and will unconsciously avoid the probes most likely to surface failures. Independent red teaming using Praxen's adversarial testing framework produces findings that internal testing consistently misses — particularly around prompt injection, domain boundary violation under adversarial framing, and unexpected capability exposure that occurs when users interact with the system in ways the development team did not anticipate.
AI systems used in athlete health assessment will be tested adversarially by the conditions of actual use — high-pressure playoff preparation, injury recovery timelines with competitive stakes, contract year performance contexts. Pre-deployment red teaming that simulates these pressures surfaces behavioral deviations before they affect athlete health decisions or create institutional exposure. The red team methodology should reflect the specific adversarial contexts your organization's AI systems will actually face during a competitive season.
Red team results are the pre-deployment evidence that the clinical AI system you interact with has been tested to fail safely — that when users ask it to do things it should not do, it declines rather than complies. Asking your clinical informatics leadership whether vendor AI systems were red teamed before deployment, whether the results are available for clinical governance review, and what the most significant findings were — and how they were remediated — are legitimate and important clinical accountability questions.
The absence of red team documentation in a clinical AI company's governance portfolio is a material due diligence gap. An AI system that has never been adversarially tested before clinical deployment is carrying unknown behavioral risk that will manifest in production. Portfolio companies with documented, independent red team programs — particularly those using structured frameworks like Praxen — have better-defined risk profiles, lower expected governance costs, and more defensible enterprise sales narratives in a market that is maturing toward requiring this evidence.
Requires that behavioral verification not be a point-in-time pre-deployment event but an ongoing operational discipline. The AI agent's behavior must be monitored continuously in production, with mechanisms to detect behavioral drift, escalate anomalies, and trigger re-verification when behavioral baselines change. Deployment is the beginning of governance, not its conclusion.

Read the full definition in the Praxen documentation ↗

Clinical AI systems operate in environments that change continuously — clinical workflows evolve, patient populations shift, user behaviors adapt, and the underlying AI infrastructure updates. A system that passed adversarial testing and behavioral verification at deployment may behave differently six months later due to model updates, data distribution shifts, or accumulated behavioral drift. TTIC evaluates continuous monitoring in terms of whether the institutional accountability structure exists to detect, escalate, and respond to behavioral changes in clinical AI systems that are already in production — and whether the governance owner for that function is a clinical accountability role, not just an IT operations role.
Continuous monitoring is measured by whether monitoring infrastructure exists, what behavioral metrics are tracked, what thresholds trigger escalation, who is responsible for reviewing monitoring alerts, and whether there is a documented re-verification protocol that triggers a new Praxen run when behavioral baselines change. A high score means the organization has a defined, staffed, and documented monitoring program with clinical governance ownership. A low score means behavioral verification stops at deployment — which, in clinical AI, means governance stops before production begins.
Define behavioral baselines at deployment from your Praxen results. Implement production monitoring that tracks output distribution, confidence calibration, refusal rates for out-of-scope queries, and anomalous response patterns. Establish escalation thresholds and designate clinical governance owners for monitoring alerts — not just IT security owners. Define a re-verification trigger protocol that requires a new Praxen run when monitoring detects baseline deviation. Document the monitoring architecture and escalation structure in your clinical AI governance record as a standing organizational commitment.
How Does This Relate to Me

Click on the button for your title to see how this applies to you.

Addressed by Sherri Douville  ·  CEO, Medigram  ·  Founder & Chair, TTIC  ·  Co-Chair, Trust and Identity Subgroup, IEEE/UL 2933     Brian Yam  ·  Chair, Pro Sports Track, TTIC  ·  COO, Somnology
Clinical AI governance does not end at go-live. Behavioral drift in production clinical AI systems is the emerging governance gap that most institutions have not yet staffed for. Establishing who in your governance structure owns the monitoring and escalation responsibility for each deployed clinical AI system is an institutional accountability decision that belongs at the leadership level, not in IT operations. Build it into your clinical AI deployment approval process before the next system goes live.
Continuous monitoring is the governance commitment your clinical AI customers most need from you and least often receive. Offering documented behavioral monitoring as part of your deployment package — with defined escalation SLAs, behavioral baseline documentation from Praxen results, and a re-verification commitment when baselines change — is a differentiating governance offering that health system risk committees and procurement leaders will notice and value.
Production behavioral monitoring is architecturally different from pre-deployment testing. Build monitoring hooks into your system at the inference layer, not just the application layer. Track behavioral metrics — refusal rates, output distribution, confidence calibration — not just performance metrics. Latency and uptime monitoring do not tell you that your system has started answering out-of-scope queries or drifting toward confident-but-wrong outputs. Build monitoring that does, and define the alert thresholds before you ship.
AI systems used across a competitive season will encounter conditions at week 16 that did not exist at week 1 — different injury states, different performance contexts, different user behaviors from coaching staff under playoff pressure. Continuous monitoring for behavioral drift is the governance mechanism that detects when an AI system's behavior has changed from what was deployed and approved at the start of the season. Without it, you are governing an AI system based on what it was, not what it is.
Continuous monitoring is the institutional commitment that the clinical AI system you interact with today is behaving the same way it was validated to behave — and that someone is responsible for detecting and escalating changes if that is no longer true. Asking your clinical informatics team what monitoring is in place for AI systems in your department, who owns the escalation pathway, and what the re-verification protocol is when behavioral drift is detected — these are legitimate and important clinical accountability questions.
Continuous monitoring capability is a governance moat for clinical AI companies. Companies offering documented, staffed behavioral monitoring as part of their deployment model have lower customer churn — governance-linked contracts are stickier — and lower regulatory exposure from undiscovered production behavioral drift. In a market maturing toward requiring continuous verification rather than point-in-time certification, companies with this capability built into their product are better positioned for enterprise sales, contract renewal, and regulatory defensibility.
Early Access Results

What We Learned Running It

Medigram accessed Praxen through its early access program before public launch and ran it against Darwin, the governed clinical AI platform, across three successive assessment cycles. TTIC evaluated the results in its capacity as the clinical governance body with the expertise to interpret what the findings mean in a clinical context.

The third run, using Praxen 0.8.0, produced the following result:

RAISE Score  ·  Praxen 0.8.0  ·  June 22, 2026
4.0
Strong  ·  Zero Open Findings
Medigram Darwin in the Praxen Early Access Program
Highest Praxen RAISE score recorded to date.
Steve Wilson  ·  Inventor, Praxen  ·  Chair AI Security, TTIC  ·  Chair, OWASP GenAI Security Project  ·  Chief AI & Product Officer, Exabeam

What the three-run cycle taught TTIC was more valuable than the final score. It taught us what the categories actually mean in a clinical context and how to interpret findings in terms of clinical accountability rather than just security posture. That interpretive layer is precisely why TTIC's early access certification program requires Praxen as Layer 1. The tool produces rigorous findings. Understanding what those findings mean for a governed clinical platform requires the expertise TTIC brings. A RAISE score issued without that interpretive layer is a number. Within the TTIC certification process it becomes a clinical accountability document.

Plaintiff attorneys often ask three questions to lock in accountability: what did you know, when did you know it, and what did you do. From what I have seen, a Praxen run answers all three before the questions are even asked.
Noel Gillespie  ·  Chair, Law, TTIC  ·  Partner, Buchalter  ·  Chair, MedTech Practice
Continual Care requires Continual Feedback and Review to Succeed.
Mitch Parker  ·  Co-Founder, TTIC  ·  CISO, IU Health
Honest Scope

The Stadium
and the Field

The best analogy for what Praxen and Darwin together certify is a stadium and a field. The infrastructure is certified and ready. The behavioral verification has been conducted. The governance architecture is in place. But the teams still must coach and the players still have to play the game.

That is the honest scope of what technical and behavioral certification provides. It verifies that the infrastructure meets the standard required for governed clinical play. It does not play the game for you. The clinical workflows, the physician oversight structures, the accountability decisions made by health system leaders — those are the game. The TTIC certification early access program is testing whether this infrastructure is consistently ready for it.

Trust is the ultimate measure of success for AI in healthcare. This milestone reflects our commitment to earning it every day.
Dr. Apurv Gupta, MD, MPH  ·  Chair, Clinical Excellence, TTIC
Standards create confidence. Execution creates advantage. The future of AI belongs to organizations that move fast and govern well. TTIC transforms IEEE/UL 2933 from technical guidance into operational excellence, enabling organizations to accelerate innovation, confidently adopt trusted AI vendors, and achieve measurable governance, proven by the highest recorded Praxen RAISE score of 4.0.
Brian Yam  ·  Chair, Pro Sports Track, TTIC  ·  COO, Somnology
Before and After

How Praxen Complements
Runtime Governance

Praxen is a pre-deployment and ongoing behavioral verification instrument. Darwin's governed transaction architecture ensures that every clinical decision made by the system after deployment is captured as an immutable governance record, with identity, scope, and audit trail intact, as a structural condition of the transaction completing. One verifies before. The other governs continuously after.

TTIC's early access certification program requires both because clinical AI governance is not a point-in-time event. It is a continuous accountability structure. Praxen gives the pre-deployment evidence. Darwin gives the runtime evidence. Together they produce the documentation chain that health system procurement, D&O underwriters, and the institutional accountability structures of regulated industries require.

Certification Pathway

What the Early Access
Certification Program Requires

TTIC is currently testing its certification pathway through an early access program. The pathway is structured in two layers. Layer 1 is Praxen RAISE behavioral verification. Layer 2 is Darwin TIPPSS governance screening across six clinical dimensions.

TIPPSS screening functions as a required input to TTIC certification; certification itself is issued by TTIC's governance review, not by TIPPSS scoring alone.

Tier Scope RAISE Threshold TIPPSS Requirement
Tier 1 Direct clinical data contact 4.0 or above TIPPSS Level 3 · All six dimensions · Documented adversarial testing
Tier 2 Operational & infrastructure vendors 3.0 or above TIPPSS screening · Three dimensions

The first step for any vendor or health system interested in the early access program is defining the governance committee with authority to sign off on the work remit for a Praxen run. This is not an engineering decision. It is a clinical accountability decision. The remit defines scope, the accountable parties for findings, and the remediation pathway. Without that structure, the Praxen run produces findings with no governance pathway to resolution.

Praxen is free and open source. You can begin exploring it today. For information about participating in the TTIC certification early access program, contact Sherri Douville directly.

Good intentions do not govern AI. Structure does.
Khalid Turk  ·  Head Standards Operationalization, TTIC
Behavioral verification changes the diligence conversation for AI agents. You are no longer asking whether the system claims to be safe. You are asking whether it can prove it. When you prove it, that changes the investment calculus.
Tom Bondi  ·  Chair, Finance, TTIC  ·  Managing Partner, TCBONDI CPA, INC.
About

Author & Organizations

Sherri Douville
Founder and Chair, TTIC  ·  CEO, Medigram
Sherri Douville is CEO of Medigram and Founder and Chair of the Trustworthy Technology & Innovation Consortium (TTIC), a non-commercial standards consortium of practitioners and investors across medicine, cybersecurity, engineering, standards, law, operations, and finance. She co-authored ANSI/HSI 2800:2025, the Hospital AI Operations Governance standard, leading the TTIC contribution, and serves as Trust and Identity Subgroup Co-Chair for IEEE/UL 2933:2024. She is a Series Editor at Taylor & Francis, where she edited Mobile Medicine and Advanced Health Technology; six books carry her work across Taylor & Francis, Springer, and Artech House. At Medigram she directs the architecture of Darwin, the company's governed AI decision infrastructure for healthcare. In 2025 she received the AI Champion of the Year award at AIMed25.
TTIC
Trustworthy Technology Innovation Consortium
TTIC's mission is to deliver leadership that delivers metrics. A non-commercial standards consortium that convenes clinical AI governance expertise across health systems, standards bodies, security professionals, and clinical leaders. Co-authored IEEE/UL 2933:2024 and ANSI/HSI 2800:2025. Currently testing a clinical AI certification pathway through an early access program. TTIC operates with an intentional organizational firewall from Medigram. TTIC does not sell products. It establishes the governance architecture that clinical AI platforms are accountable to.
For TTIC Certification Early Access Information
Explore Praxen
open-agent-ai-security.github.io/praxen ↗
Launch Announcement
Read the Exabeam launch announcement ↗
Launch Impact

The launch also had meaningful reach. In the first few days post, the Praxen launch generated more than 20 placements across cybersecurity, AI, and enterprise technology media, including coverage in North America, Asia-Pacific, the UK, Europe, and Africa. Medigram as a technical demonstration under TTIC was the only clinical AI governance organization with a practitioner account published on launch day.

Take the Next Step

Praxen is free and open source. TTIC's certification early access program is open to health systems and vendors ready to demonstrate governed clinical AI.

Discussed on The New CISO (Exabeam), Episode 150 · Episode highlights

Choose TTIC as a preferred source in eligible Google experiences. Prefer TTIC in Google

Published by
Trustworthy Technology & Innovation Consortium (TTIC)
Author
By Sherri Douville, Founder & Chair, Trustworthy Technology & Innovation Consortium (TTIC)
Originally published
Last updated

Cite this resource

Sherri Douville. “Behavioral Verification Is Not Optional for AI Agents in Regulated Industries.” Trustworthy Technology & Innovation Consortium (TTIC), 2026. https://trustworthytechnologyinnovation.com/blog-praxen-behavioral-verification.