This article draws on foundational work that will define how clinical AI security is practiced at scale. Steve Wilson, Chief AI & Product Officer at Exabeam, has done what great security leaders do: taken hard-won field knowledge and codified it into repeatable, open-source tools that any organization can run. Praxen and Observra exist because his team built them. Mitch Parker, CISO at Indiana University Health, has done the same at the framework layer: the Indiana Executive Council on Cybersecurity architecture gives every CISO a working rubric, not a position paper. Together they cover the stack, infrastructure security and AI behavioral verification, in a way that no one else has. Both co-founded TTIC. Darwin by Medigram, the worked example throughout, is the commercial implementation of the TTIC governance model. The author chairs TTIC and leads Medigram.
Start with the tool, not the theory. Agent Behavior Verification is the practice of probing an AI system's actual behavior against its intended behavior under adversarial conditions, before deployment and on an ongoing basis after it. Praxen is the open-source reference implementation, built by Steve Wilson and released under Apache 2.0 by Exabeam. It is free. A CISO can run it this week.
Praxen organizes behavioral testing into six categories, known as RAISE: domain limitation, knowledge base balance, zero trust implementation, supply chain governance, adversarial testing, and continuous monitoring.
The tactical distinction that matters for a working CISO: SOC 2, HIPAA technical safeguards, and most of ISO 27001 measure whether controls were documented and in place during an audit window. That is a different question from whether the system behaved correctly under the full range of conditions it will meet in production, including adversarial ones. A clean audit and a system that holds under attack are not the same finding.
Two things a CISO can act on immediately. First, run three successive assessment cycles and treat each run's findings as a remediation roadmap rather than expecting to close everything in one pass. Second, before the first run, define the governance committee with authority to sign off on the work remit. Without that structure, the run produces findings with no pathway to resolution. That is a clinical accountability and operational accountability decision, not an engineering one, and it has to be made before the tool runs, not after.
The most common gap we see is not bad code. It is good code in the wrong environment. A control that passes every review, cites the right files, and is implemented correctly can still fail to enforce in production if the environment it was written for is not the environment it was deployed into. That gap is invisible to documentation review. It is only visible when someone makes the control fail.
The two failure modes that matter most cannot be found by reading anything. The first is a control that is correctly configured and not enforcing, discoverable only by load testing or adversarial simulation. The second is a filter built on pattern matching that accepts a reworded version of the attack it was written to block, discoverable only by trying to break it.
Neither of these is a coding error. Both are invisible until tested.
A control that is correctly written, correctly bound, and not enforcing is indistinguishable from a working control until someone makes it fail.
The rule that follows: a review that reports a claim as verified without naming the environment it was verified against and the date is reporting that the claim is verified in code. Scope-tag every claim: verified in code, verified in staging, verified in production.
The second rule: correct in isolation, wrong in the situation it lives in, is the shape of nearly every one of these failures. Test the situation, not the artifact.
3. The Economy and the RoleWhy are senior leaders moving back onto the production surface? In agentic environments the systems are immature, the abstractions are unreliable, and the consequences of subtle errors are high. The leverage of an exceptional individual contributor has increased sharply. A technically strong executive may create more enterprise value in three hours correcting an agent architecture than in three hours managing a layer of managers. This is not a regression to task work. It is founder-level verification of the production system.
The court, briefly:
- CEO: player-coach who is also the general manager. Should not take every shot, but must recognize a broken play immediately and be able to take control.
- CTO or CIO: point guard. Reading the defense, directing agents, changing the workflow in real time, taking the difficult shot when automation breaks down.
- CISO: defensive captain on the floor, not a defensive coach supervising automated controls. Watching the whole system, calling switches, detecting abnormal behavior, testing whether controls actually work, reviewing traces and exceptions, personally investigating novel failures, and stopping play when the system becomes unsafe. Automation handles routine defense. The CISO engages personally when the opponent changes tactics or an agent behaves in a way the controls were never designed to detect.
- CFO: salary-cap strategist. Which agents are worth deploying, which workflows actually save money, where verification cost exceeds automation value, and how much operational and liability exposure is accumulating.
- GC: team legal counsel who reads every contract the franchise signed and makes sure the organization can prove it honored them. When a dispute goes to arbitration, the evidence chain either exists or it does not.
- Board: owners with a governance committee. Not in every possession, and not in the owner's suite reading the scoreboard.
Automation can handle routine defense. The CISO has to be on the floor when the opponent changes tactics.
The transition is a progression, not a break: perform tasks manually, build tools, verify tools directly, supervise tools, intervene by exception, redesign the system from evidence. Leadership moves repeatedly between abstraction levels. The new executive skill is selective depth.

Your technical knowledge still matters, but task execution alone is no longer the product. Your value increasingly comes from deciding what should be built, verifying whether it works, understanding consequences, integrating it into real operations, and accepting responsibility for the outcome. This is also why many strong technical people are returning to individual-contributor work. The highest-value role is no longer necessarily a narrow specialist or a detached manager. It may be a deeply technical operator who can use agents, inspect their work, make decisions, and close the loop.
Institutions have a responsibility here. Employers, universities, and public leaders should stop presenting STEM as an automatic economic guarantee. They should provide credible pathways from technical execution into system ownership, field operations, safety, governance, sales engineering, product leadership, and applied implementation. Otherwise, we risk creating a large population of capable people who were trained for a bargain the economy no longer intends to honor.
The economics of the role are worth stating briefly. Healthcare data breaches averaged $7.42 million in 2025 and took 279 days to detect and contain, the longest detection window of any industry (Healthcare Analytics Statistics 2026, knowi.com). AI-related class action filings more than doubled in 2024, and AI is now the largest category of event-driven securities class actions (Risk and Insurance, October 2025).
That gives a CISO a budget argument that did not exist three years ago: continuous verification is cheaper than reconstruction under adversarial conditions. The cost of discovering a problem after a decision has been made and contested is not a software cost.
4. Observra and Continuous MonitoringVerification at a point in time is a snapshot. Praxen answers whether the system behaves the way its documentation claims, at this point in time, across six behavioral categories. It does not answer whether that is still true in week fourteen.
The seasonal framing is the most relatable version of this for a CISO or a team physician. The system used in week two is not guaranteed to behave identically in week fourteen. Models update. Input patterns shift. Outputs get shorter, or longer, or start to look different, and nobody notices because nobody is watching at the right layer. In a clinical environment that drift is a patient safety problem. In pro sports it is a franchise liability problem.
Observra is an open source agent telemetry and observability framework created and open-sourced by Exabeam. It captures, normalizes, and routes runtime activity from AI agents into a standardized telemetry layer, providing a consistent record of agent behavior without custom integrations for every agent framework.
In the worked example, Darwin, Observra sits above the individual monitoring modules as the classification and routing layer, receiving findings from health checks, integrity checks, and environment parity systems, classifying each by severity, and routing it to the appropriate notification tier with automated digest and weekly summary reports. The point is what it is not: not a dashboard someone has to remember to check.
Praxen verifies that a system like Darwin behaves as claimed before it reaches a patient or an athlete. Observra confirms it still does after.
This is a requirement, not an operational preference, and four bodies with real authority say so. OWASP Agentic AI Top 10 identifies agent behavior drift, prompt injection, and ungoverned tool use as the defining risk categories for agentic systems in production. NIST AI RMF treats continuous monitoring as a core function rather than an enhancement. ISO 42001 requires continual monitoring of AI system performance and behavior under clause 9.1. ANSI/HSI 2800:2025 establishes continuous behavioral oversight as non-negotiable for clinical AI deployment.
Praxen verifies before, and on re-verification. Observra monitors after. They answer different questions, and neither replaces the other.
The Indiana Executive Council on Cybersecurity AI Security System Architecture Layers is the most operationally specific state-level AI security framework published to date. Its own stated purpose is to provide a framework and rubric organizations can use to evaluate AI systems for security and risk. That is a working instrument for a CISO, not a position paper.
The framework names six layers: Model/Agent Infrastructure Boundaries, Input Sanitization and Context Isolation, Multi-Agent Verification, Output Validation, Temporal Controls, Continuous Monitoring.
Its Basic Expectations cite IEEE/UL 2933:2024 by name for data provenance. The standard and the framework are not two separate conversations. The framework says so itself.
The framework was designed by Mitch Parker, CISO at Indiana University Health and a co-founder of TTIC. That is a relationship, not an endorsement, and the framework does not certify anyone. What it gives a reader is a framework written by somebody who runs security in the same environment their prospective vendors deploy into.
The honest scorecard from the worked example, at summary level: five of the six layers implemented as specified, one deliberate deviation, two components in flight, one named limitation.
The deviation, stated as a deviation and not as alignment: the framework calls for multiple heterogeneous agents with distinct identities running concurrently. The worked example uses a different approach for documented operational reasons. That is an argument for the instrument, not a claim of alignment.
The input filtering approach has a documented limitation. What still holds is that any attempt lands in a sealed, hashed, independently verifiable record, becoming evidence rather than a silent event. That is not prevention and should not be described as prevention.
For a CISO evaluating a vendor, the question is not whether the vendor claims alignment with a framework. It is whether they can show the work, including the parts that are not finished.
The same standard applied inward is harder and more useful. If your own controls have been verified in code and never in production, you have documentation rather than governance. That distinction is invisible until an incident makes it visible, and by then the record either exists or it does not.
Praxen is free and open source: explore it.
The shift is not from technical to non-technical. It is from execution to accountability. The people who make that transition well will not look like they changed jobs. They will look like they finally have the authority to match what they already knew.
Configuration is a claim. Behavior is evidence. The defensive captain's job is to produce the second one, personally, before anyone asks for it.
TTIC | Trustworthy Technology & Innovation Consortium
Pick one control in your environment that has been verified in code or documentation but not in production. Make it fail. Record what you find. That is the first step from governance as documentation to governance as evidence.
If your organization has not run a behavioral verification cycle, download Praxen this week. It is free and open source. Run the first cycle against your highest-consequence AI system. Treat the findings as a remediation roadmap, not a report card.
If you are a technical leader being asked to manage rather than operate, consider whether your highest-value contribution right now is organizational or on the production surface. The answer may have changed.
Run behavioral verification
Praxen is free and open source. Install it, run it against your highest-consequence AI system this week, and treat the findings as a remediation roadmap. Three successive cycles produce a baseline. That baseline is your evidence.
Install Praxen →Add continuous monitoring
Observra captures, normalizes, and routes runtime activity from AI agents into a standardized telemetry layer. Verification answers what the system did at one point in time. Observra answers whether that is still true in week fourteen.
Install Observra →Use the CISO's worksheet
The Indiana Executive Council on Cybersecurity AI Security System Architecture Layers is the most operationally specific state-level AI security framework published to date. Six layers. A working rubric, not a position paper.
Read the framework →- What is the last control you personally made fail, rather than read about?
- Where in your environment is a control correctly configured and not enforcing right now, and how would you know?
- Who signs the work remit for adversarial testing in your organization, and do they have authority over the remediation budget?
- What behavioral baseline did you record at deployment, and who owns the alert when it moves?
- When did your organization last state a security claim externally, and what environment was it verified against?
This piece was shaped in part by conversations sparked through The New CISO podcast. Thank you to Kelly Buckman (producer) and Steve Moore (host) for the thought, care, and interviewing quality that goes into that show. It helped crystallize a framework for how the CISO role is evolving: judgment, authority, execution, verification, organizational integration, and ownership of the outcome.

Image: Open Agent and AI Security Community, August 2026.
Featured by the Open Agent and AI Security Community, August 10, 2026.
This piece was spotlighted by the Open Agent and AI Security Community, which described Medigram’s Darwin deployment as having “posted the highest Praxen behavioral score recorded to date.”
Discussed on The New CISO (Exabeam), Episode 150 · Episode highlights