
Where compliance breaks down when using AI for work
How to right-size your AI oversight
Hey friends,
I gave a talk last month at the DFW RAPS meeting called AI Fluency as a Compliance Control, and if you weren’t able to catch it live, here’s my main point.
Most of the AI risk in regulated work isn't where companies are looking for it.
If you ask a Quality or Reg leader at a life sciences company how they're managing AI, you'll usually hear about vendor risk assessments, IT security reviews, and the bolierplate “acceptable-use policy” that everyone (and I mean everyone- GMP and non-GMP employees, admin, finance, the guy who answers the phone, etc.) signed during onboarding and nobody has opened since. Don’t get me wrong, this is all real work, and all necessary. But none of it actually controls what happens at 4pm on a Tuesday when someone is racing to finish a deviation summary and decides to paste it into ChatGPT.
That gap, between the policy on paper and the practice at the desk, is what I've started calling the last mile. It's where a lot of the risk actually shows up, and it's almost completely uncontrolled in most organizations I see.
The reason is structural. Governance and procurement operate at the level of the tool: did we vet it, is it approved, what's the data agreement? But the failure modes that get you in trouble aren’t always at the tool level, they’re at the use level. A perfectly well-vetted tool used poorly is the same risk as an unapproved tool used poorly, and maybe worse, because the institutional comfort with "approved" means nobody looks twice.
So what does "used poorly" actually look like in practice? In the talk I broke it into five failure modes:
Confident errors, where the output sounds great and is wrong.
Missing traceability, where you can't reconstruct where a claim came from.
Over-reliance, where the human stops actually reading.
Data boundary violations, where something sensitive ends up somewhere it shouldn't.
And inconsistent practice, where two people doing the same task get there in completely different ways (which might be fine) but neither is documented.
The thing to notice is that none of those are about the AI being bad. The AI is just doing what AI does. The exposure is created by what we are or aren't doing on either side of it — the workflow around the AI, which in most cases is completely undefined.
Which brings me to the part I think regulated organizations are going to have to come around to. For regulated work, AI fluency has to be more than a personality trait or a sign that someone is "good with tools." It's a competency, the same kind of competency as aseptic technique or batch record review. Defined, trainable, assessable, periodically reviewed. If you cannot say what good AI use looks like for a given role and a given task, you cannot say whether anyone on your team is doing it well, and you certainly cannot say so to an auditor.
That sounds like a lot, but it really isn't. It comes down to standard work: maybe it’s an approved prompt structure for a given task, a defined boundary for what data goes in, a verification checklist sized to the risk of the output, and a lightweight record of what happened. Most of this looks suspiciously like the QMS habits regulated teams already have. The difference is applying them to AI-assisted work the same way we apply them to everything else.
The risk-tiering part is where this gets practical. Not every AI use needs the same level of oversight, and acting otherwise is how you either stall adoption entirely or invite drift everywhere. Two questions get you most of the way there: what's the impact if this output is wrong, and how close is it to a controlled deliverable?
Here are some examples:
Internal brainstorm? Go nuts, self-review is fine (within data boundary control of course).
Drafting something that will go inside a regulatory response? Someone better be reviewing every word (at least with today’s processes).
Most real work lands in the yellow zone in the middle, which is exactly the zone where structured verification matters most and is least likely to be happening.
The “last mile” is ties very nicely to the expertise piece I was on about last week. The verification has to happen there, in the last mile, by the person doing the work, with the domain knowledge to know when the output is off. You can't push that upstream into procurement and you can't push it downstream into a QA audit six months later. It has to be the practitioner, in the moment, with the right standard work and the judgment to use it.
That's the whole reason AI fluency is a compliance question and not just an IT question. The control lives in the human doing the work, which means the control either exists or it doesn't, depending on what that person actually knows how to do.
This is exactly what AI-Proof Your Practice is built around. Half-day virtual workshop on May 29. We work through the failure modes in the context of your actual work, and you leave with the start of your own standard work, not just a list of principles.
Thanks for reading!
Alexa
FDA Wants to Watch Your Trial Happen, Not Read About It Later
What's New: On May 15, the FDA held a virtual info session on its Real-Time Clinical Trials (RTCT) pilot program. About a thousand of us showed up (they had to buy more Zoom licenses). The session followed the agency's April 28 announcement of two proof-of-concept RTCTs already running with AstraZeneca and Amgen, plus an RFI seeking industry input on a broader pilot launching this summer. Comments on the RFI are being accepted through May 29, 2026, with final selection criteria in July and pilot selections in August.
How It Works:
RTCT replaces the traditional "site → sponsor → analysis → submit → wait" cadence with a continuous feed of pre-agreed signals from the trial to the FDA, ideally within 24 hours of data capture.
The pilot is scoped tightly: phase 1 and early phase 2 only. Late-phase, blinded, or formally hypothesis-tested trials are out of scope for now.
The FDA is explicit that it does not want patient-level data. Jeremy Walsh, FDA's Chief AI Officer, said this multiple times during the session. The agency wants signals, defined upfront in partnership with each sponsor.
Sponsors submit, not technology vendors. Sponsors pick their own platform (homegrown or a partner) that can meet the real-time transmission requirements.
Pilot trials are scoped to run roughly 3 to 6 months so the FDA can iterate on the framework quickly.
The proof-of-concept trials are already piping JSON data into an FDA API pulled from sponsor EDCs.
Why It Matters: This is the regulatory architecture catching up to how trials could actually be run with modern tooling. The current system was built for paper-era data flows, and everyone knows it. What's notable about RTCT isn't the technology, which is honestly not that exotic. It's the regulatory posture. The FDA is essentially saying: we'll define what "good" looks like by running it, not by writing guidance about it first. For sponsors, that means the way you instrument your trial and structure your data flows is starting to have regulatory weight, not just operational weight.
My Take: I respect that the FDA is willing to learn by doing here, because the alternative is another decade of guidance documents about a problem everyone already understands. The "signals not patient data" line is the right call and they were disciplined about repeating it. What I'm watching for is the gap between the agency's vision and what sponsors are actually ready to deliver. Real-time data flows assume a level of internal data discipline that a lot of organizations do not have, and you cannot retrofit that in 3 to 6 months. The pilot will probably end up being a useful filter for who's actually serious about this versus who just wants the optics of participating. If you're thinking about submitting, the RFI deadline is May 29 — that's two weeks out, and the session made it clear they're not extending it.
Get the next edition in your inbox
Clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.
No spam. Unsubscribe anytime.


