
Is the FDA making AI do their job (and handing us the bill)?
Something weird is happening at the FDA, and I bet it’s related to AI.
Howdy Friends,
Something weird is happening at the FDA, and I bet it’s related to AI.
If you haven't been watching the agency as closely as I have, let me catch you up. Over the last 18 months they've gone all in on AI: a January 2025 guidance laying out a risk-based framework for AI in regulatory decision-making, a scientific-review pilot that spring, the June 2025 launch of their internal "Elsa" tool for reviews and inspections, updated guidance on good AI practice in early 2026, a real-time clinical data submission pilot in April, and the Elsa 4.0 platform in May. You get the idea. The FDA is more all-in on AI than a hyped-up biotech CEO raising a seed round for his three-person startup out of a garage.
But behind every flashy announcement, I've been trying to watch how the agency actually executes in the real world. So I was beyond intrigued when a colleague called me last week about a pattern of strange stuff showing up in FDA comments—across multiple products, multiple sponsors, and different parts of the submission process. "New" requests for control strategy. Stability testing nobody had been asked for before.
Here's why that's weird. Consultants who have been at this a long time have seen just about everything there is to see in getting a BLA or NDA out the door. They often personally know the reviewers: how they work, which areas they dig into, and the recurring requests that are pretty much guaranteed to come. A good CMC/RA consultant is a sherpa, guiding sponsors across the harsh, snowy peaks of the submission process. They've walked these trails a hundred times. And now it's like the trails have shifted. New boulders in the path—some big, some small—that they've just never had to climb over before.
So the first question I get is: "Is FDA using AI for this?"
Let's be real, we all know the answer. And those of you who've read my stuff know I have no problem with AI use—I encourage it where it fits. What I have a problem with is people using AI irresponsibly: without thinking critically, without a "license". So when someone asks me "Is FDA using AI?" what they're really asking is: "Are they outsourcing their critical thinking to a machine, kicking back, reducing their workload while increasing mine?"
I'll admit it—for a hot second, I was alarmed. The news is full of exactly this across every industry. Part of it is that none of us really knows how humans are supposed to work with AI yet, because the technology is improving faster than we can update mental frameworks we've been building for thousands of years.
Then I looked closer. I went hunting for published evidence behind these "new" comments—the additions to control strategy, the extra testing, whatever. Can you guess what I found? Real, recent research supporting the agency's concerns. In other words: dang, the science appears reasonably sound. And it was recent—the oldest paper I found was from 2019, which might be exactly why seasoned consultants have never seen these requests before. The basis for them didn't exist the last time they walked this trail.
So if the FDA isn't "lazily outsourcing" the way I suspected for a moment, what's actually going on?
My guess is that this is what it looks like when everyone in the process—reviewers and consultants alike—gets augmented by AI at the same time, and the patterns we've relied on to work efficiently with each other start to break. The reviewer's effective knowledge just jumped past the historical baseline that consultants spent decades calibrating against. The expertise is still needed, but the map that expertise relied on is suddenly less reliable.
And this isn't just a biopharma thing. Take the patent office. In late 2025 they started piloting AI tools that search for prior art—older patents and publications that can sink a claim. Suddenly examiners are surfacing references they never would have found on their own, and patent attorneys are getting rejections in places they didn't see coming, outside the patterns they'd built their drafting strategy around for years. The examiner still makes the final call. Their reach just got a lot longer.
Same story in science. Journals and conferences are scrambling to figure out AI-assisted peer review, and the recurring theme is that once a reviewer's toolkit changes, the predictability authors had counted on goes out the window, and nobody yet knows what the new normal looks like.
So whats my final take? No, I don't think the FDA is cheating with AI. But yes, they're redrawing the map faster than we can keep up. So what do we do about it?
The honest version of how this went: the "are they lazily outsourcing?" alarm was mine, not his. My colleague, who's been at this longer than I have, didn't jump there at all. He asked two much better questions, and they're the ones I’d like address.
First: does the FDA know something we don't?
Historically, when the agency started asking for something new, it traced back to a real-world signal, like a product failure, a recall, or an adverse event pattern that showed up in the field. The new request was the agency reacting to something that actually happened. What's different now is that the trigger seems to sit upstream of any failure. The concern is coming from the literature, not the field. The agency isn't reacting to a problem that surfaced in a marketed product, it's surfacing a problem that recent research suggests could exist before anyone's seen it go wrong.
That's a real shift in where new comments come from and it changes how you should prepare. You can't wait for the field signal anymore, because the field signal isn't always going to be the source. The science is, and the science is moving faster than before.
Second: is there something we can do to prepare, if this keeps up?
Yes, and it's the same move the agency just made. They augmented their reach into the literature, and we can do exactly the same thing on our side of the table, before we file.
It has never been easier to run deep research on a control strategy and stress-test it for gaps, to ask what the recent literature says someone might worry about here before the agency asks it for us. The old preparation was pattern memory: known reviewers, known requests, and all those boulders you'd climbed over before. That's still valuable, but it's depreciating, because the map is being redrawn from sources that didn't exist last time we walked the trail. What holds up is the habit of re-surveying, checking the current science against your strategy every time, even on the hundredth climb.
If you want it in a simplified version: When it feels like the other side suddenly has the upper hand, it's usually not because they're trying to overpower you. It's that they're augmented and you're not yet. The discomfort is the asymmetry, and the fix isn't to resent the new boulders, it's to go find them first.
That's what keeps a sherpa useful as the mountain changes. Not having memorized the old trail, but being the one who keeps re-walking it with fresh eyes, and now, fresh tools.
Thanks for reading!
Alexa
News

Anthropic's Most Powerful Model Goes Public, With a Catch for Life Sciences
What's New: On June 9, Anthropic released Claude Fable 5, the first publicly available model from its "Mythos" class, the most capable tier the company makes. Fable 5 is a version of Mythos, the model Anthropic had previously said was too dangerous to release, now wrapped in guardrails that block responses in high-risk areas like cybersecurity and biology. For those of us in life sciences, that biology guardrail is the part to pay attention to. The model is excellent, reportedly compressing months of engineering and research work into days, but the safety system sitting on top of it can get in the way of legitimate technical work, sometimes invisibly.
How It Works:
The safeguards aren't baked into the model itself. Separate classifier systems inspect every request before Fable 5 responds, and when one spots a query about cybersecurity, biology, chemistry, or distillation, the response gets handled by Claude Opus 4.8 instead (the prior top model).
Anthropic says the fallback kicks in for fewer than 5% of sessions, and openly admits the classifiers are tuned cautiously, so benign requests will sometimes trip them. Reducing those false positives is a stated post-launch priority.
There's a separate track for researchers: selected life science researchers can receive Fable 5 with the biology and chemistry safeguards removed, though the cyber restrictions stay in place for that group.
On pricing and access: Fable 5 is included at no extra cost on Pro, Max, Team, and seat-based Enterprise plans through June 22. On June 23 it comes off those plans, and continued access runs through usage credits billed at API rates. Anthropic says it intends to restore it as a standard plan feature when capacity allows, but has named no date.
One wrinkle worth knowing: within two days of launch, Anthropic reversed course on an invisible guardrail, one aimed at preventing users from training competing models, that silently altered prompts to produce faulty results rather than refusing outright. After user backlash, the company agreed to make that safeguard visible.
Why It Matters: This is the tension we keep circling in this newsletter, now shipped as a product. The most capable model available has guardrails that specifically target biology and chemistry, the exact domains a lot of life science work lives in. In some cases, when the model classified a query as sensitive, it would provide a lower-quality answer without telling the user it had downgraded. If you can't see when the tool quietly switched to a weaker model, you can't account for it in your verification, which is precisely the kind of silent degradation that's hard to catch and harder to document in a regulated workflow. The flip side: for non-sensitive work like drafting, summarizing, analysis, and coding, you're getting access to a genuinely frontier model on a normal subscription, at least for a couple of weeks.
My Take: I'd try it before June 22 while it's free on the paid plans, with eyes open. For the bulk of knowledge work it's an impressive tool and worth forming your own opinion on. But know going in that if your prompt brushes up against biology or chemistry, you may get bounced to a different model or a quieter, lower-quality answer, and you won't always be told. That's not a reason to avoid it, it's a reason to use it the way we talk about using all of these tools: knowing what tool you're actually holding at any given moment. If you do serious bench or CMC work and find the guardrails genuinely in the way, the researcher-access track is worth looking into rather than fighting the classifier. After the 22nd it shifts to usage-based pricing, so the free window is the low-stakes time to take it for a spin.
Source: TechCrunch — Anthropic releases Claude Fable 5 | Anthropic — Claude Fable 5 and Mythos 5 | Aardwolf Security
Get the next edition in your inbox
Clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.
No spam. Unsubscribe anytime.


