
I went to update a few slides...and rebuilt the deck
How much has AI grown in the last 12 months?
Hey friends,
Last week I kicked off a new and refreshed life science consultant bootcamp. It doesn't feel like it, but the last one I ran just for consultants was exactly one year ago. While preparing for this one, I figured I'd need to update a few facts and figures to reflect the current state of AI tools available to the public. Then I realized I needed a full-blown presentation overhaul. It's staggering how much of the landscape has changed in twelve months. The review was eye-opening for me, so I thought I'd share some of the most interesting facts about AI capability and how its effects are being felt well beyond the tech world.
A look back at the last twelve-ish months (August 2025 to August 2026)

About a year ago, the leading labs released models that started to change the game for knowledge workers and for anyone who wanted to write code, including non-programmers like me. For me, this is the most exciting use case. Like many others, I've spent countless hours building webpages, apps, and small programs for myself and my colleagues to make tasks faster and higher quality.
While many professionals and their organizations were still trying to get access to any AI tools for work, the market was reacting to the potential of the technology. Anthropic was valued at around $183 billion a year ago; this spring it raised money at $965 billion, and OpenAI isn't far behind. User adoption was on the uptick, and twelve months ago most of us weren't considering that limits of the physical world (compute, power, chips) would become the next puzzle to solve.
To anyone who likes to keep repeating that AI is not such a big deal
Social media has been a fun place to point out the wacky things that come out of AI use, including hallucinations and letter-counting snafus. Those findings matter too, because it turns out that whatever users complain about, the problem is almost guaranteed to be improved, if not solved, within the next model release or two. At the very least, problems are usually an area of active research.
Hallucination rates still aren't zero, but they've dropped a whole bunch. That makes them trickier to spot, which is exactly why human review is the skill we need to keep sharpening. And the real skill with AI is increasingly about deciding what to use it for, because "everything" is not the answer.
How fast is fast when it comes to model improvement?
One of the better measurements we have comes from METR, a nonprofit that tests how long a task a model can finish on its own, measured in the time it would take a human expert. When I ran the first bootcamp, the best models could reliably handle a task of about two to three hours. This spring, Claude Mythos Preview crossed sixteen hours. What’s wild about that is that the number comes with an asterisk; METR says its test suite can't reliably measure anything beyond that. The pace has picked up too: from 2019 to 2024, that task length doubled roughly every seven months. Since 2024, it's been doubling every three to four.

METR running out of long-horizon tasks to test AI with was not on my 2026 bingo card.
Model capability and the government's stance
The road has been bumpy this year when it comes to model capability and national security, and it goes to show that the speed of this technology means none of us, not even the AI labs, know exactly what we're doing or the best way to manage it. A year ago, the US government was all about deregulation. Hit fast forward, and executive orders on frontier models are being signed because capability has reached a point where oversight has to at least be considered.
What we do know is that we can't put the genie back in the bottle (though some will try), and even if models stopped advancing today, they've already changed the world dramatically.
What does the next twelve months hold for us? Not a clue. I just know it'll be crazy.
Thanks for reading!
Alexa
News
FDA Opens the Conversation on Regulating Generative AI Medical Devices
What's New:
On August 18, the FDA's Center for Devices and Radiological Health released a discussion paper on how it might regulate generative AI-enabled medical devices, covering risk assessment, premarket evaluation, postmarket monitoring, and the tricky question of devices built on third-party foundation models. Comments are open through October 19, 2026, under docket FDA-2026-N-7874 on Regulations.gov.
How It Works:
A two-axis risk framework. Risk is assessed along two dimensions: what the device does (from non-directive information, to action-directing information, to HCP-supervised action, to fully autonomous action) and the consequences of relying on an incorrect output (limited to severe). A chatbot suggesting hydrocortisone for poison ivy and one directing an insulin dose change both give advice, but they land in very different places on the grid.
Competency-based premarket evaluation. Instead of exhaustively testing every possible input (impossible for open-ended GenAI), CDRH is floating an approach inspired by how human clinicians are credentialed: structured "benchmarking" of the device's knowledge, safety behavior, and generalizability, followed by "clinical confirmation" in real or representative conditions. Notably, clinical confirmation might not require a prospective clinical study in every case; options range from retrospective evaluation and shadow deployment up through RCTs, scaled to risk.
The final device gets evaluated, not the foundation model alone. Testing would cover the deployed configuration, including prompts, guardrails, and orchestration, not just the underlying model.
Heavier reliance on postmarket monitoring. CDRH is considering accepting greater premarket uncertainty in exchange for robust postmarket surveillance: periodic re-benchmarking, sample-based clinician review, and drift monitoring, potentially even using machine-based supervisory agents to do some of the watching.
Foundation Model Master Files. A proposed voluntary mechanism where model developers (think Anthropic, OpenAI, Google) could submit confidential model information directly to FDA, which device sponsors could then reference in their submissions, addressing the problem that device makers often can't see inside the models they build on.
Why It Matters:
The document reads like a straight port of concepts our industry already knows. Benchmarking against prespecified acceptance criteria is qualification testing. Re-benchmarking after a change is revalidation. The Predetermined Change Control Plan is change control. Drift monitoring is continued process verification. The Foundation Model MAF is literally the Drug Master File concept applied to AI vendors. For anyone working in CMC, QA, or regulatory, the vocabulary FDA is reaching for here is our vocabulary, which means quality and regulatory professionals are better positioned to engage with this framework than they might assume. It also signals how FDA may eventually think about GenAI tools used elsewhere in the product lifecycle, not just in devices.
My Take:
What caught my eye here was the approach to premarket evaluation. The argument I keep hearing against generative AI in any regulated application is "it doesn't give you the same exact answer every time." What they really mean is the same exact words every time; the answer itself is often the same, or equally valid, and that's the wrong way to think about it. Ask two different experts a question and their answers won't match word-for-word either, yet both are credentialed to give advice on the topic. FDA seems to have noticed the same thing, and building the evaluation around competency rather than exact-match outputs is a meaningful shift. This topic still has to be handled with care, but we need to make sure we're looking at the correct mechanisms for control and trust, just like we've trusted the clinicians who have treated us humans for decades.
Source: FDA: Considerations for the Regulation of Generative AI-Enabled Medical Devices
Get the next edition in your inbox
Clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.
No spam. Unsubscribe anytime.


