
The FDA Made History This Week. I Made a $25 Mistake.
This week I want to share a slightly embarrassing story — partly because it's funny, and partly because there's a real lesson buried in it for anyone working with AI outside of a chat app.
Howdy friends,
This week I want to share a slightly embarrassing story — partly because it's funny, and partly because there's a real lesson buried in it for anyone working with AI outside of a chat app.
I've been building a data collection script for biologics annual revenue. The project involves pulling specific financial values from a large stack of SEC filings, and I've been working on it in Cursor (my favorite AI-assisted coding environment) with one of OpenAI's frontier coding models. After a few days of iteration, I had the script dialed in — tested on a single year of filings, tweaked until it was extracting exactly what I wanted. Things were clicking.
Quick aside for those less familiar: when you use AI through an API — meaning you're calling a model directly from code rather than chatting with it in a subscription app — you pay per token. This is why you hear so much about AI labs releasing smaller, highly capable versions of their biggest models. Not every task needs the most powerful brain available, and the cost difference can be enormous.
During my test runs, I noticed that the coding model helping me write the script wasn't aware of OpenAI's latest cost-effective model releases (ironic, since I was using an OpenAI model to write the code). Its pre-training cutoff meant it didn't know about gpt-5-nano*, which was perfect for my use case — small, cheap, and more than capable of reading SEC filings. After I manually pointed the script to the right model, my full test run across a year of filings cost a grand total of $0.67. Beautiful.
Feeling confident, I decided to backfill two additional years of revenue data. Double the stack of filings. I kicked off the job and went to eat dinner.
I came back to find I had burned through my API credits.
In a panic, I checked my usage dashboard: $25 spent in under an hour. Thank God I had set a spending limit — otherwise the job would have just kept running until it was done, and the bill would have been significantly worse.
So what happened? When I told the coding assistant to run the script on the new batch of files, the context window had cleared and a new chat session started. That meant the model reverted to the default configuration in my codebase — the older, more expensive model I'd been using earlier. I had never removed the reference to it. The result: each token was more than 10x more expensive than my previous run, and I didn't notice because I was already eating pasta.
I sat there staring at my screen, trying to figure out what I'd actually learned
The first lesson was obvious: don't stop thinking critically just because your AI workflow appears to be cruising. And I get it — it is so hard to stay vigilant, especially when things start clicking and you can really feel the advantage of working with these tools. But that's exactly when you're most vulnerable to a dumb, expensive mistake.
The second lesson is a bit more nuanced. For those of us working with AI and trying new things every week — this is messy. It's going to continue to be messy for a while, and we shouldn't be surprised by it. I like to think of myself as one in a legion of researchers helping humanity figure out this fascinating technology. The ones who get their hands dirty, mess things up, and yes, make costly mistakes — those are the people who are going to drive this forward.
I invite you to do the same. Just keep an eye on your usage dashboard.
Thanks for reading!
Alexa
*Side note: Even Claude Opus, a front runner for amazing reasoning models and the one I use to assist in drafting my newsletter “corrected” my reference to gpt-5-nano and changed it to gpt-4.1 — um, excuse me bro.
News

FDA's Elsa AI Tool Forced to Swap Models — What It Means for Sponsors
What's New: The FDA's internal AI assistant Elsa — built on Anthropic's Claude and used by reviewers to summarize submissions, flag data gaps, and support scientific review — is being forced to migrate to Google's Gemini. This isn't a planned upgrade. On February 27, 2026, President Trump directed all federal agencies to stop using Anthropic's technology following a dispute between Anthropic and the Department of War over autonomous weapons and mass surveillance. The Department of Health and Human Services has already told employees to stop using Claude, and internal FDA communications confirm Gemini is being positioned as the replacement within Elsa.
How It Works:
Elsa is not a simple chatbot — it's a custom retrieval-augmented generation (RAG) system built by Deloitte, originally evolved from CDER-GPT, and deployed in AWS GovCloud (FedRAMP-High).
The system integrates FDA-specific document stores, embedding models, vector databases, and retrieval workflows — all originally built around Claude's behavior.
Swapping the underlying model is not plug-and-play. The entire RAG pipeline — how documents are retrieved, interpreted, and summarized — needs to be re-engineered and revalidated for Gemini.
If Gemini can't run within AWS GovCloud, the migration may also require moving to Google's own FedRAMP-High cloud, introducing new data residency and security validation questions.
Anthropic is challenging the supply chain risk designation in court, meaning the transition could potentially be paused or reversed — adding yet another layer of uncertainty.
Why It Matters:
This is a big deal for sponsor companies with active or pending FDA submissions. First, there's the review consistency question: if a submission is reviewed partly under Claude-powered Elsa and partly under Gemini-powered Elsa, the administrative record may contain inconsistent AI-generated analyses. Legal experts have flagged this as a potential vulnerability under the Administrative Procedure Act. Second, the speed and political nature of this migration means sponsors should not assume this is a controlled, fully validated process. FDA employees have expressed concern internally — one anonymous reviewer told reporters that the switch would effectively erase 18 months of training and prompt development efforts. Third, while the FDA maintains that Elsa does not train on sponsor submissions, the shift to a new model and potentially a new cloud environment means data handling, prompt logging, and metadata generation controls all need to be revalidated. Sponsors are being advised to mark sensitive content as Confidential Commercial Information, request written clarification on which model is being used for their reviews, and document everything.
My Take: I want to be careful not to catastrophize here — there are some breathless takes floating around, and the reality is that Elsa was always positioned as a productivity tool for reviewers, not an autonomous decision-maker. The FDA has been clear that AI doesn't replace human judgment in their process. But the disruption is real, and the timing is not so good. I can imagine the agency had been building genuine momentum with AI-assisted review, and a forced, politically-motivated model swap in the middle of that ramp-up introduces exactly the kind of uncertainty that slows things down. For those of us in the life sciences space who have been advocating for thoughtful AI adoption, this is a reminder that the technology is only as stable as the ecosystem around it.
Sources: Clinical Leader — FDA's Elsa AI Switches From Claude To Gemini | NOTUS — HHS Tells Employees to Stop Using Anthropic's Claude | Fierce Biotech — HHS Bans Claude AI Tool
FDA Launches AI-Powered Adverse Event Monitoring System, Replacing Seven Legacy Databases
What's New: On March 11, 2026, the FDA launched the Adverse Event Monitoring System (AEMS) — a unified, AI-assisted platform that consolidates the agency's fragmented patchwork of adverse event databases into a single real-time dashboard. AEMS replaces legacy systems including FAERS (drugs and biologics), VAERS (vaccines), and AERS (animal products), with additional systems like MAUDE (medical devices) and the Human Foods Complaint System scheduled to migrate by the end of May. The agency processes roughly 6–7 million adverse event reports per year and was previously running them through seven separate databases at a cost of about $37 million annually.
How It Works:
AEMS provides a single streamlined dashboard where adverse event reports for drugs, biologics, vaccines, cosmetics, and animal food can all be searched and analyzed together.
The platform uses AI for two key functions: automated redaction of personally identifiable information and digitization of reports — both intended to speed up case processing and improve data accuracy.
Reports are now published in real-time rather than the previous quarterly cadence, which is a major shift for anyone doing postmarket surveillance.
The system also incorporates standardized reporting protocols across all product categories and includes advanced analytics tools and APIs for external researchers.
Notably, AEMS expands beyond traditional adverse event reporting to also centralize consumer complaints, regulatory misconduct reports, and whistleblower submissions across all FDA centers.
The FDA reported a 3,000% increase in users during a pilot program launched last September, and estimates AEMS will save approximately $120 million over the next five years.
A front-end submission tool is also in development — the agency estimates that 80% of adverse event reports are never filed due to the complexity of the current reporting process.
Why It Matters:
For sponsor companies, CROs, and anyone in pharmacovigilance, this changes the tempo of postmarket safety monitoring. When adverse event data was published quarterly, there was a built-in lag between signal emergence and visibility. Real-time publication compresses that window significantly, meaning sponsors need to be monitoring more continuously and responding faster when patterns emerge. The cross-product surveillance capability is also worth watching — with all product categories in one system, it becomes easier for the FDA (and external researchers) to spot safety signals that span product types or identify trends that would have been invisible across siloed databases. For QA and regulatory teams, the standardized reporting protocols should eventually reduce administrative burden, but the transition period will require attention as legacy data is migrated and new workflows are established.
My Take: This is genuinely impressive infrastructure work from the FDA, and the timing makes for an interesting juxtaposition with the Elsa model-swap story above. On one hand, the agency is dealing with political disruption to its AI-assisted review tools. On the other, it just shipped what its Chief AI Officer called the biggest technical transformation in agency history. The real-time reporting shift is the headline, but I think the sleeper story is that 80% non-filing rate. If the new front-end submission tool actually makes it easy for healthcare professionals and consumers to report adverse events, we could see a dramatic increase in report volume — which is great for safety surveillance, but also means sponsors need to be prepared for more noise alongside the signal. If you're in pharmacovigilance, now is the time to make sure your monitoring processes can keep pace with real-time data. The quarterly rhythm we've all been working around just ended.
Sources: FDA Press Release — FDA Launches New Adverse Event Look-Up Tool (March 11, 2026) | Pharmaceutical Commerce — FDA Vows Improved Adverse Event Reporting With AEMS (March 12, 2026)
Get the next edition in your inbox
Clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.
No spam. Unsubscribe anytime.


