All editions
    Share this edition:
    Your prompt is doing less than you think

    Your prompt is doing less than you think

    Six things are shaping your AI output, but most of us are only watching two.

    Hey friends,

    This week I want to cover a question I get consistently (which means it must be a good question): "how do we know if our AI output is using the right information?" And I'm going to answer it in my annoying way, by telling you that's not the question you should be asking (sorry).

    Here's why I think it comes out that way. A lot of us are still treating AI as a magical, unknowable box that lands on the right answer most of the time because of everything it absorbed in pretraining. If you ask an AI model about a recent agency guidance or regulation, there's an increasingly good chance the answer is correct, but as it stands today, it is still not guaranteed. The methods of training (pre-, post-, and whatever else goes on as models get better) have improved the general intelligence of the frontier models so fast that governments aren't sure what to do with them now. That's not a personal problem for me at the moment, so I'd rather put my energy into figuring out the best way to get AI outputs that actually help me with the tools I have access to. Smart models are great, but the key to the best outputs is still the context and not so much how many trillions of parameters went into pretraining. That's what I want to talk about today.

    So, "how do we know if our AI output is using the right information?" Well, we know because we brought the context to the interaction, and we don't make a habit of relying on training data for high-risk or important tasks.

    So the question we should be asking is, "did I bring all the pieces of the puzzle to the interaction?" You see what I did there? It's up to the human to make sure the correct data is available for the task at hand. Way less opportunity to blame the AI for hallucinations, because it has no need to fabricate an answer when you already brought the answers.

    And yes, I understand that regulations and guidances are public data on the internet, and it's almost certain that the latest models have trained on very recent versions of that data. I still say bring the regs to the conversation. You want your context as close to the chat as possible, and it gives the human a chance to know what the latest actually says, so the two of you are singing off the same sheet of music (I often refer to this as 'evidence-first drafting').

    But what about web search? That's good if we have it, but it's still not a guarantee that the most official source was the one pulled in to answer the question. Do you believe every interpretation you read in a blog? Those are on the internet too, where the model can find them. I'd argue AI isn't nearly as skeptical as I am when it goes looking for sources. Call me paranoid, but I don't believe most people, including myself. All I can do is gather sufficient evidence to get to a good enough answer, understanding that even that answer is temporary and may be immediately disputed.

    If I haven't lost you yet, fabulous! Let's get into what I want to cover today, which is the types of context we use when working with generative AI models. There's a lot more to it than what you put in the prompt and even the documents you upload. So here we go.

    Task and role. This is usually what we picture when we sit down to write a prompt: what you want done, and who you want the model to be while it does it. It isn't necessary to explicitly state the task and role these days, but both do need to be inferable from what you wrote. Role tends to shape tone and vocabulary, and task shapes the structure and the depth of the answer. All I'm saying is, be an effective communicator. If a new human contractor would have to come back and ask you three clarifying questions before starting, so will the model.

    Attached documents. The source material you hand over: the guidance PDF, the protocol, the draft deviation, last year's report. This is the highest-value context you control, because it turns a general-knowledge question into a reading comprehension question, which is a thing these models are very good at. Scope matters here too. Twelve documents when three would do doesn't make the answer more accurate, it just gives the model more to weigh and more chances to pull from the wrong one. Bring what's relevant, and be able to say why each piece is in there.

    Instructions. There are two layers and they're both always on. System instructions come from whoever built or configured the tool, and they sit underneath the whole conversation setting defaults for behavior, format, and boundaries. User instructions are yours: the standing preferences you've saved, project-level instructions, and whatever you typed just now. Most people never look at the system layer, which is exactly why an enterprise-configured tool behaves differently from the consumer version of the same underlying model (which is well-known source of personal frustration when I talk about tools like MS Copilot). Knowing what's in your instruction layer explains a lot of "why on earth did it do that?"

    Conversation history. Every prompt and every completion gets added to context over time. That's mostly a good thing, because the model accumulates your situation and stops needing to be re-briefed every five minutes. It's also how a chat starts to drift, because a correction you made forty messages ago, an early assumption you've since abandoned, and that tangent about something unrelated are all still sitting in there carrying weight. When an answer starts feeling stale or slightly off, the history is usually the culprit, and starting a fresh chat is often faster than arguing with the old one.

    Retrieved context. Anything a tool went and pulled in on your behalf, whether that's web search results, a connector reaching into SharePoint or Drive or your email, or a RAG system searching a document repository. You didn't hand-pick any of it. The retrieval step did, based on how you happened to word things, and then it handed the results over as plain text. The model does not inherently know that an agency page is more authoritative than a blog post that used similar words (though it might know if that’s in the standing instructions). If retrieval is doing the choosing, you should be checking what it chose.

    Artifacts and output documents. Whatever the model has produced and kept around: canvas documents, generated files, code, and earlier drafts you asked it to revise. These stay in the mix and the model starts referring back to its own work as if it were source material, which is where 'context creep' comes from. More on that next week.

    Only the first two are the ones most of us are consciously managing. All the others are contributing and accumulating right alongside them, and it's important to remember that.

    Gone are the days of "prompt templates," because that's not going to cut it anymore, at least not for real or important work. It's shaping up to be a competency rather than a cool trick for the tech savvy.

    So the next time you catch yourself wondering whether the output used the right information, back up one step and audit what you handed it. What did I attach? What's already been said in this chat? What did the tool go get on its own while I wasn't looking? And what's still floating around in here from an hour ago? If you can answer those, you can defend the output. If you can't, then it was never really an AI problem, it was an evidence problem, and that one has always been ours to solve.

    Thanks for reading!

    -Alexa


    mhra logo

    MHRA Says AI Slop in Inspection Responses Has Gone From Theoretical to Real

    What's New: On June 29, the MHRA Inspectorate published a blog post titled "Use of AI for GXP inspection responses: setting standards without stifling innovation," and it is essentially a field report on what AI-generated material looks like when it arrives at a GxP compliance team's desk. The inspectorate says the inappropriate use of AI has "shifted from theoretical risk into actual, realised risk." MHRA isn't banning AI in responses, but it has spelled out what happens when one comes back wrong.

    How It Works:

    • The named problems are specific: references to MHRA guidance that does not exist, citations of regulatory frameworks that don't apply to the situation, and responses to serious deficiencies that appear written to obscure the underlying problem rather than fix it.

    • One cited response ran over 90 pages and still failed to address the deficiencies it was answering.

    • In at least one case with patient safety impact, an AI-generated response contained material inaccuracies and non-existent references. Review time went from roughly 4 hours to over 20, and it pulled in a multidisciplinary team plus a full re-analysis of prior MHRA guidance to untangle.

    • The expectations MHRA restates are not new: submissions must be factually accurate, verifiable, technically reviewed by experienced people, signed off by an accountable individual, and supported by evidence.

    • What is new is voluntary disclosure. You can identify which sections were AI-assisted and confirm human verification took place.

    • Consequences for getting it wrong scale up from there. Poor, inaccurate, or overly generic responses may be rejected or returned for revision, treated as a marker of higher inspection risk, or escalated for regulatory action.

    Why It Matters: A deficiency response is a compliance record, and the credibility of the whole exchange depends on the regulator being able to trust what you send back. What MHRA is describing is that trust being spent down: a reviewer who used to spend four hours now spends twenty, because they can no longer assume a citation is real. That cost doesn't stay with the one firm that submitted the bad response, it gets spread across everyone in the queue behind them. And the escalation path goes further than a returned document, because "marker of higher inspection risk" means a bad response changes how the agency looks at you going forward. The voluntary disclosure option matters too. Right now it's optional, and optional disclosure mechanisms in this industry have a way of becoming expected practice once enough people use them.

    My Take: With the arrival of do-it-for-you submission writing tools, we're now running a real risk of slop in the submission and response space, which is about the worst place to put it. We need to resolve to be intentional about AI use here, and I'm not convinced we'll get it right the first time. For example, that 90-page response. Nobody set out to write 90 pages of nothing. Somebody had a deadline, a tool that produces volume on demand, and no checkpoint between the draft and the send button. That's a process problem, and the fix is the boring stuff we already know how to do: scope what the tool is allowed to draft, make somebody technically competent own the review, and be able to point at the evidence behind every claim. If you can't do that, don't send it.

    Source: MHRA Inspectorate blog (June 29, 2026) | ECA Academy / GMP-Compliance | Pink Sheet (July 6, 2026)


    saradasish pradhan OtolOiM9V0Q unsplash

    India's Regulator Is Putting AI on the Inspector's Side of the Table

    What's New: At the Global Drug Regulatory Conclave in New Delhi on July 29-30, Drugs Controller General of India Dr. Rajeev Singh Raghuvanshi said CDSCO is piloting AI across its regulatory processes to build a faster, more data-driven system. The same week, CDSCO issued guidance clarifying how software-based medical devices, including AI-enabled ones, are regulated under the Medical Devices Rules, 2017. So India is deploying AI internally and defining the rules for AI products in the market at roughly the same moment.

    How It Works:

    • Reuters reports the pilots include AI tools that generate pre-inspection checklists, recommend regulatory actions, draft inspection reports, classify manufacturing facilities by risk level, and produce reports on serious adverse events.

    • CDSCO says more than 99% of its regulatory processes are now digitized, which is the part that makes the rest of this possible.

    • Phase one of an end-to-end digital drug regulatory platform is planned within 18 months, spanning research, clinical development, manufacturing, approvals, distribution, and post-market surveillance.

    • GMP and CoPP certificates will carry QR-code digital authentication so overseas authorities can verify them directly.

    • The device guidance classifies software into risk Classes A through D by intended use and takes a lifecycle approach: software design, verification and validation, cybersecurity, updates and maintenance, post-market surveillance, and clinical evidence.

    • For context on the enforcement climate this lands in, Reuters reported that India has taken action against roughly 90% of the high-risk facilities it inspected, 860 out of more than 960 risk-based inspections since late 2022.

    Why It Matters: Most of our conversations about AI in regulated work assume the AI is on our side of the table, drafting our documents while the regulator reads them the old-fashioned way. India is describing the opposite arrangement, where the checklist that shapes your inspection, the risk tier that determines how often you get one, and the first draft of the report about you are all AI-assisted. That changes what your data profile is actually doing. Your filing history, your inspection record, and your adverse event reports stop being an archive and start being a live input into a classification that has consequences for you. The 99% digitization came before the AI, which is the same lesson as the Amgen dossier a couple weeks back. The AI story is always sitting on top of a long, unglamorous data cleanup nobody wrote a press release about.

    My Take: The near-term thing I'd watch is what happens the first time a firm wants to contest an AI-informed risk classification. If a facility gets tiered high-risk and the tiering came out of a model, what does the appeal look like, and what does CDSCO have to show you about how it got there? Nobody has a good answer to that yet, and India will probably have to work one out in public before the rest of us do. The symmetry with MHRA above is hard to miss. One agency is warning us about the AI-generated material we send in, and another is building AI to process what arrives. Both are pointing at the same requirement, which is that somebody accountable has to be able to explain how the output came to exist. That expectation is going to keep showing up on both sides of the exchange, and I'd rather we practice it now while it's still a voluntary disclosure checkbox.

    Source: Reuters / SRN (July 30, 2026) | The Hindu BusinessLine (July 30, 2026) | The Hindu — CDSCO guidance on AI/software devices (July 29, 2026)

    Get the next edition in your inbox

    Clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.

    No spam. Unsubscribe anytime.

    Newsletter

    Get the AI advantage for life science professionals—delivered weekly

    Cut through the AI hype in life sciences — clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.

    No spam. Unsubscribe anytime.