All editions
    Share this edition:
    Why I don't like Copilot: the full rant

    Why I don't like Copilot: the full rant

    You asked why I don't like Copilot. Buckle up.

    Greetings from the BIO International Convention 2026!

    It's the last day of the convention and I'm not going to lie, I'm beat. That's not a complaint, it's been a fabulous event. I got to spend quality time at the booth with my CMC consulting colleagues at BPTG, gave a talk about getting ready for agents in CMC and Bioprocessing, and learned about some fantastic organizations.

    This week's topic comes out of a few different conversations I had about choosing AI tools. There's a lot to be said here, so I'm going to do my best to stay focused on one piece of it. And I want you to know up front that you can take or leave my advice, it's entirely up to you. This one reflects my opinions and nobody else's, but I have real good reasons for them.

    So here I go: If you or your team are new to AI and trying it out for work tasks, do not use Microsoft Copilot. *gasp*

    "But we just bought the seats!" you say. "We finally got an AI tool approved, and it's Copilot!" I hear it almost daily. My sympathies.

    I'm not going to stop there, because if I'm going to say something like that to all of you, you deserve a clear justification for why, and why I'm so adamant about repeating it to everyone who asks.

    Microsoft applications have shaped the way billions of people work. We have Word documents and Excel spreadsheets and PowerPoint presentations, and with very few exceptions, anyone who has touched a computer for work has used these products, for better or worse. So it makes sense that on sheer distribution advantage alone, Microsoft has been able to roll out their AI wrapper to all those same businesses already locked into Office. If you're already paying for the applications you can't do your job without, why wouldn't you just tack on the AI solution from the same vendor?

    There are reasons you shouldn't, and I have three of them for you.

    One — your prompt gets rewritten before the model ever sees it.

    Most people don't realize that when you type a prompt into Copilot, it doesn't go to the model. Not your version of it, anyway. There's an orchestration layer that intercepts what you wrote first. It runs your prompt through safety and "Responsible AI" checks, and if it doesn't like what it sees, it can kill the request before the model is ever involved. If it passes, the orchestrator goes and pulls data from across your organization (emails, files, chats, whatever it thinks is relevant) and bundles all of that together with Microsoft's own instructions into something called a "meta-prompt." That package is what gets sent to the model, and then the model's answer gets run back through another filter on the way to your screen.

    So when you select a shiny frontier model in the dropdown, guess what…you don’t even get to talk to that model. You're talking to it through a Microsoft-built wrapper that has rewritten your prompt, stapled its own instructions on top of yours, and injected a bunch of context you didn't ask for. As one detailed write-up puts it, Copilot runs on the same underlying GPT model as ChatGPT, but because it's embedded in Microsoft 365 it has its own system prompt (instructions defined by Microsoft) and its own content filter. Your carefully worded instructions are now just one voice in a crowded room, and they aren't the loudest one.

    Two — the model you picked might not be the model you think.

    This one matters less than the other two, but I’m always tell my readers to “know your tools” so this is part of it. Even when you choose a model in the interface, you may be getting an older version of it. Copilot is locked to a dedicated, vetted version of the model, while the consumer tools like ChatGPT roll out new versions rapidly. There's a lag between when a new model ships and when Microsoft can actually deploy it inside Microsoft 365, because they have to make sure all their function-calling plumbing still works with the newer version first.

    I'll be fair here: the last couple of model generations are all genuinely good, so a slightly older frontier model isn't the end of the world. If this were the only problem, I wouldn't be writing this newsletter. But it stacks on top of everything else, and it doesn't matter how good the model is if your original prompt never reaches it in one piece.

    Three — there's a hidden budget, and your instructions are the first thing cut

    This is the one that actually explains the daily frustration, and it's the one that makes me a little crazy because of how invisible it is. The model has a fixed amount of room to work with, a token budget, and that budget has to be split between Microsoft's system instructions, your conversation history, all that organizational data the orchestrator injected, your actual prompt, and the space reserved for the answer. Think of a fixed pie that someone else has taken the liberty of slicing for you.

    When the pie gets crowded (a long conversation, a verbose prompt, a lot of retrieved documents) something has to give, and the things that give are your chat history and the depth of the answer. Microsoft's own guidance recommends keeping input to around 3,000 words or less for focused editing tasks, and users have complained on Microsoft's own forums that they can't even fit a single example into a prompt the way they can with other tools. So your instructions aren't being ignored because the AI is dumb, they're being ignored because they got squeezed out of the room by content you never saw and weren't told about.

    And that's the part that gets me, because all of this happens silently. Nobody tells you your prompt got rewritten, or that the model is a version behind, or that your instructions got truncated to make room for an email the orchestrator decided was relevant. You just get a worse answer than you asked for, and you're left to assume the problem is you.

    Which, conveniently, is about what Microsoft has tended to suggest. Back when this first surfaced in 2024, customers complained that Copilot wasn't as good as ChatGPT, and the company's response was essentially that users didn't understand the difference between the products, and per reporting on internal feedback, that a lot of people were just bad at writing prompts. "It's a copilot, not an autopilot," went the line.

    Now, that was two years ago, and in fairness the tool has had plenty of time to improve since. So let me give you something more current. I am, if I may say so, an above-average AI user. Earlier this year I sat there and tried every trick I know to get Copilot to follow a set of instructions, and it flat-out would not do it. And while I was writing this very issue, a friend was texting me in real time, fighting to get it to export a file to PDF. Now, I'll be honest, I don't actually know what tripped her up, it could've been a permissions thing on her tenant rather than the tool itself. But export to PDF is about as basic as a request gets, and I have a hard time picturing ChatGPT or Claude leaving someone that stuck on it.

    So I'm sure "you're holding it wrong" is true for some people some of the time. But when you've built a system that silently mangles a competent user's input and then suggests the user is the problem, you've got something no amount of prompt training is going to fix.

    So here's where I land. I’m not just going to say that Copilot is bad software because it's a perfectly reasonable enterprise tool built to solve enterprise problems like permissions, governance, and keeping your data inside the tenant. Those are real and worth solving. The danger is what it does to a person's first impression of AI.

    Because here's how it actually plays out… a team finally gets an AI tool approved after months of waiting, and it's Copilot, because of course it is. They try it, it doesn't follow their instructions, it gives them mush, and they walk away thinking "that's it? that's what everyone's losing their minds over?" They don't conclude they got handed a heavily-wrappered tool with a constrained context window and a silent rewrite layer.

    They conclude AI is overhyped.

    And now you've got a person who isn't behind because they lack access, they have access, they're behind because the access to something they think doesn’t work.

    Now, there are people who will be fine. These are the ones who go get a personal subscription to a frontier model and find out what these tools can really do. Some of your colleagues will do exactly that, and good on them. But you can't build an organization's AI adoption on the assumption that everyone will privately go around the tool you gave them (that’s actually the opposite of what you want). They'll take the one bad impression and file the whole thing under "not for me."

    So if you're a team lead, or you're the person who got handed the Copilot seats and told to be grateful: by all means use it for what it's good at (and let me know what that is!). But before you let anyone decide what they think about AI, get them in front of a tool that actually sends their words to the model. Let them form their opinion based on what these things can really do, not on what survives the trip through the wrapper. That's not a knock on Microsoft, it's just being honest about what you can and can't conclude from the tool in front of you.

    Thanks for reading!

    Alexa


    News

    BMS Bristol Myers Squibb Logo 500x281

    BMS Bets the Whole House on One AI Platform

    What's New: Bristol Myers Squibb signed a strategic agreement with Anthropic to roll out Claude across its entire global operation — research, clinical development, manufacturing, commercial, and corporate functions. The deployment puts the tool in front of more than 30,000 employees. What makes this one worth your attention isn't the size, though, it's where they're pointing it. The agreement specifically calls out quality systems and production operations: root-cause investigations, CAPA documentation, and data-driven batch release decisions. If you've read this newsletter for more than a few weeks, you know that's exactly the territory I get twitchy about.

    How It Works:

    • The rollout prioritizes three areas: deploying Claude Code within engineering and data science teams, embedding agents into workflows, and connecting the model to BMS's "institutional knowledge" spread across its systems and data repositories.

    • It's explicitly an agentic deployment, meaning the pitch is autonomous multi-step tasks, not just a chat window. BMS frames it as moving beyond the conversational tools that defined the first wave of enterprise AI.

    • On the R&D side, research teams get access to help identify and optimize drug targets across oncology, neuroscience, hematology, and immunology, and the tool will be used in drug development to automate trial documentation and regulatory submissions.

    • The stated motivation from BMS leadership is unlocking value "trapped behind decades of data silos."

    Why It Matters: A major pharma is putting AI agents into deviation handling, CAPA writing, and batch release — the exact GMP-critical workflows where a confident, wrong output does real damage. It can go one of two ways, and which way depends entirely on the standard work wrapped around it. If BMS has the verification, traceability, and human-in-the-loop controls sized to the risk, this is the kind of deployment that makes the rest of us look slow. If they don't, they've just scaled their existing process — good or bad — across 30,000 people.

    My Take: The use cases here are the right ones — root-cause work, CAPA, batch-release support are all places where a well-governed AI can take real weight off your people, and I'm glad to see it pointed at meaningful problems instead of slide decks. What gives me pause isn't the what, it's the how much, from one source. Wiring a single provider this deeply into research, manufacturing, and quality systems all at once is a concentration risk we don't talk about enough. We spend enormous energy on supplier qualification and second-sourcing for raw materials and components, and then we'll happily make a single AI vendor load-bearing across the entire operation without the same scrutiny. So I have to ask the boring continuity questions: what's the plan when the service goes down mid-shift? If batch-release decisions have come to lean on a tool that's suddenly unreachable, does everyone get the day off? Do we fall back to paper and hope someone remembers how it worked? Is there a degraded-mode procedure, or just a faith that the uptime number holds? None of this is a reason not to do it. It's a reason to treat your AI provider like any other critical supplier — with a qualified backup, a documented manual fallback, and a tested plan for the day the thing you depend on isn't there. "It was working great right up until it wasn't" is not a continuity strategy.

    Source: Bristol Myers Squibb press release | Pharm Tech analysis

    Buying Ai scientists

    Everybody's Buying an "AI Scientist" This Month

    What's New: The BMS deal didn't happen in a vacuum. June has been a steady drumbeat of pharma companies signing multi-year agreements to embed AI agents into drug R&D, and the pattern is more interesting than any single deal. The clearest example: Sanofi expanded its long-running partnership with Owkin into a five-year collaboration to co-develop biopharma agents, backed by a license for Owkin's "AI Scientist" platform, K Pro. Around the same stretch, Alnylam signed an AI collaboration with Inceptive Nucleics worth a potential $2 billion, and Merck struck a $1 billion Google Cloud deal to build out its "AI-enabled enterprise."

    How It Works:

    • The Sanofi-Owkin agents are pitched as intelligent assistants that can autonomously perform complex tasks across drug R&D, deployed inside Sanofi's existing workflows rather than as a standalone tool.

    • The relationship isn't new, which is the point. Sanofi and Owkin have worked together since 2021 under a €90 million partnership focused on oncology target identification and patient subgrouping, later expanded into immunology. This is an existing collaboration leveling up to agents, not a cold start.

    • Owkin is signing the same kind of deal elsewhere. It recently announced a similar multi-year licensing agreement with AstraZeneca to deploy AI agents for competitive intelligence and R&D workflows.

    • The common thread across all of these: companies are moving away from isolated analytics tools toward embedding AI systems into broader enterprise workflows. The word doing the heavy lifting in every press release is "embedded."

    Why It Matters: Step back from the dollar figures and you can see the industry placing a collective bet on a specific idea — that the value isn't in a chatbot you visit, it's in agents wired directly into the systems where the work already happens. For most of us that aren't Sanofi or BMS, this matters as a signal of where the floor is heading. The capabilities these companies are paying billions to build custom will, in some watered-down form, show up in the tools the rest of us can buy off the shelf in a couple of years. It also tells you something about where the expertise bar is moving. When the agent can autonomously run a target-identification task, the human value shifts to framing the right question and verifying the output, which is the same story I've been telling about domain authority all along.

    My Take: I'll admit the phrase "Biological Artificial Superintelligence" in Owkin's own description of itself made my hype allergy flare up a little. But underneath the branding, the structural move here is the sound one — purpose-built tools grounded in specific pharmaceutical data, deployed inside real workflows, rather than a general assistant asked to wing it. That's the opposite of the Copilot problem I opened this issue with, and it's not a coincidence. The companies that can afford to do AI right are building exactly the kind of grounded, scoped, embedded systems that actually work, while everyone else gets handed a generic wrapper and told to be grateful. The gap between those two experiences is going to define who feels like AI "works" over the next few years, and it has almost nothing to do with the underlying models, which are all pretty good now. It has everything to do with the engineering around them.

    Source: Owkin/Sanofi announcement | MobiHealthNews

    Get the next edition in your inbox

    Clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.

    No spam. Unsubscribe anytime.

    Newsletter

    Get the AI advantage for life science professionals—delivered weekly

    Cut through the AI hype in life sciences — clear, practical takes on what matters for your CMC, QA, and regulatory work, once a week.

    No spam. Unsubscribe anytime.