Agent sprawl in life sciences - part 2: how to bring it back into a coherent picture

Agent sprawl in life sciences - part 2: how to bring it back into a coherent picture

Writer

Susu Zhang

Date

July 23, 2026

Contact us

Why full standardisation isn't the answer, and what a pragmatic path forward looks like

In part one, we looked at how agent sprawl quietly takes shape across pharma R&D: computational biology, clinical data science, regulatory, safety, each building reasonable solutions to real problems, none of them with visibility into the others. We looked at why that's particularly complex in a GxP context, where the line between a convenient internal tool and a regulated system can get blurry without anyone noticing, and at the five data modalities that make a one-size-fits-all platform approach unrealistic.

The diagnosis is the easier half. The harder question is what to do about it. Because the instinct, once you've mapped the sprawl, is to want to clean it all up: one platform, one model provider, one governance framework. In pharma R&D, that instinct is almost always wrong.

What good consolidation actually looks like

However, the more useful framing is a thin shared platform with strong shared standards. Not standardising what teams build, but standardising how it's instrumented, governed, and visible.

Concretely, that means three things:

• A shared instrumentation layer. Every agent emits consistent logs, usage data, and evaluation signals. This doesn't require migrating to a common platform. It requires agreeing on a standard. Once you have it, you can measure quality and performance, track cost and usage, and debug problems across the portfolio.

• A model access layer. Central visibility into which models are in use, which versions, and under what contracts. Not to restrict choice, but to know what's there. When a model provider deprecates a version or changes pricing, you want to know which systems are affected before they break.

• An observability layer. A single place where someone can see all agents, who owns them, what data they touch, and whether they're still being used. This is less sophisticated than it sounds. A shared registry and a lightweight dashboard get you most of the way there.

What stays decentralised: which models, and which chunking and embedding strategies each team chooses to use, where agents are deployed, and more importantly who owns the agent and the data behind it. Those decisions stay with the people who understand the context. The shared layer doesn't override that. It just makes it visible.

What this looks like in practice

Take the four agents from part one. The computational biology team's PubMed scanner, the clinical team's safety-signal agent, regulator’s dossier drafting tool, and the quality team's stability monitor don't need to merge into one system, and they don't need to share a model provider. What they need is to show up in the same registry, log to the same instrumentation standard, and route through the same model access layer, so somebody can finally answer basic questions: what exists, what does it touch, and what happens when one of the underlying models changes. None of that requires touching how any of the four teams actually built their agent. It just means the four of them stop being invisible to each other.

The confidentiality question from part one gets easier too. Once every agent is registered with what data it touches, someone can actually check whether the quality team's stability monitor and the regulatory team's dossier drafter are handling data at the classification level they're supposed to, instead of hoping nobody's quietly drifted from Tier 3 into Tier 2, or from internal into confidential.

The spend problem is the easiest one to point to. Three teams calling separate external model APIs, with no shared view of the combined bill, was one of the earliest symptoms in part one. A shared model access layer doesn't consolidate those contracts, teams can still pick different providers, but it does mean someone can finally add the numbers together and see what the sprawl is actually costing, instead of finding out at invoice season.

The instinct, once you've mapped the sprawl, is to want to clean it all up: one platform, one model provider, one governance framework. In pharma R&D, that instinct is almost always wrong.

The organisational piece nobody talks about enough

Technology aside, agent sprawl exists because building independently has been faster and easier than building together. That's a rational response to how most pharma R&D organisations are structured. Teams are accountable for their own timelines, their own budgets, and their own deliverables. Coordinating across functions adds friction.

Changing that requires making shared solutions easier to adopt than building alone. Which means:

• Shared platform capabilities that are faster to plug into than to rebuild from scratch

• Clear agreements about what each team is responsible for when they build on shared infrastructure

• Quality and compliance stakeholders involved early, not called in for sign-off at the end

That last point is the one that tends to get skipped. Compliance is treated as a gate rather than a design input. When it comes in late, it causes delays and creates resentment issues on both sides. When it comes in early, it shapes decisions that are much cheaper to make up front: what tier does this sit in, what change control applies, what does validation look like for this type of system.

The teams that have navigated this well tend to have someone sitting at the intersection of data science and quality, someone who acts as a translator rather than a gatekeeper. That role is undervalued and underrepresented in most organisations.

What to aim for in the near term

Full consolidation isn't a realistic near-term goal, and it doesn't need to be. The more useful target is visibility and control: knowing what exists, who owns it, and what it touches.

A reasonable near-term state:

• Every agent is registered somewhere, with an owner and a basic description of what it does and what data it accesses.

• New agents are built to a common instrumentation standard and share one evaluation setup and orchestration framework, even if the underlying stack varies: AWS or Azure, whichever vector store, whichever model provider.

• There's a shared layer for model access, ideally through common MCP servers for shared data sources, so version changes and contract renewals aren't surprises

• Cost and usage across every agent and contract lands in one place, so total spend is something you can just look up, not something you find out about when the invoices come in.

• At least one high-value use case, ideally something like stability monitoring or deviation triage where the business case is straightforward, is in production and generating real evidence.

That last point matters more than it sounds. Abstract arguments for governance don't move organisations. A concrete example that saved two weeks of investigation time, or caught an OOS trend three weeks early, does.

Get one of those right, and the conversation about shared infrastructure becomes much easier.

A closing thought

Agent sprawl isn't the result of bad decisions. It's what you get when good people solve real problems independently, under pressure, without anyone stepping back to look at the whole. In most industries, the main cost of that is some duplication and inefficiency. In a regulated environment, the cost can be higher: a system operating in a validated context that nobody realised was there, a prompt change that triggered a revalidation obligation nobody was expecting, a compliance gap that surfaces during an inspection.

None of that means don't build agents. It means build them the way you'd build anything else that touches regulated data: knowing what you're building, who owns it, and what happens when it needs to change. The organisations that get this right probably won't have the flashiest agents. They'll be the ones who took stock of what they already had before adding to the pile.

Talk to our experts

Let's create real impact together with data and AI

Andreas Kjær

Life Science Lead

Andreas Kjær