All Models are Wrong, but Some are Useful 

I gave a deliberately provocative talk at this year’s Digital Bioprocessing Symposium, and I want to bring the same argument here, because I think it’s one our industry needs to sit with: the model alone will not lead to productivity gains.

I know that’s an odd thing for the CTO of a so called “modeling company” to say. But I don’t think it’s controversial once you look at how other transformative technologies actually played out.

 

What personal computers taught us about productivity 

In the 1980’s, businesses started buying personal computers at scale. And then, for almost two decades, nothing happened to productivity statistics. Robert Solow, who later won the Nobel Prize in Economics, put it bluntly: you could see the computer age everywhere except in the productivity numbers. Labor productivity actually dipped slightly at first. 

It took companies close to twenty years to figure out why. The answer wasn’t a better computer. It was that most organizations had simply bolted PCs onto their existing ways of working, paving the cow paths instead of building new roads. The technology itself turned out to be roughly one-fifth of the actual cost of the transformation. The rest was new workflows, new roles, new skills, and a genuine redesign of how work got done. Only once that reorganization happened did the productivity gains show up. 

I think we’re watching the same pattern with large language models today. The frontier models are converging and perform similarly on most tasks. And yet most enterprise GenAI pilots show no measurable return. The users who are succeeding aren’t the ones with the best model. They’re the ones doing what I’d call harness engineering: building rich context around the model, giving it tools to act, creating loops for checking and correcting its work, and giving it memory. 

A simple example: if you ask an LLM to “write me an email,” you’re not really saving time; you’re drafting, rereading, re-prompting, reformulating. You might even be slower than just writing it yourself. But if that model is wired into your knowledge base, your past email threads, your calendar, and your actual figures, it can draft the email and book the meeting and prepare the agenda. That’s when the time savings show up. The model didn’t change. What surrounds it did. 

 

Bioprocessing is making the same mistake 

I’d argue we’re repeating this exact pattern in bioprocess modeling. Much of the discussion about model-based bioprocessing gets stuck on the models themselves. Benchmark comparisons. RMSE. We are chasing accuracy as if it were the whole game. 

I think, at times, we’re asking ourselves the wrong questions. 

We ask: is the model accurate enough? I’d rather we asked: can we make a robust decision from the derived insights? Otherwise, we end up optimizing an accuracy metric in a direction we didn’t actually need. 

We ask: how many experiments will this model save us? I’d rather we asked: what workflow would we design if we actually trusted the model? Right now, most of us place the model on top of an existing workflow just to compare results, as a safety net. The better question is what needs to surround the model so we can continuously compound what we know, rather than just checking its homework once. 

To be clear, I’m not arguing the model doesn’t matter. It’s the single most important component in this system. At DataHow, we build on three pillars: hybrid modeling as a mechanistic backbone, transfer learning so you can transfer historical insight when your own data is thin, and uncertainty quantification, which I think is chronically undervalued. Knowing what you don’t know is often more useful than the prediction itself. 

And we need a mindset shift alongside it: we don’t need a perfect model, we need a useful one. George Box’s line, All models are wrong, but some are useful, gets quoted so often it’s almost a cliché, but it’s true. A young model early in process development doesn’t need to describe everything precisely. It needs to point you in the right direction and get more certain as development progresses. 

 

From models, to insights, to actions 

Here’s where I think the real work lies. A process model on its own, honestly, answers very little. The next layer has to turn a model into insights: uncertainty, sensitivity, criticality, robustness for process understanding; Pareto-based trade-offs and headroom assessments for development progress. And the layer after that has to turn insights into actions: concrete recommendations on whether to iterate, what to adjust, when to move to the next stage. This is the hardest part to get right, because it requires real domain knowledge, not just statistics. It’s also where the actual value lives. 

 

 

Done consistently, this becomes a loop: historical data in, models trained, insights generated, a gate-review decision made, and (if needed) another experiment designed and run. Done well, it also feeds Quality by Design almost for free: if the design space, criticalities, and parameter ranges are traceable back to the model, you’re not building quality into your regulatory filing as an afterthought – it’s already there. 

 

The unglamorous layer that actually determines success 

This is the part that rarely makes it into a conference slide, but it’s where most projects actually succeed or fail. We need the model connected to historical and real-time data, ideally via a direct connection versus a manual export. We need to harmonize units, naming conventions, and metadata. Glucose expressed in g/L in one system and mmol/L in another is a genuinely painful, unglamorous problem, but it’s the one that determines whether transfer learning works at all. We need lifecycle management, so every model version is traceable and auditable back to screening. We need security, and we need democratized access and usage of model-based technologies across bioprocessing. A subject-matter expert without coding skills or a data science background should still be able to participate in the process, not be locked out of it. 

 

 

Agents will have a role here, but I’d push back on the idea that they should be wired directly on top of our data, replacing robust process models and rigorous, traceable methods. I think their real value is as an interface, automating the repetitive, deterministic steps of retrieving data, training, and evaluating, while a person stays in the loop for the decisions that matter. Agentic technologies will be commoditized within a year or two. Teaching them the deterministic steps and providing a robust portfolio of domain-relevant tools that create value is the actual work in front of us. 

 

Why the urgency 

The number of AI-discovered molecules entering clinical trials keeps climbing. In 2024, 64% of drug-launch delays were linked to CMC. It’s no longer just about how fast you can discover a molecule; it’s about how fast you can build a robust process around it. 

Bioprocessing has a very solid foundation in hybrid modeling, but we need to kick on from here. I don’t think the next real advance in the industry will come from focusing on better models, but from the workflows, the connectivity, and the change management we build around it. By taking a more holistic, integrated perspective to model-based digital bioprocessing, we will accelerate our journey to realizing the returns of these transformative technologies. 

See the full presentation at the Digital Bioprocessing Symposium 2026 here

BACK TO POSTS
Transforming Digital
Bioprocessing