Articles

Harnessing Explainable AI to Navigate Process Intensification Data with Single-Step Mixed-Mode Chromatography

Dr. Mark Schofield and Christina Caporale discuss how explainable machine learning models, data collation, and targeted transport frameworks uncover hidden process trends and advance single-step polishing strategies.
Written byShiama Thiageswaran
InterviewingMark Schofield and Christina Caporale
A bioprocess researcher wearing safety glasses analyzes experimental data from mixed-mode chromatography on a laboratory computer.

iStock

Register for free to listen to this article
Listen with Speechify
0:00
5:00

While the push for process intensification has left few areas of biomanufacturing untouched, downstream polishing remains difficult to streamline. Consolidating multiple purification steps into a single unit operation is highly desirable for footprint and efficiency, but multi-variable chromatography systems have historically been challenging to optimize. The interconnected, overlapping variables inherent in mixed-mode sorbents frequently produce noisy datasets that resist standard experimental analysis.

With pipelines remaining highly unpredictable and operations requiring greater efficiency, finding smarter ways to evaluate complex process data is becoming essential. This challenge was recently explored at Cytiva.

Dr. Mark Schofield, Director of Science at Cytiva, emphasizes that the motivation for their recent work was to establish better platform guidance. “We came into this with an interest in process intensification," Dr. Schofield explains. "We see process intensification throughout the whole mAb manufacturing process, all the different steps. But for polishing, we've seen people use single-step polishing, and we wanted to give our customers better recommendations on how to do that. We wanted to understand it more ourselves.”

Working in analytical science?

Register for a FREE Separation Science account to subscribe to the Separation Science Newsletter.

Subscribe for free

That desire for a deeper, more predictive understanding led the team to explore how to condense traditional multi-step polishing down to a single, highly efficient step. “Here we're using CaptoTM adhere, which is our mixed-mode chromatography sorbent,” Dr. Schofield outlines. “We really wanted to understand how to use that as a single step to get what we'd normally do in two steps using anion exchange and cation exchange or anion exchange and hydrophobic interaction chromatography (HIC).”

The Analytical Bottleneck: The Need for Domain Knowledge

Unlocking the true potential of mixed-mode chromatography requires more than just generating vast quantities of data; it requires a highly specialized approach to data modeling.“There are so many challenges that scientists have to face throughout the entire development of the pipeline and modeling process,” highlights Christina Caporale, Bioprocess Engineer at Cytiva. “Even from the initial experimental challenges with data quality and experimental execution, you really need a lot of domain knowledge about the data generation itself, the analytics, and what's a reasonable data set.”

Without this close alignment between data science and bench reality, critical nuances are easily missed or misinterpreted. She cautions that a distinct skills gap emerges when data scientists work in a vacuum. “It's hard when you have data scientists who are building models about things they don't have intimate knowledge of," Caporale adds. "You need a balance of expertise—people who understand the experimental design, have hands-on bench experience, and possess the analytical skills to maintain a close and continuous dialogue.”

Choosing EBM Over "Black Box" AI

When evaluating modeling frameworks for interpreting their mixed-mode data sets, the team deliberately avoided both highly complex, opaque neural networks and overly simplistic linear regressions. Caporale notes that the team chose an explainable boosting machine (EBM) framework because it perfectly fit their specific scientific purpose.

“Regression-based models are what we use for the most part," Caporale shares. "But when you're combining datasets from different experiments, operating ranges, and starting conditions, the relationships are often more complex than a simple linear model can capture." On the other end of the spectrum, neural networks present their own hurdles. “While model performance is important, we prioritize understanding why the model makes certain decisions or how it arrives at its predictions—a need met by the transparency and interpretability of the EBM model.”

The spectrum of process modeling approaches, highlighting where interpretable, data-driven explainable boosting models (EBMs) fit relative to mechanistic white-box and black-box machine learning methods.

Cytiva

Because the EBM model is open-source and easily accessible in Python, it bridges the gap for bioprocess teams without extensive machine learning backgrounds. More importantly, it directly enhances mechanistic understanding. "These predictions help identify patterns and driving factors that can be translated into mechanistic hypotheses, which is what made the model useful for our case," Caporale states.

This transparency is also a vital component of regulatory compliance. Dr. Schofield highlights the FDA's recent regulatory focus on integrating artificial intelligence into decision-making workflows. He notes that the FDA's 19-page draft guidance on AI decision-making emphasizes 'credibility'—a term that appears 75 times throughout the text.

For biomanufacturers, this makes explainable modeling an operational necessity. “That intelligibility and credibility from this kind of model is the key,” Dr. Schofield emphasizes. “It gives you credible outputs because you can understand how the model works, where there is data, and where there is no data. That's really key to the credibility of the model.”

Uncovering Surprising Trends via Data Aggregation

The true power of the EBM approach became evident when the team aggregated data across multiple distinct experiments rather than analyzing them in isolation. This compilation allowed the team to cross-examine variables that traditional literature searches struggle to reconcile.

Caporale recalls her past frustrations when reviewing historical literature in which experiments were conducted under disparate conditions. “Experiments typically prioritize pH and conductivity, but comparing results across studies is difficult when other variables, such as residence time or buffer composition, differ. “Previously, there hasn't been a reliable, consistent way to normalize those differences,” Caporale explains. “This approach to modeling allows us to aggregate all available data and factors, providing a clear indication of which variables are truly significant."

Aggregating disparate experimental datasets into a unified EBM model to rank feature importance and map specific feature effects across key bioprocess variables.

Cytiva

The Molarity Surprise and the Shift in pH

A key finding from their feature importance analysis was that buffer molarity emerged as a more critical driver for aggregate removal than bulk conductivity—a trend the team hadn't explicitly isolated in past DoEs. While Caporale notes that scientists should remain skeptical and always run experiments to validate machine learning hypotheses, the linear trend was remarkably clear.

“We hadn't prioritized buffer molarity in our previous experiments," Dr. Schofield admits. "This analysis highlights it as a critical factor, signaling that we must design future experiments with greater precision. Having this insight enables us to refine our methodology and improve our research outcomes.”

Furthermore, looking at the aggregated data across different mAbs reshaped their understanding of pH optimization. "Our model, incorporating data from three distinct mAbs, identifies a clear global optimum around pH 8," reveals Dr. Schofield. "Performance declines at both higher and lower pH levels. This was a surprising finding, as we previously focused on lower pH ranges to better align with adjacent process steps. However, the aggregated data demonstrates that if aggregate removal is the priority, targeting this higher pH is essential."

A Pivot to Mechanistic Modeling

The EBM model didn't just spot isolated trends in variables; it uncovered systemic behaviors that necessitated a deeper thermodynamic evaluation. Specifically, the data showed a clear divergence in impurity clearance behaviors. Caporale observes that aggregate clearance depended solely on residence time, whereas retrovirus-like particle (RVLP) clearance depended on both bed height and residence time.

Because this feature proved highly robust and repeatable across disjointed data sets, it provided the team with a definitive anchor point.

“That surprising observation led us down that path of exploring different modeling techniques,” Dr. Schofield continues. “But from the EBM, we pulled out that this was a consistent, generalizable feature across multiple datasets. Knowing it was reproducible gave us something concrete to chase down, which ultimately led us down the mechanistic modeling path.”

Demystifying Machine Learning for the Bench

For bioprocess teams looking to maximize their legacy data, Caporale emphasizes that adopting these frameworks is far less intimidating than it seems.

“Machine learning might seem intimidating, but it is more accessible than you’d expect. Using Python and current tools, it’s quite manageable to build a model on a small dataset—this one took me only a day or two to get running. The process is straightforward, especially with AI tools to assist with coding.”

For organizations running platform purification processes across dozens of molecules, the operational payoff of this approach is substantial. Aggregating historical data allows teams to quickly establish optimization boundaries without having to start from scratch for each new molecule.

“This approach offers quick guidance on where to focus your efforts, helping you reach platform conditions with minimal work,” Dr. Schofield concludes. “While you'll always need to verify boundaries, aggregated data helps you identify those conditions much faster.”

Add Separation Science as a preferred source on Google

Add Separation Science as a preferred Google source to see more of our trusted coverage

Meet the Author(s):

Interviewing

  • Mark Schofield

    Mark Schofield leads a team of scientists and engineers to fundamentally understand how separations work and apply that knowledge to challenging purifications. He is happiest guiding technology and innovation through an understanding of applications. But more recently, has been driving thought leadership and building bridges to better serve science and the biopharmaceutical industry!

    View Full Profile
  • Christina Caporale

    Christina Caporale is an Engineer in the R&D Bioprocessing group at Cytiva, where she develops downstream process intensification strategies and data‑driven approaches to improve biotherapeutic manufacturing. She holds a bachelor’s degree in chemical engineering and is currently pursuing her M.S. in Data Science at the Università degli Studi di Padova.

    View Full Profile

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...