AI mobile phase optimisation in HPLC is, for many method development scientists, the most immediately practical application of machine learning in the analytical workflow. Mobile phase and gradient conditions interact in ways that resist one-variable-at-a-time exploration, making this the stage of development where systematic, multivariate methods have always had the clearest advantage over intuition alone. Combining design of experiments with machine learning prediction compresses the iteration cycle further still, and recent work has taken this toward fully autonomous closed-loop optimisation. This article maps the approaches, from the established DoE foundations to adaptive and Bayesian methods, and addresses the regulatory context that governs their use.
It sits within the AI-assisted chromatographic method development guide, which covers the full method development workflow. For the earlier scouting phase that feeds conditions into optimisation, the companion article on automated method scouting is a useful starting point.
Key Takeaways
|
Why Mobile Phase Optimisation Is Time-Consuming
The number of variables involved in gradient HPLC method development quickly makes the experimental space intractable by sequential, empirical exploration. A method developer adjusting organic modifier type, gradient start and end composition, gradient time, pH, buffer type, and column temperature faces a combinatorial problem: varying each factor while holding the others constant misses interaction effects and requires an impractical number of runs to cover the space systematically.
The interaction between pH and selectivity is the clearest example. A change in mobile phase pH affects the ionisation state and, therefore, the retention of ionisable analytes, but that effect depends on the organic modifier concentration and the column chemistry simultaneously. One-variable-at-a-time exploration cannot map this reliably. The result, in labs that do not use multivariate methods, is that optimisation often stops at 'good enough' rather than reaching a genuinely robust method, with fitness for purpose confirmed within a narrow window that creates problems when conditions drift.
Optimisation Variable | Typical Range of Influence | Interaction Effects |
Organic modifier (% B start/end) | Retention, selectivity, peak shape | Coupled with pH for ionisables |
Mobile phase pH | Selectivity of ionisable analytes | Strong interaction with column chemistry and modifier |
Gradient time/slope | Resolution, run time, throughput | Interacts with organic start composition |
Column temperature | Selectivity, viscosity, pressure | Modifier-dependent; affects ionisation equilibria |
Buffer type and concentration | Peak shape, ion suppression in MS | Dependent on pH and organic composition |
Mobile phase optimisation is a multivariate problem. Any approach that varies one factor at a time cannot reliably capture the interactions that determine whether a method is robust.
Design of Experiments Approaches for HPLC Optimisation
Design of experiments has been part of analytical method development for decades, and it remains the most widely used and most defensible approach to systematic gradient optimisation. ICH Q14, the 2024 analytical procedure development guidance, explicitly supports science- and risk-based method development that accommodates DoE within an analytical quality-by-design (AQbD) framework, alongside the established ICH Q2(R2) validation requirements. Understanding the DoE designs in common use, and their relative efficiency, is the starting point for understanding where ML adds value on top.
The DoE designs used most frequently in HPLC optimisation:
- Central composite designs (CCD). The classic response-surface design: a factorial core augmented with axial and centre points, which provides a full quadratic model and allows curvature to be captured. Requires a moderate number of experiments and produces a well-characterised design space.
- Box-Behnken designs. A response-surface design that avoids extreme multi-factor combinations (vertices of the cube), making it practical when those extreme conditions are analytically risky or physically impossible. Often slightly more efficient than CCD for three to four factors.
- Doehlert designs. Provide uniform space-filling across the experimental region with particularly good efficiency for two to three factors, valued for their flexibility in extending the design without replication.
- D-optimal designs. Computer-generated designs that maximise the information content of the experiment set for a given model and constraints. Particularly useful when standard designs cannot be applied because of experimental restrictions.
All of these produce a response surface model, the mathematical description of how the separation quality responds across the experimental space, from which an optimum can be identified. The model is fit to the experimental data, not derived from first principles, which is why the quality of the training experiments matters as much as the design choice.
How Machine Learning Augments DoE in Gradient Optimisation
Machine learning enters gradient optimisation primarily by improving the quality and resolution of the response surface model, and by making it possible to predict separation quality at conditions that were not directly run. A 2025 study published in ACS Analytical Chemistry demonstrated this in silico approach, using QSPR models combined with linear solvation energy relationship theory to predict retention factors across mobile phase compositions without any new experimental runs, enabling optimisation of both isocratic and gradient methods from molecular structure alone.
What ML adds over classical response-surface modelling:
- Denser prediction. A classical response-surface model predicts separation quality at a fixed resolution determined by the design. An ML model trained on the same data can predict more densely across the space, resolving the optimum more precisely.
- Non-linear capability. Classical polynomial response-surface models handle quadratic interactions well but can miss sharper non-linear behaviour. ML models, particularly gradient boosting and neural network approaches, capture this more flexibly.
- Incorporation of in-silico prediction. Where retention models are built from molecular descriptors rather than solely from experimental points, the ML layer can predict across conditions that were never run physically, dramatically reducing the experimental burden.
- Uncertainty quantification. Probabilistic ML models, including Gaussian processes, quantify how confident the prediction is at each point, which helps direct experimental effort toward regions where uncertainty is highest.
For SepSci's chemometrics-fluent audience, the parallel is direct: classical response-surface modelling is, in effect, a chemometric approach, and ML augmentation extends it with greater predictive flexibility and the ability to incorporate structural chemistry. The interpretability advantage of the classical approach, a response surface that can be plotted, examined, and explained, is partly retained by keeping the ML as a layer on top of an experimental design rather than replacing the design entirely.
Bayesian and Active Learning Approaches
The most sophisticated current approach to gradient optimisation is Bayesian optimisation, which treats the optimisation as a sequential experimental design problem: at each step, the algorithm selects the next experiment to run based on what it has learned so far, balancing exploration of uncertain regions against exploitation of promising ones. A 2024 paper in RSC Digital Discovery, a collaboration between the University of Leeds and Pfizer, demonstrated a fully autonomous closed-loop Bayesian optimisation system that developed impurity-screening HPLC methods in as few as 35 experiments, using multi-objective optimisation to balance resolution and run time simultaneously.
Active learning takes a related approach: rather than following a fixed experimental design, the model continuously updates its predictions and recommends the next most informative experiment. A 2024 study in ACS Analytical Chemistry applied assisted active learning to LC method development, updating model parameters iteratively and selecting the most informative conditions for each subsequent experiment.
The practical implications for a method development lab:
- Bayesian and active learning approaches reach a robust optimum with fewer experiments than fixed DoE designs, which is meaningful when the analyte supply is limited or instrument time is constrained
- They handle high-dimensional optimisation, many factors simultaneously, more gracefully than classical designs, which scale poorly past five or six factors
- The tradeoff is complexity: implementing a Bayesian or active-learning workflow currently requires more computational and integration effort than applying a standard DoE design in existing method-development software
- For most routine pharmaceutical method development, a well-designed central composite or Box-Behnken study remains the practical standard, with Bayesian approaches most compelling for complex multianalyte separations or automated development platforms
pH, Organic Modifier, and Temperature: What the Models Cover
The choice of which variables to include in the optimisation design is itself an analytical decision, and one where method knowledge matters more than the algorithm. Not all variables are equally influential for all analyte types, and including too many factors in a response-surface design inflates the number of required experiments without proportionate information gain.
Guidance on factor prioritisation by analyte type:
- Ionisable analytes. pH is the highest-priority factor: it controls the ionisation state and therefore the retention and selectivity of acids and bases. For a mixture containing ionisable species, pH and organic modifier must both be included, and their interaction modelled.
- Neutral analytes. Organic modifier concentration and gradient slope carry most of the predictive weight. pH is secondary unless the mobile phase additive affects peak shape. Temperature becomes more relevant when selectivity is otherwise similar across candidate conditions.
- High-resolution separations (UHPLC impurity profiling). Gradient time and slope interact strongly with resolution at the critical pair. For impurity methods, an objective function that captures both resolution and run time is more useful than one that optimises either alone.
- LC-MS methods. Buffer type and concentration interact with ionisation suppression and must be included if MS compatibility is an acceptance criterion. Purely chromatographic optimisation without accounting for the detection constraint produces methods that look good on UV and fail on mass spectrometry.
The objective function used to evaluate candidate conditions is at least as important as the model itself. Resolution of the critical pair, peak symmetry, run time, and robustness to small perturbations can all be incorporated, and multi-objective optimisation, which the Bayesian Leeds/Pfizer work used explicitly, allows tradeoffs between them to be made explicitly rather than by intuition.
Regulatory Considerations for AI-Assisted Optimisation
For laboratories operating in pharmaceutical development or quality control, the regulatory framework for method development has evolved to accommodate data-driven optimisation. ICH Q14, finalized in 2024, provides the harmonised guidance on analytical procedure development and explicitly supports science- and risk-based approaches that include DoE and multivariate modelling within an AQbD framework. Together with ICH Q2(R2) on analytical procedure validation, it describes the lifecycle approach that governs development, validation, and post-approval change management.
The practical regulatory implications for AI-assisted optimisation:
- Document the development rationale. The design choices, the factors included and excluded, the objective function, and the basis for selecting the final conditions must be documented in a way that demonstrates a science- and risk-based development approach.
- Define the method operable design region (MODR). ICH Q14 introduces the MODR, the region within which the method continues to meet its performance criteria. An ML or Bayesian optimisation that identifies an optimum can also characterise the boundaries of acceptable performance around it, which supports MODR definition.
- Interpretability and explainability. A response surface model whose contour plots can be examined by a reviewer is more straightforward to defend than a black-box neural network. The interpretability advantage of DoE-grounded response surface modelling is one reason it remains the dominant approach in regulated method development, even as more sophisticated ML tools become available.
- Validation requirements do not relax. An AI-assisted method is validated to the same criteria: specificity, accuracy, precision, linearity, and range, under ICH Q2(R2) as a conventionally developed one. AI accelerates the development; it does not replace the validation.
What This Means for Your LabIf mobile phase optimisation is currently your development bottleneck, design of experiments augmented by machine learning is the most mature and most accessible point to adopt AI in the method development workflow. Start with a well-chosen response-surface design, define an objective function that captures what the method actually needs to achieve, and use ML prediction to map the optimum at higher resolution than the experimental grid alone provides. For complex analyte sets or automated platforms, Bayesian and active-learning approaches offer further efficiency gains. Throughout, keep the documentation and interpretability that a regulated environment requires in mind: the goal is not just to find the optimum, but to understand the space around it well enough to characterise a robust method. For the full method development framework, the AI-assisted method development guide covers each stage, and the AI in analytical science overview maps the wider landscape. |
This article was produced under Separation Science’s AI Editorial Guidelines



