Key Takeaways
Predictive formulation models can guide the selection of excipients, ingredient proportions, and processing conditions for defined product targets.
High-throughput screening supplies experimental data that machine-learning models can use to predict amorphous solid dispersion formation.
Bayesian optimization uses new experimental results to select subsequent formulation experiments, with laboratory testing used to assess the resulting candidates.
Model evaluation should address intended use, assumptions, validation, and uncertainty, with oversight proportionate to the model’s contribution to product quality.
Where Prediction Enters Formulation
Turning a drug candidate into a product requires decisions about ingredients, their proportions, and the conditions under which they will be processed. Those choices must deliver the intended product properties while supporting a workable manufacturing process. Formulation development builds that understanding through experiments that examine how materials and processing conditions affect the resulting dosage form. Each round of testing raises another question: which combination should be investigated next, and what information is needed to justify moving it forward?
Predictive formulation brings computational methods into those decisions. Mechanistic models describe relationships using physical principles, while data-driven models identify patterns in experimental observations. Hybrid approaches combine the two. High-throughput experiments can supply the data needed to evaluate a wider range of formulations, and iterative optimization can use new measurements to guide subsequent selections. These approaches offer different ways to focus laboratory work on combinations most likely to meet defined targets.
The opportunity these models present is to narrow the search while retaining the experiments needed to establish product performance. The value of a given model depends on what it predicts, how researchers test that prediction, and whether it remains useful when materials or processing conditions change. Assessing predictive formulation therefore begins with a specific development decision and the evidence that a model can improve it.
Learning from Small-Scale Experiments
Limited quantities of active pharmaceutical ingredient (API) make the choice of experiments consequential early in development. To study amorphous solid dispersions (ASDs), a team at the University of Macau reported using small-scale, high-throughput screening to make 1,272 binary and ternary formulations.1 It then characterized whether each had formed an amorphous material: 188 had, while 1,084 remained crystalline. Those results supplied the training data for models predicting dispersion formation.
The best-performing model achieved 96.7% accuracy and approximately 87.9% precision on the in-house data set. The team also assessed its predictions against larger-scale formulations reported in the literature. These findings illustrate that dispersion formation is a candidate for predictive screening, provided the result stays attached to its measured endpoint. Whether a selected dispersion meets other development requirements calls for other tests.
For oral lipid-based nanoparticles containing cannabidiol as a model drug, another team, from the University of Toronto, prepared 10% of the formulations in a defined design space.2 Measurements from that subset trained machine-learning models that predicted properties across 1,215 formulations. The team then made selected candidates to check their predicted performance. The initial experiments informed which combinations to investigate further, while the subsequent experiments tested the selections.
Making the Next Experiment Count
Once an initial data set exists, each new result can guide another choice. In an oral microemulsion study, a Virginia Commonwealth University team spanning pharmaceutics and chemical engineering began with 22 experiments and used batch Bayesian optimization to select five successive groups of five experiments.3 The search considered ingredients and processing conditions while seeking formulations that met specified physicochemical criteria.
The researchers identified five high-performing blank microemulsions, four of which remained physically stable during storage for up to 30 days. After selected formulations were loaded with model drugs, three showed the combination of drug loading, stability, and in vitro permeability reported in the study. The workflow used laboratory results to direct the next round of work and then tested the candidates it identified. The authors describe later-stage translation as a possibility for further investigation.
From Digital Formulation to Tablets
Selecting a candidate still leaves the task of making a dosage form with the desired properties. In a tableting study reported by a University of Strathclyde-led collaboration with pharmaceutical and equipment companies, a digital formulator used material information to choose excipients, their proportions, and an initial compaction pressure against defined manufacturing criteria.4 An automated system made and tested the tablets, and Bayesian optimization refined the pressure. The integrated workflow was demonstrated in nine use cases involving six APIs.
The authors reported using 65% less API material than an approach described in the literature. This was a comparison with a published benchmark, rather than a measured saving across parallel drug development programs under identical conditions. Along with the six-hour result, it shows what the connected workflow achieved for its defined manufacturing task.
The platform’s limits matter to its use. Its models were trained on a finite collection of materials, and the current equipment accommodates a restricted range of tablet geometries. It does not explicitly assess some scale-up risks, including over-lubrication, segregation, and sticking. The study established that the platform could meet its specified tablet manufacturing criteria. Applying the approach to other materials or production conditions requires further evaluation.
Does Machine Learning Improve the Prediction?
A new model should improve on the method available for the same decision. In a study of API solubility in binary solvent mixtures from a team affiliated with Nicolaus Copernicus University in Poland, a physics-based model was compared with an approach that added a machine-learning correction.5 The correction offered no meaningful improvement for a relatively homogeneous group of compounds but improved the reported prediction metrics for a more diverse set of APIs and related compounds. The researchers evaluated how the methods performed on compounds withheld from model training.
The solubility comparison illustrates why a model should be evaluated under conditions relevant to its intended use. In protein and peptide formulation, a team from LEUKOCARE, a provider of biopharmaceutical formulation and development services, developed Excipient Prediction Software (ExPreSo), training it on information from 335 approved peptide and protein drug products to suggest commonly used excipients.6 The tool could help prioritize ingredients for screening, with experimental testing needed to establish the properties of the resulting formulation.
Measuring the Limits
Prediction becomes harder when a formulation behaves differently from the cases a model handles well. In an academic and clinical research collaboration on polymeric long-acting injectables, a method combining machine learning with release theory was tested using a commercial product and newly prepared microspheres.7 The authors found prediction more challenging for slow-release formulations than for the intermediate-release formulations they examined. Adding three early measurements reduced the error in the slow-release cases.
Those measurements exposed a weakness and supplied information to address it. U.S. Food and Drug Administration (FDA) guidance identifies considerations for evaluating a model, including its intended use, assumptions, supporting experiments, validation, uncertainty, and contribution to assuring product quality.8 The guidance helps frame the evidence needed for a proposed use of a predictive formulation model.
How Much Empirical Work Can Be Reduced?
The evidence is strongest for reducing untargeted screening within defined tasks. Models can help select ingredient combinations, choose subsequent experiments, and guide work toward measured formulation or manufacturing targets. The published results do not establish a single saving that applies across products or stages of development. Their practical value is in making the next experimental decision better informed, then checking that decision against the properties the product must achieve.
References
1. Lu, Tianshu, et al. “Combining High-Throughput Screening and Machine Learning to Predict the Formation of Both Binary and Ternary Amorphous Solid Dispersion Formulations for Early Drug Discovery and Development.” Pharmaceutical Research. 42: 697–709 (2025).
2. Bao, Zeqing, et al. “Data-driven development of an oral lipid-based nanoparticle formulation of a hydrophobic drug.” Drug Delivery and Translational Research. 14: 1872–1887 (2024).
3. Gunawardena, Madeline, et al. “Machine learning guided optimization of an oral microemulsion system: a Bayesian optimization approach.” AAPS Open. 12: 34 (2 Jul. 2026).
4. Abbas, Faisal, et al. “Accelerated drug development using a digital formulator and a self-driving tableting data factory.” Nature Communications. 17, article 4739 (1 Apr. 2026). Nature Communications https://doi.org/10.1038/s41467-026-71204-6
5. Przybyłek, Maciej, et al. “When Does Machine Learning Add Value over Theory? Predicting API Solubility in Binary Mixtures with COSMO-RS and DOOIT2 Across Diverse and Homogeneous Systems.” Molecules. 31: 1566 (2026).
6. Vidal-Henriquez, Estefania, et al. “Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo).” Computational and Structural Biotechnology Journal. 27: 4517–4525 (2025).
7. Wang, Tianqi, et al. “Predicting drug release from polymeric long-acting injectables using a machine learning approach to decode formulation-performance relationships.” Drug Delivery and Translational Research. 13 Jul. 2026.
8. “Q8, Q9, & Q10 Questions and Answers—Appendix: Q&As from Training Sessions (Q8, Q9, & Q10 Points to Consider).” U.S. Food and Drug Administration. Aug. 2012.












