Abstract
Oral bioavailability (%F) is a key determinant of drug efficacy and a critical decision-making parameter in early-stage drug development. Current prediction of human oral %F often relies on preclinical animal studies, which are costly, ethically disputed and do not necessarily correlate with human %F. Machine learning (ML), an emerging in silico technology, has been leveraged to predict human %F. However, previously published ML approaches have often relied on binary classification or achieved limited regression performance, with R2 values ranging from -0.36 to 0.50. In this study, we demonstrate that an ML pipeline can indeed be used to predict accurately human oral %F using preclinical data, potentially reducing the need for animal experiments. A total of 192 FDA-approved compounds were curated, which were discovered to span a diverse chemical space.
Model inputs included molecular descriptors, structural fingerprints, published animal %F and cytochrome P450-aware metabolic feature, with the latter designed to account for metabolism-driven interspecies variability. Five-fold cross-validation, repeated 100 times, yielded an R2 of 0.766 (RMSE = 16.13%), substantially outperforming all previously published models and individual animal species univariate correlation, including non-human primates (R2 = 0.691). Notably, high performance was maintained without non-human primate data, with rat and dog %F combined achieving R2 = 0.723 ± 0.016, reducing dependence on ethically and logistically demanding primate studies.
Permutation feature importance analysis identified lipophilicity, surface area and CYP features as informative predictors, providing mechanistic interpretability alongside predictive accuracy. These results demonstrate that a combination of feature engineering and a comprehensive ML pipeline, can achieve precision and reproducible prediction of human oral %F from routinely collected preclinical data. The impact of this study includes the potential to reduce late-stage attrition and accelerate the translation of drug candidates to the clinic.
Introduction
Oral drug delivery remains the most dominant and preferred route of administration, accounting for 71% of the 2.4 trillion medications prescribed globally in 2023 (White 2024). This preference is largely driven by the convenience, safety, cost-effectiveness and patient compliance associated with oral formulations. In addition, orally administered therapies are increasingly recognised as critical for both the treatment and prevention of diseases, including emerging infectious threats such as COVID-19 (Fan et al. 2022). Furthermore, the development of oral equivalents for high-value therapeutics, such as glucagon-like peptide-1 (GLP-1) receptor agonists, holds the potential to significantly transform modern healthcare by improving accessibility and adherence (Harrison 2025; Drucker 2025).
Despite these advantages, the successful development of oral medications remains a considerable scientific and technical challenge (Brown 2025). A major contributor to attrition in the development of oral drug candidates is poor bioavailability (%F), defined as the rate and extent to which an active pharmaceutical ingredient is absorbed and becomes available at the site of action (Stielow et al. 2023). Achieving adequate bioavailability is inherently complex, requiring careful optimisation within a drug’s therapeutic window while accounting for physiological barriers such as gastrointestinal pH, enzymatic degradation and first-pass metabolism. Furthermore, formulation development introduces additional layers of complexity that can significantly impact drug product performance (Obrezanova 2023). Given its substantial clinical and commercial importance, a concerted effort has been undertaken to address the challenges of %F.
Currently, predicting human oral %F requires animal models, which include both small animals and non-human primates (NHP) (Obrezanova 2023). However, there are several issues here, including poor correlation to human %F, high costs and ethical considerations (Musther et al. 2014). Recent regulatory changes are striving for animal-free research, meaning an immediate alternative will need to be found in order to prevent research downtime (Akabane et al. 2010; Olivares-Morales et al. 2014; Huh et al. 2011; Negoro et al. 2025).
An alternative to animal models is in silico models, with machine learning (ML) emerging as a powerful tool for modelling complex behaviours. ML, a subset of artificial intelligence (AI), has demonstrated considerable successes in the field, including identifying novel compound-target interactions, unearthing new biological behaviours and even generating novel drug candidates (Catacutan et al. 2024; Du et al. 2024; Okubena et al. 2026). The real gravitas of this tool has started to materialise, with the first AI-designed molecule entering human trials within 12 months of discovery (Khalid Shaikh and Affaan Shaikh 2026). Its capacity to process high-dimensional and heterogeneous datasets, including numerical, text and image-based data, enables the integration of diverse sources of information at an unprecedented scale (Abdalla et al. 2023; Wang et al. 2022, 2023; Elbadawi et al. 2021; Onoue et al. 2026). Moreover, its ability to extract meaningful patterns from large and complex datasets are contributing to its ability to expedite developments in pharmaceutical research.
ML has also been applied to modelling physiologically-based pharmacokinetic (PBPK) data, including predicting maximum concentration, clearance rate, area under the curve and volume distribution (Obrezanova 2023). It has been used to predict human intravenous, subcutaneous and oral %F, demonstrating promising potential for accelerating drug development (Zou 2023; Lou and Hageman 2021).
However, applying ML for predicting human oral %F has been challenging, with previous work predominantly treating the prediction as a classification task, labelling compounds as either “high” or “low” %F using 50% as an arbitrary cut-off threshold (Yang et al. 2024; Wei et al. 2025). This binary approach obscures clinically meaningful differences. For example, a compound with 5% bioavailability and one with 45% are treated as equivalent. Similarly, compounds with 50% and 90% bioavailability are treated as one. Such imprecision categorisation is particularly limiting at the extremes, where the practical consequences for dosing, formulation and therapeutic window are most pronounced. A compound misclassified near the decision boundary could be prematurely deprioritised or, conversely, advanced with an inappropriately designed dosing regimen. Beyond classification boundaries, the continuous distribution of %F across a compound series carries actionable information, enabling scientists to track incremental improvements during lead optimisation and allowing researchers in pharmacokinetics to estimate required doses before any human data are available.
Regression tasks, by contrast, preserve this quantitative resolution, providing a predicted numerical value that can be directly propagated into pharmacokinetic modelling and dose-exposure simulations. Therefore, an in silico tool capable of predicting %F as a continuous value is not only more informative than its classification counterpart, but more pertinent to the practical decisions that determine whether a drug candidate progresses through the development pipeline (Bassani et al. 2024). Previous work has attempted to develop regression ML models to predict human oral %F as continuous values. Wei et al. (2025) used molecular descriptors as inputs and reported a coefficient of determination (R2) of 0.08, which they stated as being unsatisfactory, and subsequently pivoted to a classification task. Yang et al. (2024) also applied a regression task using structural fingerprints as inputs but obtained R2 ranging from −0.3575 to 0.2623, which was also considered unsuitable, with negative R2 indicating predictions worse than randomly guessing. Evidently, treating the prediction of human oral %F as a regression task is indeed challenging. However, previous studies have lacked a comprehensive ML analysis, such as exploring different input features and regression methods.
Herein, we perform a comprehensive ML regression analysis to predict human oral %F as a continuous variable. We experimented with different input features, such as molecular descriptors, animal %F data and structure–activity behaviour. Furthermore, we explored different data pre-processing methods, such as data imputation, feature selection and logit transformation. The study demonstrates that a successful prediction of human oral %F using ML regression models can be achieved, thereby demonstrating the applicability of ML for advancing a key stage of the drug development pipeline.
Download the full article as PDF here Animal versus human oral drug bioavailability
or continue reading here
Mohammed, Y.A., Heaton-Ward, M., Gaisford, S. et al. Animal versus human oral drug bioavailability: machine learning says yes – they correlate. AAPS Open 12, 54 (2026). https://doi.org/10.1186/s41120-026-00193-z
25th September is World Pharmacists Day:
World Pharmacists Day 2026












































All4Nutra







