2025 Volume 14 Issue 2
Creative Commons License

Machine Learning QSAR Under Dataset Shift and Chemical Novelty


, ,
  1. Department of ML QSAR and Dataset Shift, School of Pharmacy, University College Cork, Cork, Ireland.
  2. Department of Chemical Novelty and Model Generalization, Faculty of Pharmacy, University of Galway, Galway, Ireland.
Abstract

Machine-learning quantitative structure–activity relationship models are commonly compared under validation schemes that only partially reproduce the chemical, assay, temporal, and local structure–activity changes encountered during prospective drug discovery. Apparent performance can therefore reflect the difficulty of the selected test distribution as strongly as the intrinsic capability of the model. Peer-reviewed literature published from 2017 through 2025 was systematically examined for evidence concerning molecular-property and bioactivity prediction under scaffold displacement, assay or property change, chronological separation, external chemical novelty, activity cliffs, applicability-domain analysis, and predictive uncertainty. Model representation, validation design, distribution-shift mechanism, applicability assessment, and uncertainty treatment were coded separately. Source-level review-flow counts were not reconstructed from the final bibliography and are consequently not reported as numerical selection statistics. The evidence does not support a single hierarchy of QSAR model classes or a universal definition of out-of-distribution performance. Random splitting can overstate transferable performance, but scaffold splitting is not uniformly equivalent to chronological or prospective validation. Local activity cliffs constitute a separate challenge from broad scaffold novelty. Assay context, property range, representation, transfer strategy, and chemical-space support further modify apparent generalization. Applicability-domain membership and predictive uncertainty are informative but cannot independently establish reliability. QSAR evaluation under chemical novelty should begin by defining the mechanism producing distribution shift rather than by selecting a conventional splitter by default. Model comparison becomes more decision-relevant when global chemical support, local structure–activity discontinuity, assay context, temporal separation, model representation, and uncertainty are evaluated as distinct but interacting dimensions.


How to cite this article
Vancouver
O'Connor P, Murphy G, Walsh N. Machine Learning QSAR Under Dataset Shift and Chemical Novelty. Int J Pharm Res Allied Sci. 2025;14(2):148-63. https://doi.org/10.51847/n23TOInXCc
APA
O'Connor, P., Murphy, G., & Walsh, N. (2025). Machine Learning QSAR Under Dataset Shift and Chemical Novelty. International Journal of Pharmaceutical Research and Allied Sciences, 14(2), 148-163. https://doi.org/10.51847/n23TOInXCc
Related articles:
Most viewed articles:
Issue 1 Volume 16 (2027)