Development of Novel Methods for QSAR Modeling by Machine Learning Repeatedly: A Case Study on Drug Distribution to Each Tissue

Koichi Handa; Saki Yoshimura; Michiharu Kageyama; Takeshi Iijima

doi:10.26434/chemrxiv-2023-qsrxp-v2

Theoretical and Computational Chemistry

Search within Theoretical and Computational Chemistry

Development of Novel Methods for QSAR Modeling by Machine Learning Repeatedly: A Case Study on Drug Distribution to Each Tissue

04 April 2024, Version 2

Working Paper

Show author details

This content is a preprint and has not undergone peer review at the time of posting.

Abstract

AI is expected to help identify excellent candidates in drug discovery. However, we face a lack of data as it is time consuming and expensive to acquire raw data perfectly for many compounds. Hence, we tried to develop a novel QSAR method to predict a parameter more precisely from an incomplete dataset via optimizing data handling by making use of predicted explanatory variables. As a case study we focused on the tissue-to-plasma partition coefficient (Kp), which is an important parameter for understanding drug distribution in tissues and building the physiologically based pharmacokinetic (PBPK) model, is a representative of small and sparse datasets. In this study, we predicted the Kp values of 119 compounds in nine tissues (adipose, brain, gut, heart, kidney, liver, lung, muscle, and skin), while some of these were not available. To fill the missing values in Kp for each tissue, firstly we predicted those Kp values by the non-missing dataset using a random forest (RF) model with in vitro parameters (log P, fu, Drug Class, and fi) like a classical prediction by a QSAR model. Next, to predict the tissue-specific Kp values in a test dataset, we constructed a second RF model with not only in vitro parameters but also the Kp values of other tissues (i.e. other than target tissues) predicted by the first RF model as explanatory variables. Furthermore, we tested all possible combinations of explanatory variables and selected the model with the highest predictability from the test dataset as the final model. The evaluation of Kp prediction accuracy based on the root-mean-square error and R2-value revealed that the proposed models outperformed other machine learning methods, such as the conventional RF and message-passing neural networks. Significant improvements were observed in the Kp values of adipose tissue, brain, kidney, liver, and skin. These improvements indicated that the Kp information of other tissues can be used to predict the same for a specific tissue. Additionally, we found a novel relationship between each tissue by evaluating all combinations of explanatory variables. In conclusion, we developed a novel RF model to predict Kp values. We hope that this method will be applied to various problems in the field of experimental biology which often contains missing values in the near future.

Keywords

QSAR

machine learning

tissue-to-plasma partition coefficient

random forest

message passing neural network

missing values

classification of tissues

Supplementary materials

Title

Description

Actions

Title

Supporting Information

Description

Supporting Tables

Actions

Comments

Comments are not moderated before they are posted, but they can be removed by the site moderators if they are found to be in contravention of our Commenting Policy - please read this policy before you post. Comments should be used for scholarly discussion of the content in question. You can find more information about how to use the commenting feature here .

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Version History

Apr 04, 2024 Version 2

Nov 09, 2023 Version 1

Version Notes

We have added some tables to have readers understand smoothly.

Metrics

732

395

Views

Downloads

Citations

License

The content is available under CC BY 4.0

DOI

10.26434/chemrxiv-2023-qsrxp-v2

Author’s competing interest statement

The author(s) have declared they have no conflict of interest with regard to this content

Ethics

The author(s) have declared ethics committee/IRB approval is not relevant to this content

Development of Novel Methods for QSAR Modeling by Machine Learning Repeatedly: A Case Study on Drug Distribution to Each Tissue

Authors

Abstract

Keywords

Supplementary materials

Comments

Version History

Version Notes

Metrics

License

DOI

Author’s competing interest statement

Ethics

Share