Document Type : Original Article
Authors
1
Department of Materials Engineering, Isfahan University of Technology, Isfahan, Iran
2
Department of Computer Science, Faculty of Mathematical Sciences, Shahrekord University, Shahrekord, Iran
Abstract
The accurate calculation of band gaps is necessary for the quick selection and design of inorganic materials; however, the first-principles calculations pose a computational challenge for extensive chemical space. In this work, a machine-learning framework was developed to predict optB88-vdW band gaps using composition- and structure-based descriptors obtained from the JARVIS-DFT database. After data preprocessing and feature extraction, unsupervised clustering was applied to identify the natural grouping in the materials space. Next, a series of regression models, such as Linear Regression, Random Forest, XGBoost, TabNet, and Residual Multilayer Perceptron, were developed. Evaluation metrics used for measuring model performance included MAE, RMSE, and R^2, while interpretability was additionally analyzed via intrinsic and permutation-based feature importance analyses. Out of all models that were considered, XGBoost provided the highest performance, with the achieved results of MAE=0.399 eV, RMSE=0.582 eV, and R2=0.893. The feature importance analysis revealed formation energy per atom, oxidation state descriptors, halogen descriptors, magnetic moment, and transition metal descriptors as important features. Overall, the results demonstrate that nonlinear machine-learning models can provide accurate and interpretable predictions of optB88-vdW band gaps from relatively inexpensive materials descriptors. These findings also suggest that the proposed workflow can serve as a practical screening tool for large inorganic-material datasets by helping prioritize promising compounds before more computationally expensive electronic-structure calculations are performed. In addition, combining predictive performance with feature-level interpretation can facilitate a better understanding of composition–structure–property relationships and support more efficient data-driven materials discovery.
Keywords
Subjects