Acta Phys. -Chim. Sin.

Previous Articles     Next Articles

Data distribution matters: enhancing model generalization in datadriven materials discovery by intentionally exploiting both positive and negative data

Yiheng Wu1,2, Lijun Liu1, Wei Deng1, Haodong Jiang1,2, Ruotong Li1,2, Ming Hu1, Bo Peng1, Dapeng Wang1,2   

  1. 1 State Key Laboratory of Polymer Science and Technology, Changchun Institute of Applied Chemistry, Chinese Academy of Sciences, Changchun 130022, Jilin Province, China;
    2 School of Applied Chemistry and Engineering, University of Science and Technology of China, Hefei 230026, Anhui Province, China
  • Received:2026-01-11 Revised:2026-03-21 Accepted:2026-03-22
  • Contact: Bo Peng, Dapeng Wang E-mail:bpeng2019@ciac.ac.cn;wdp@ciac.ac.cn

Abstract: The distribution of training data is a fundamental yet often overlooked factor governing the success of machine learning models in materials science. This study evaluates the influence of data distribution on predictive modeling by benchmarking seven static and adaptive sampling strategies across high-dimensional benchmark functions and the real-world materials database. The core finding is that sampling strategies that simultaneously focus on both positive and negative extremum regions significantly enhance model generalization, outperforming both uniform sampling and unidirectional strategies that target only one extremum region (either positive or negative). These results suggest that actively shaping the data distribution to explore “undesirable” material regions is neither a waste of resources nor just marginally beneficial. Instead, it is a well-considered move that paves the way for a more efficient data-driven discovery workflow.

Key words: Data science, Machine learning, Data distribution, Data-driven materials discovery