物理化学学报

上一篇    下一篇

数据分布至关重要:通过协同利用正负样本数据增强数据驱动材料发现中的模型泛化能力

武一衡1,2, 刘立军1, 邓伟1, 姜浩东1,2, 李若彤1,2, 胡铭1, 彭博1, 王大鹏1,2   

  1. 1 中国科学院长春应用化学研究所, 高分子科学与技术全国重点实验室, 吉林 长春 130022;
    2 中国科学技术大学应用化学与工程学院, 安徽 合肥 230026
  • 收稿日期:2026-01-11 修回日期:2026-03-21 录用日期:2026-03-22
  • 通讯作者: 彭博, 王大鹏 E-mail:bpeng2019@ciac.ac.cn;wdp@ciac.ac.cn
  • 基金资助:
    本研究由吉林省与中国科学院科技合作高技术产业化专项(2024SYHZ0036),国家自然科学基金(22473106)及吉林省科技发展计划项目(20240208014JH)共同资助

Data distribution matters: enhancing model generalization in datadriven materials discovery by intentionally exploiting both positive and negative data

Yiheng Wu1,2, Lijun Liu1, Wei Deng1, Haodong Jiang1,2, Ruotong Li1,2, Ming Hu1, Bo Peng1, Dapeng Wang1,2   

  1. 1 State Key Laboratory of Polymer Science and Technology, Changchun Institute of Applied Chemistry, Chinese Academy of Sciences, Changchun 130022, Jilin Province, China;
    2 School of Applied Chemistry and Engineering, University of Science and Technology of China, Hefei 230026, Anhui Province, China
  • Received:2026-01-11 Revised:2026-03-21 Accepted:2026-03-22
  • Contact: Bo Peng, Dapeng Wang E-mail:bpeng2019@ciac.ac.cn;wdp@ciac.ac.cn

摘要: 训练数据的分布是决定机器学习模型在材料科学领域应用成败的一个基础性但常被忽视的关键因素。本研究在高维基准函数和真实材料数据库上系统评估了七种静态与自适应采样策略,考察了数据分布对预测模型的影响。核心发现表明:同时覆盖正、负极值区域的采样策略可显著提升模型泛化能力,其表现优于均匀采样和仅针对单一极值区域(正或负)的单向采样策略。上述结果证明,主动塑造数据分布并探索“非理想”材料区域并非资源浪费,也并非仅带来有限的收益;相反,这是一种合理且有效的策略,有助于构建更高效的数据驱动材料发现工作流。

关键词: 数据科学, 机器学习, 数据分布, 数据驱动的材料发现

Abstract: The distribution of training data is a fundamental yet often overlooked factor governing the success of machine learning models in materials science. This study evaluates the influence of data distribution on predictive modeling by benchmarking seven static and adaptive sampling strategies across high-dimensional benchmark functions and the real-world materials database. The core finding is that sampling strategies that simultaneously focus on both positive and negative extremum regions significantly enhance model generalization, outperforming both uniform sampling and unidirectional strategies that target only one extremum region (either positive or negative). These results suggest that actively shaping the data distribution to explore “undesirable” material regions is neither a waste of resources nor just marginally beneficial. Instead, it is a well-considered move that paves the way for a more efficient data-driven discovery workflow.

Key words: Data science, Machine learning, Data distribution, Data-driven materials discovery