摘要
针对非平衡数据存在的类内不平衡、噪声、生成样本覆盖面小等问题,提出了基于层次密度聚类的去噪自适应混合采样算法(adaptive denoising hybrid sampling algorithm based on hierarchical density clustering,ADHSBHD).首先引入HDBSCAN聚类算法,将少数类和多数类分别聚类,将全局离群点和局部离群点的交集视为噪声集,在剔除噪声样本之后对原数据集进行处理,其次,根据少数类样本中每簇的平均距离,采用覆盖面更广的采样方法自适应合成新样本,最后删除一部分多数类样本集中的对分类贡献小的点,使数据集均衡.ADHSBHD算法在7个真实数据集上进行评估,结果证明了其有效性.
As imbalanced data are exposed to problems such as intra-class imbalance,noise,and small coverage of generated samples,an adaptive denoising hybrid sampling algorithm based on hierarchical density clustering(ADHSBHD)is proposed.Firstly,the clustering algorithm HDBSCAN is introduced to perform clustering on minority classes and majority classes separately;the intersection of global and local outliers is regarded as the noise set,and the original data set is processed after noise samples are eliminated.Secondly,according to the average distance between clusters of samples in minority classes,the adaptive sampling method with broader coverage is used to synthesize new samples.Finally,some points that contribute little to the classification of majority classes are deleted to balance the dataset.The ADHSBHD algorithm is evaluated on six real data sets,and the results can prove its effectiveness.
作者
姜新盈
王舒梵
严涛
JIANG Xin-Ying;WANG Shu-Fan;YAN Tao(School of Mathematics,Physics and Statistics,Shanghai University of Engineering Science,Shanghai 201620,China)
出处
《计算机系统应用》
2022年第10期206-210,共5页
Computer Systems & Applications
关键词
不平衡数据
分类
聚类
混合采样
imbalanced data
classification
cluster
hybrid sampling
作者简介
通信作者:姜新盈,E-mail:jxynovelty@163.com