This paper introduces and analyzes several feature extraction algorithms. These algorithms use linear or non-linear feature extraction methods to project high-dimensional objects into lower dimensional space, thus the...This paper introduces and analyzes several feature extraction algorithms. These algorithms use linear or non-linear feature extraction methods to project high-dimensional objects into lower dimensional space, thus the complexity of the operations upon them, such as clustering, the nearest-neighbor search, visualization and etc can be reduced. The paper also presents some comparative experimental results of these algorithms and analyzes briefly their advantages or shortcomings.展开更多
提出一种新的稀疏谱聚类算法——基于PAM算法的HSSPAM聚类(high-dimensional sparse spectral clustering based on partitioning around medoids).该算法先用高相关系数过滤及主成分分析降维方法以有效减小甚至消除维度灾难对高维数据...提出一种新的稀疏谱聚类算法——基于PAM算法的HSSPAM聚类(high-dimensional sparse spectral clustering based on partitioning around medoids).该算法先用高相关系数过滤及主成分分析降维方法以有效减小甚至消除维度灾难对高维数据处理的影响,再采用Minkowski距离指数变换函数及稀疏化算法来构建分块对角矩阵以重新解释样本之间的相似度;然后构造新颖的拉普拉斯矩阵以实现进一步压缩数据矩阵,进而结合partitioning around medoids(PAM)算法取代传统谱聚类中的K-means算法对特征向量聚类以提高算法的聚类稳定性;最后引入高维基因数据设计了实验,并以不同的聚类评价指标来衡量该研究算法的聚类质量,实验结果表明,新算法能够更精确、更稳定地对基因数据聚类.展开更多
文摘This paper introduces and analyzes several feature extraction algorithms. These algorithms use linear or non-linear feature extraction methods to project high-dimensional objects into lower dimensional space, thus the complexity of the operations upon them, such as clustering, the nearest-neighbor search, visualization and etc can be reduced. The paper also presents some comparative experimental results of these algorithms and analyzes briefly their advantages or shortcomings.
文摘提出一种新的稀疏谱聚类算法——基于PAM算法的HSSPAM聚类(high-dimensional sparse spectral clustering based on partitioning around medoids).该算法先用高相关系数过滤及主成分分析降维方法以有效减小甚至消除维度灾难对高维数据处理的影响,再采用Minkowski距离指数变换函数及稀疏化算法来构建分块对角矩阵以重新解释样本之间的相似度;然后构造新颖的拉普拉斯矩阵以实现进一步压缩数据矩阵,进而结合partitioning around medoids(PAM)算法取代传统谱聚类中的K-means算法对特征向量聚类以提高算法的聚类稳定性;最后引入高维基因数据设计了实验,并以不同的聚类评价指标来衡量该研究算法的聚类质量,实验结果表明,新算法能够更精确、更稳定地对基因数据聚类.