Skip to content

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Description

As datasets continue to grow in size and complexity, the use of traditional supervised learning techniques for large-scale research efforts is becoming increasingly expensive and time consuming. To address these challenges, many projects have begun to look towards unsupervised learning methods which are better equipped for unlabeled and complex data. Among these methods, K-Means clustering has perhaps been the most well-adopted, largely due to its simplicity, flexibility, and performance. However, the default K-Means algorithm is not without its limitations. Most notably, the random initialization of cluster centroids exposes the algorithm to getting caught in local, non-optimal minima. This can severely harm clustering performance, especially with highly-clustered or high-dimensional data (where clusters may be very close together). Our research attempts to better understand the impact of this problem and how it might be addressed by comparing the performance between K-Means and two alternative implementations: K-Means++ and the Hartigan Wong K-Means algorithm.

Taken from CSCI 347 Final Project Report

Repository Structure

The repository is broken up into four sections: algorithms, data, utilities, and the experiment.

  • Algorithms: This contains an implementation of the base K-Means algorithm as well as our implementation of the three initialization methods under consideration
  • Data: This contains the Original Breast Cancer Wisconsin dataset, sourced from the UCI Machine Learning Repository
  • Utilities: This directory contains a set of reusable utility classes that abstract common and complex functionality used throughout our experiments
  • Experiment: This Jupyer notebook contains the implementation of our experiment as documented in the CSCI 347 Final Project Report

About

A custom implementation of the KMeans algorithm and three initialization methods (K-Means++, Hartigan Wong, and random initialization) that was used to explore the impact of initialization methodologies in the clustering process.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages