Skip to content

Repository files navigation

Super K-Means
Paper Paper PyPI License License GitHub stars

A super-fast clustering library for high-dimensional vector embeddings

SuperKMeans vs FAISS and Scikit Learn

Why Super K-Means?

  • Faster clustering of vector embeddings (Cohere, OpenAI, MXBAI, CLIP, MiniLM) than FAISS.
  • Index 10M embeddings of 1024 dimensions in less than a minute on a single CPU.
  • Faster without compromising clustering quality.
  • Support for quantized clustering (8-bit Scalar Quantization, LVQ, and RabitQ)
  • Efficient on CPUs (ARM and x86) and GPUs.

Our secret sauce

  • Carefully interleaving GEMM routines and pruning kernels that prune dimensions efficiently
  • In the benchmarks you see in the cover image, all algorithms are clustering the same data: No dimensionality reduction, no sampling, no early-termination.

Usage

from superkmeans import SuperKMeans

data = ... # Numpy 2D matrix
k = 1000
d = 768

kmeans = SuperKMeans(
    n_clusters=k,
    dimensionality=d
)

# Run the clustering
centroids = kmeans.train(data) # 2D array with centroids (k x d) 

# Get assignments
assignments = kmeans.assign(data)

Then, you can use the centroids to create an IVF index for Vector Search, for example, in FAISS.

Usage in C++
#include <vector>
#include <cstddef>
#include "superkmeans/superkmeans.h"
#include "superkmeans/hierarchical_superkmeans.h"

int main(int argc, char* argv[]) {
    std::vector<float> data; // Fill
    size_t n = 1000000;
    size_t k = 10000;
    size_t d = 768;

    auto kmeans = skmeans::SuperKMeans(k, d);

    // Or Hierarchical Super K-Means for extreme performance:
    // auto kmeans = skmeans::HierarchicalSuperKMeans(k, d);
    
    // Run the clustering
    std::vector<float> centroids = kmeans.Train(data.data(), n);
    
    // Assign points
    std::vector<uint32_t> assignments = kmeans.Assign(data.data(), centroids.data(), n, k);
}

Check our examples for fully working examples in Python and C++.

Quantized Clustering

SuperKMeans supports clustering quantized vectors, substantially accelerating clustering and barely affecting clustering quality. We support 8-bit scalar quantization, LVQ, and RabitQ. You give us float32 vectors and we handle the rest:

kmeans = SuperKMeans(
    n_clusters=k,
    dimensionality=d,
    quantizer='rabitq'
)

centroids = kmeans.train(data) # We return float32 centroids
assignments = kmeans.quantized_assign(data, centroids)

Check our fully working examples in Python or C++.

Documentation

Check our wiki for advanced usage and API reference.

Installation

Python

pip install superkmeans

Tip

For maximum performance, we recommend compiling from source.

C++

As a header-only library with CMake FetchContent:

FetchContent_Declare(
    superkmeans
    GIT_REPOSITORY https://github.com/cwida/superkmeans
)
FetchContent_MakeAvailable(superkmeans)

target_link_libraries(myapp PRIVATE superkmeans)
Compiling Python Bindings from source

Prerequisites

  • Clang 17 or GCC 13
  • CMake 3.26
  • OpenMP
  • A BLAS implementation
  • Python 3 (only for Python bindings)
git clone https://github.com/cwida/SuperKMeans.git
cd SuperKMeans
git submodule update --init
pip install .

# Run plug-and-play example
python ./examples/simple_clustering.py

# Set a value for n, d and k
python ./examples/simple_clustering.py 200000 1536 1000
Compiling C++ library from source

Prerequisites

  • Clang 17 or GCC 13
  • CMake 3.26
  • OpenMP
  • A BLAS implementation
git clone https://github.com/cwida/SuperKMeans.git
cd SuperKMeans
git submodule update --init

# Set proper path to clang if needed
export CXX="/usr/bin/clang++-18" 

# Compile
cmake .
make examples

# Run plug-and-play example
cd examples
./simple_clustering.out

# Set a value for n, d and k
./simple_clustering.out 100000 1536 1000

For a more comprehensive installation and compilation guide, check INSTALL.md.

Getting the Best Performance

Check INSTALL.md.

Roadmap

We are actively developing Super K-Means and accepting contributions! Check CONTRIBUTING.md.

Benchmarking

To run our benchmark suite in C++, refer to BENCHMARKING.md.

Adoption

SuperKMeans' ideas have been adopted in:

Other implementations of SuperKMeans

Research behind SuperKMeans

A Super Fast K-means for Indexing Vector Embeddings

@article{kuffo2026super,
  title={A Super Fast K-means for Indexing Vector Embeddings},
  author={Kuffo, Leonardo and Hepkema, Sven and Boncz, Peter},
  journal={arXiv preprint arXiv:2603.20009},
  year={2026}
}

Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings

@article{kuffo2026superquantized,
  title={Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings},
  author={Kuffo, Leonardo and Boncz, Peter},
  journal={VLDB 2026 Workshop: The 2nd Workshop on Vector Databases},
  year={2026}
}

About

⚡ Super fast clustering for high-dimensional vectors on CPUs (x86, ARM) and GPUs — for Python and C++. 100x faster clustering of vector embeddings than FAISS

Topics

Resources

Contributing

Stars

73 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages