You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Parallel implementations of Lloyd's K-Means clustering algorithm using CUDA and OpenMP, with relative study of the performance and bottolneck given the architecture and code.
Run exactly N iterations instead of waiting for convergence. Makes timing reproducible across datasets. Without it the algorithm stops when centroids stop moving.
-O3
Optional (OMP only)
Enables vectorization and heavy optimizations.
-arch=sm_XX
Recommended (CUDA)
Target GPU architecture. Use sm_89 for RTX 40-series, sm_86 for RTX 30-series, sm_80 for A100. If omitted nvcc defaults to an older architecture and may miss optimizations.
Optional OpenMP environment variables
# Pin threads to physical cores (recommended for reproducibility)
OMP_PROC_BIND=close OMP_PLACES=cores OMP_NUM_THREADS=16 ./omp-k-means K in out
About
Architectures for AI (UniBo) - Optimization and parallization OpenMP and CUDA