Project Name
Sparse Operations for Large-Language Models and Other Machine Learning Methods
Project Goal
This project aims to optimize core computations in common machine learning algorithms, such as clustering, classification, and large language model (LLM) inference. The goal is to develop efficient implementations for low-level operations and to establish general techniques that enhance the efficiency of model training and inference in real-world applications.
Project Description
1. Optimize operation efficiency * Identify computational bottlenecks in the K-Means algorithm and in deep learning models for both sparse and dense data. * Develop efficient implementations for low-level operations, such as efficient attention mechanisms for transformer architectures. 2. Enhance the efficiency of model training and inference * Reduce Gradient Bias: Investigate the bias in stochastic gradient descent (SGD) and develop debiasing techniques. * Exploit Model Sparsity: Identify inference bottlenecks of large language model (LLM) inference and create a library for model sparsification.