Presentation
An Algorithm-Hardware Co-design Based on Revised Microscaling Format Quantization for Accelerating Large Language Models
DescriptionThe narrow-bit-width data format is crucial for reducing the computation and storage costs of modern deep learning applications, particularly in large language models (LLMs) based applications. Microscaling (MX) format has been proven as a drop-in replacement for the baseline FP32 in existing inference frameworks, with low user friction. However, deploying such a new format into existing hardware systems is still challenging, and the dominant solution for LLM inference at low precision is still low-bit quantization. This particularly limits the strategic applications of such LLMs in real deployment on a large scale. In this work, we propose an algorithm-hardware co-design that adopts a two-level Revised MX Format Quantization (RMFQ) and a Revised MX Format Accelerator (RMFA) architecture design. RMFQ proposes the revised MX (RMX) format and provides a novel quantization framework with innovative group direction. Also, RMFA provides an RMX adaptive hardware architecture and an RMX encoding scheme. As a result, RMFQ pushes the limit of 4-bit and 6-bit quantization to a new state-of-the-art, and RMFA surpasses the existing outlier-aware accelerator such as OliVe, achieving a 1.28× speedup and a 1.31× energy reduction.
Event Type
Research Manuscript
TimeTuesday, June 245:15pm - 5:30pm PDT
Location3000, Level 3
Similar Presentations


