Presentation
Rethinking the Distribution of Outliers in Large Language Models: An In-depth Study
DescriptionOutliers in large language models (LLMs) significantly impact performance, particularly in quantization and compression. These outliers often cause substantial quantization errors, degrading model accuracy and limiting deployment on edge devices or specialized hardware. Two common types, massive activations and channel-wise outliers, pose significant challenges. While various quantization algorithms aim to mitigate their effects, few studies deeply explore their root causes.
This paper investigates the formation mechanisms of these outliers and introduces strategies to address them. We propose efficient methods to eliminate most massive activations and channel-wise outliers, enhancing the quantization process and facilitating more effective and accurate model deployment.
This paper investigates the formation mechanisms of these outliers and introduces strategies to address them. We propose efficient methods to eliminate most massive activations and channel-wise outliers, enhancing the quantization process and facilitating more effective and accurate model deployment.
Event Type
Networking
Work-in-Progress Poster
TimeMonday, June 236:00pm - 7:00pm PDT
LocationLevel 2 Lobby


