Close

Session

Research Manuscript
:
Optimizing Large Language Models: Speed, Size, and Smarts
DescriptionThis session highlights advancements in optimizing large language models (LLMs) for more efficient inference and communication. The papers cover strategies such as pruning, quantization, and adaptive sparse gradient compression, along with methods to reduce communication costs and improve inference speed. These techniques address challenges in long-context processing, multi-modal inference, and chip-level optimizations, all crucial for scaling LLMs in real-world applications.
Event Type
Research Manuscript
TimeMonday, June 231:30pm - 3:00pm PDT
Location3000, Level 3
Topics
AI
Tracks
AI1: AI/ML Algorithms