BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260402T024533Z
LOCATION:3000\, Level 3
DTSTART;TZID=America/Los_Angeles:20250624T171500
DTEND;TZID=America/Los_Angeles:20250624T173000
UID:dac_DAC 2025_sess111_RESEARCH2166@linklings.com
SUMMARY:An Algorithm-Hardware Co-design Based on Revised Microscaling Form
 at Quantization for Accelerating Large Language Models
DESCRIPTION:Yingbo Hao (South China University of Technology), Huangxu Che
 n (Hong Kong University of Science and Technology (HKUST)), and Yi Zou and
  Yanfeng Yang (South China University of Technology)\n\nThe narrow-bit-wid
 th data format is crucial for reducing the computation and storage costs o
 f modern deep learning applications, particularly in large language models
  (LLMs) based applications. Microscaling (MX) format has been proven as a 
 drop-in replacement for the baseline FP32 in existing inference frameworks
 , with low user friction. However, deploying such a new format into existi
 ng hardware systems is still challenging, and the dominant solution for LL
 M inference at low precision is still low-bit quantization. This particula
 rly limits the strategic applications of such LLMs in real deployment on a
  large scale. In this work, we propose an algorithm-hardware co-design tha
 t adopts a two-level Revised MX Format Quantization (RMFQ) and a Revised M
 X Format Accelerator (RMFA) architecture design. RMFQ proposes the revised
  MX (RMX) format and provides a novel quantization framework with innovati
 ve group direction. Also, RMFA provides an RMX adaptive hardware architect
 ure and an RMX encoding scheme. As a result, RMFQ pushes the limit of 4-bi
 t and 6-bit quantization to a new state-of-the-art, and RMFA surpasses the
  existing outlier-aware accelerator such as OliVe, achieving a 1.28× speed
 up and a 1.31× energy reduction.\n\nTopics: AI\n\nTracks: AI3: AI/ML Archi
 tecture Design\n\nSession Chairs: Ziang Yin (Arizona State University) and
  Yulhwa Kim (Sungkyunkwan University)\n\n
END:VEVENT
END:VCALENDAR
