BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260402T024534Z
LOCATION:3001\, Level 3
DTSTART;TZID=America/Los_Angeles:20250623T113000
DTEND;TZID=America/Los_Angeles:20250623T114500
UID:dac_DAC 2025_sess113_RESEARCH1090@linklings.com
SUMMARY:DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FP
 GA-based LLMs Acceleration
DESCRIPTION:Zhuoquan Yu, Huidong Ji, Yue Cao, Junfu Wu, Xiaoze Yan, Lirong
  Zheng, and Zhuo Zou (Fudan University)\n\nQuantization enables efficient 
 deployment of large language models (LLMs) on FPGAs, but its presence of o
 utliers affects the accuracy of the quantized model. Existing methods main
 ly deal with outliers through channel-wise or token-wise isolation and enc
 oding, which leads to expensive dynamic quantization. To address this prob
 lem, we introduce DuoQ, an FPGA-oriented algorithm-hardware co-design fram
 ework. DuoQ effectively eliminates outliers through learnable equivalent t
 ransformations and low-semantic token awareness in the quantization scheme
  part, facilitating per-tensor quantization with 4-bits. We co-design the 
 quantization algorithm and hardware architecture. Specifically, DuoQ accel
 erates end-to-end LLM through a novel DSP-aware PE unit design and encoder
  design. In addition, two types of post-processing units assist in the rea
 lization of nonlinear functions and dynamic token awareness. Experimental 
 results show that compared with platforms with different architectures, Du
 oQ's computational efficiency and energy efficiency are improved by up to 
 8.8x and 23.45x. In addition, DuoQ has achieved accuracy improvements comp
 ared to other outlier-aware software and hardware works.\n\nTopics: AI\n\n
 Tracks: AI4: AI/ML System and Platform Design\n\nSession Chairs: Chaojian 
 Li (Georgia Institute of Technology) and Zhongzhi Yu (Nvidia)\n\n
END:VEVENT
END:VCALENDAR
