BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260402T024533Z
LOCATION:3001\, Level 3
DTSTART;TZID=America/Los_Angeles:20250625T164500
DTEND;TZID=America/Los_Angeles:20250625T170000
UID:dac_DAC 2025_sess132_RESEARCH196@linklings.com
SUMMARY:OutlierCIM: Outlier-Aware Digital CIM-Based LLM Accelerator with H
 ybrid-Strategy Quantization and Unified FP-INT Computation
DESCRIPTION:Zihan Zou, Shikuang Chen, Chen Zhang, Xing Wang, Zhichao Liu, 
 Haoran Du, Xin Si, Hao Cai, and Bo Liu (Southeast University)\n\nActivatio
 n outliers in Large Language Models (LLMs), which exhibit large magnitudes
  but small quantities, significantly affect model performance and pose cha
 llenges for the acceleration of LLMs. To address this bottleneck, research
 ers have proposed several co-design frameworks with outlier-aware algorith
 ms and dedicated hardware. However, they face challenges balancing model a
 ccuracy with hardware efficiency when accelerating LLMs in a low bit-width
  manner. To this end, we propose OutlierCIM, the first algorithm and hardw
 are co-design framework for compute-in-memory (CIM) accelerator with outli
 er-aware quantization algorithm. The key contributions of OutlierCIM are 1
 ) an outlier-clustered tiling strategy that regulates memory access and re
 duces inefficient workloads which are both introduced by outliers, 2) a hy
 brid-strategy quantization and a reconfigurable double-bit CIM macro array
  that overcome the low storage utilization and high latency of outlier-bas
 ed LLM quantization, and 3) a quantization factor post-processing strategy
  and a dedicated quantizer that efficiently unify the multiplication and a
 ccumulation of outlier-caused FP-INT workloads. Implemented in a 28nm CMOS
  technology, OutlierCIM occupies an area of 2.25 mm². When evaluated at co
 mprehensive benchmarks, OutlierCIM achieves up to 4.54× energy efficiency 
 improvement and 3.91× speedup compared to the state-of-the-art outlier-awa
 re accelerators.\n\nTopics: Design\n\nTracks: DES2B: In-memory and Near-me
 mory Computing Architectures, Applications and Systems\n\nSession Chairs: 
 Steve Dai (NVIDIA) and Haitong Li (Purdue University)\n\n
END:VEVENT
END:VCALENDAR
