BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260402T024533Z
LOCATION:3000\, Level 3
DTSTART;TZID=America/Los_Angeles:20250624T154500
DTEND;TZID=America/Los_Angeles:20250624T160000
UID:dac_DAC 2025_sess111_RESEARCH1279@linklings.com
SUMMARY:A Memory-Efficient LLM Accelerator with Q-K Correlation Prediction
  using Cluster-Based Associative Array for Selective KV Accessing
DESCRIPTION:Zikang Zhou, Kaiqi Chen, Xuyang Duan, and Jun Han (Fudan Unive
 rsity)\n\nAttention-based LLMs excel in text generation but face redundant
  computations in autoregressive token generation. While KV cache mitigates
  this, it introduces increased memory access overhead as sequences grow. W
 e propose Sella, a hardware-software co-design using cluster-based associa
 tive arrays to predict Q-K correlations, enabling selective KV cache acces
 s and reducing memory access without retraining. Sella includes a speciali
 zed accelerator featuring a prediction engine to improve performance and e
 nergy efficiency. Experiments show Sella achieves 2.1x, 93.8x, 31.4x, and 
 53.5x speedup over SpAtten, Sanger, TITAN RTX GPU, and Xeon CPU, respectiv
 ely, reducing off-chip memory access by up to 66% with negligible accuracy
  loss.\n\nTopics: AI\n\nTracks: AI3: AI/ML Architecture Design\n\nSession 
 Chairs: Ziang Yin (Arizona State University) and Yulhwa Kim (Sungkyunkwan 
 University)\n\n
END:VEVENT
END:VCALENDAR
