BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260402T024532Z
LOCATION:3000\, Level 3
DTSTART;TZID=America/Los_Angeles:20250624T170000
DTEND;TZID=America/Los_Angeles:20250624T171500
UID:dac_DAC 2025_sess111_RESEARCH2105@linklings.com
SUMMARY:BBAL: A Bidirectional Block Floating Point-Based Quantization Acce
 lerator for Large Language Models
DESCRIPTION:Xiaomeng Han (Southeast University); Yuan Cheng (Nanjing Unive
 rsity); Jing Wang, Junyang Lu, and Hui Wang (Southeast University); Xuanxi
  Zhang (Jilin Normal University); Ning Xu (Southeast University); Dawei Ya
 ng (Houmo); and Zhe Jiang (Southeast University)\n\nLarge language models 
 (LLMs), with their billions of pa-\nrameters, pose substantial challenges 
 for deployment on edge devices,\nstraining both memory capacity and comput
 ational resources. Block\nfloating-point (BFP) quantisation reduces memory
  and computational\noverhead by converting high-overhead floating-point op
 erations into low-\nbit fixed-point operations. However, BFP requires alig
 ning all data to the\nmaximum exponent, which causes loss of small and mod
 erate values,\nresulting in quantisation error and degradation in the accu
 racy of\nLLMs. To address this issue, we propose a Bidirectional Block Flo
 ating-\nPoint (BBFP) data format, which reduces the probability of selecti
 ng\nthe maximum as shared exponent, thereby reducing quantisation error.\n
 By utilizing the features in BBFP, we present a full-stack Bidirectional\n
 Block Floating Point-Based Quantisation Accelerator for LLMs (BBAL),\nprim
 arily comprising a PE array based on BBFP, paired with\nour proposed cost-
 effective nonlinear computation unit. Experimental\nresults show BBAL achi
 eves a 22% improvement in accuracy compared\nto an outlier-aware accelerat
 or at similar efficiency, and a 40% efficiency\nimprovement over a vanilla
  BFP-based accelerator at similar accuracy.\n\nTopics: AI\n\nTracks: AI3: 
 AI/ML Architecture Design\n\nSession Chairs: Ziang Yin (Arizona State Univ
 ersity) and Yulhwa Kim (Sungkyunkwan University)\n\n
END:VEVENT
END:VCALENDAR
