BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260402T024532Z
LOCATION:3002\, Level 3
DTSTART;TZID=America/Los_Angeles:20250624T154500
DTEND;TZID=America/Los_Angeles:20250624T160000
UID:dac_DAC 2025_sess117_RESEARCH795@linklings.com
SUMMARY:VSpGEMM: Exploiting Versal ACAP for High-Performance SpGEMM Accele
 ration
DESCRIPTION:Kai Shi (Beijing University of Posts and Telecommunications); 
 Zhe Lin (Sun Yat-sen University); and Xinya Luan, Jianwang Zhai, and Kang 
 Zhao (Beijing University of Posts and Telecommunications)\n\nSparse genera
 l matrix-matrix multiplication (SpGEMM) serves as a fundamental operation 
 in real-world applications such as deep learning. Different from general m
 atrix multiplication, matrices in SpGEMM are highly sparse and therefore r
 equire a compact representation. This places an additional burden on data 
 preprocessing and exchanging and also causes irregular memory access patte
 rns, which can in turn lead to communication and computation bottlenecks. 
 To break these bottlenecks, we present VSpGEMM, a hardware accelerator for
  SpGEMM that is tailored and optimized on Versal ACAP. Firstly, a new stor
 age format called BCSX is proposed in VSpGEMM, which offers a unified and 
 block-wise compression strategy to deal with both row-major and column-maj
 or representation of non-zero data, enabling fixed-pattern memory accesses
  and effective data preloading. Secondly, a multi-level tiling mechanism i
 s introduced to decompose the holistic SpGEMM into multiple computation gr
 anularities that fit into the AI Engines (AIEs) on Versal in a hierarchica
 l manner, enhancing data reuse. Thirdly, a hybrid partitioning scheme is p
 resented to orchestrate both the AIEs and programmable logic (PL) for inte
 rmediate product merging, which together resolve the issues of high memory
  utilization and communication demand. Experimental results demonstrate a 
 2.65× speedup over state-of-the-art (SOTA) GEMM design on Versal and an av
 erage 33.62× improvement in energy efficiency compared to cuSPARSE on RTX 
 4090 GPU, showing the efficacy of VSpGEMM.\n\nTopics: Design\n\nTracks: DE
 S1: SoC, Heterogeneous, and Reconfigurable Architectures\n\nSession Chairs
 : Tianhao Cai (Beihang University) and Dirk Stroobandt (Ghent University)\
 n\n
END:VEVENT
END:VCALENDAR
