BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260402T024509Z
LOCATION:3002\, Level 3
DTSTART;TZID=America/Los_Angeles:20250624T153000
DTEND;TZID=America/Los_Angeles:20250624T173000
UID:dac_DAC 2025_sess117@linklings.com
SUMMARY:Back to the Future: Where Speed Meets Efficiency
DESCRIPTION:This session explores cutting-edge advancements in hardware ac
 celeration, focusing on optimizing computation, memory access, and paralle
 lism in modern architectures. Featuring research on heterogeneous reconfig
 urable accelerators and FPGA optimization, the papers highlight novel appr
 oaches to accelerating key computational tasks. Topics include efficient s
 parse matrix multiplication, inter-tile parallelism, adaptive tree computa
 tions, large number modular reduction,  compiler mapping strategies for CG
 RAs and physical design for nonvolatile FPGAs. Together, these works demon
 strate how innovations in hardware and algorithm design are driving the fu
 ture of high-performance computing, pushing the boundaries of speed, effic
 iency, and scalability in diverse applications.\n\nALLMod: Exploring \unde
 rline{A}rea-Efficiency of \underline{L}UT-based \underline{L}arge Number \
 underline{Mod}ular Reduction via Hybrid Workloads\n\nModular arithmetic, p
 articularly modular reduction, is widely used in cryptographic application
 s such as homomorphic encryption (HE) and zero-knowledge proofs (ZKP). Hig
 h-bit-width operations are crucial for enhancing security; however, they a
 re computationally intensive due to the large number of m...\n\n\nFangxin 
 Liu, Haomin Li, and Zongwu Wang (Shanghai Jiao Tong University); Bo Zhang,
  Mingzhe Zhang, and Shoumeng Yan (Ant Group); and Li Jiang (Shanghai Jiao 
 Tong University)\n---------------------\nHiSpTRSV: Exploring Tile-Level Pa
 rallelism for SpTRSV Acceleration on FPGAs\n\nSparse Triangular Solve (SpT
 RSV) is a critical level-2 kernel in sparse Basic Linear Algebra Subprogra
 ms (BLAS). While Field-Programmable Gate Array (FPGA) accelerators for SpT
 RSV focus on optimizing individual tiles, they overlook inter-tile paralle
 lism. Designing an inter-tile parallelism accelera...\n\n\nFan Sun, Fang D
 ong, and Dian Shen (Southeast University)\n---------------------\nGPS: GNN
 -Based Two-Stage Pre-Scheduling Loop Mapping Method on CGRAs\n\nCoarse-gra
 ined reconfigurable architecture (CGRA) has emerged as a promising solutio
 n for accelerating computationally intensive applications, particularly in
  the field of artificial intelligence. One of the primary challenges for C
 GRA compilers is generating effective mapping results for complex ap...\n\
 n\nMingyang Kou and Weiqing Ji (University of Science and Technology of Ch
 ina), Shouyi YIN (Tsinghua University), and Hailong Yao (University of Sci
 ence and Technology of China)\n---------------------\nA Data-Centric Hardw
 are Accelerator for Efficient Adaptive Radix Tree\n\nThe Adaptive Radix Tr
 ee (ART) is a widely used tree index structure prevalent in various domain
 s such as databases and key-value stores. Despite many solutions have been
  proposed to improve the performance of ART, they still suffer from signif
 icant redundant tree traversals and serious synchronizati...\n\n\nJin Zhao
 , Yu Zhang, Jun Huang, weihang yin, Hui Yu, Hao Qi, and Zixiao Wang (Huazh
 ong University of Science and Technology); longlong lin (Southwest Univers
 ity); and Xiaofei Liao and Hai Jin (Huazhong University of Science and Tec
 hnology)\n---------------------\nHeteroSVD: Efficient SVD Accelerator on V
 ersal ACAP with Algorithm-Hardware Co-Design\n\nSingular value decompositi
 on (SVD) is a matrix factorization technique widely used in signal process
 ing and recommendation systems, etc. In general, the time complexity of SV
 D algorithms is cubic to\nthe problem size, making SVD algorithms difficul
 t to meet stringent performance requirements in real-...\n\n\nXinya Luan (
 Beijing University of Posts and Telecommunications); Zhe Lin (Sun Yat-sen 
 University); and Kai Shi, Jianwang Zhai, and Kang Zhao (Beijing University
  of Posts and Telecommunications)\n---------------------\nRewire: Advancin
 g CGRA Mapping Through a Consolidated Routing Paradigm\n\nCoarse-Grained R
 econfigurable Arrays (CGRA) balance the performance and power efficiency i
 n computing systems. Effective compilers play a crucial role in fully real
 izing its potential. The compiler maps Data Flow Graphs (DFG), which repre
 sent compute-intensive loop kernels, onto CGRAs. However, exis...\n\n\nZha
 oying Li, Dan Wu, Dhananjaya Wijerathne, Dan Chen, and Huize Li (National 
 University of Singapore); Cheng Tan (Google); and Tulika Mitra (National U
 niversity of Singapore)\n---------------------\nVSpGEMM: Exploiting Versal
  ACAP for High-Performance SpGEMM Acceleration\n\nSparse general matrix-ma
 trix multiplication (SpGEMM) serves as a fundamental operation in real-wor
 ld applications such as deep learning. Different from general matrix multi
 plication, matrices in SpGEMM are highly sparse and therefore require a co
 mpact representation. This places an additional burden...\n\n\nKai Shi (Be
 ijing University of Posts and Telecommunications); Zhe Lin (Sun Yat-sen Un
 iversity); and Xinya Luan, Jianwang Zhai, and Kang Zhao (Beijing Universit
 y of Posts and Telecommunications)\n---------------------\nRoutability-awa
 re Packing for High-density Nonvolatile FPGAs\n\nNonvolatile field-program
 mable gate arrays (NVFPGAs) can use multi-level cell (MLC) nonvolatile mem
 ories (NVMs) to enhance their logic density. However, the high-density des
 ign of NVFPGAs degrades the intra-routability of configurable logic blocks
  (CLBs), which significantly prolongs the time consum...\n\n\nHuichuan Zhe
 ng, Yuqing Xiong, Jian Zuo, Hao Zhang, Zhenge Jia, and Mengying Zhao (Shan
 dong University)\n\nTopics: Design\n\nTracks: DES1: SoC, Heterogeneous, an
 d Reconfigurable Architectures\n\nSession Chairs: Tianhao Cai (Beihang Uni
 versity) and Dirk Stroobandt (Ghent University)
END:VEVENT
END:VCALENDAR
