Presentation
INSTA: An Ultra-Fast, Differentiable, Statistical Static Timing Analysis Engine for Industrial Physical Design Applications*
DescriptionExisting GPU-accelerated Static Timing Analysis (GPU-STA) efforts aim to build standalone engines from scratch but result in poor correlation with commercial tools, limiting their industrial applicability.
In this paper, we present INSTA, a tool-accurate, differentiable, GPU-STA framework that overcomes these limitations by initializing timing graphs directly from reference STA tools (e.g., Synopsys PrimeTime).
INSTA's core engine utilizes two custom CUDA kernels: a forward kernel for statistical arrival time propagation, and a backward kernel for gradient backpropagation from timing endpoints, enabling two unprecedented capabilities: (1) high-fidelity, rapid timing analysis for incremental netlist updates (e.g., gate sizing), and (2) gradient-based, global timing optimization at scale (e.g., timing-driven placement).
Notably, INSTA demonstrates a near-perfect 0.999 correlation with PrimeTime on a 15-million-pin design in a commercial $3nm$ node with runtime under 0.1 seconds.
In the experiments, we showcase INSTA's power through three applications: (1) serving as a fast evaluator in a commercial gate sizing flow, achieving 25x faster incremental update_timing runtime with almost no accuracy loss; (2) INSTA-Size, a gradient-based gate sizer that achieves up to 15% better Total Negative Slack (TNS) than PrimeTime's default engine by sizing 68% fewer amount of cells; and (3) INSTA-Place, a differentiable timing-driven placer that outperforms the state-of-the-art net-weighting placer by up to 16% in Half-Perimeter Wirelegnth (HPWL) and 59.4\% in TNS.
We will open-source INSTA upon acceptance.
In this paper, we present INSTA, a tool-accurate, differentiable, GPU-STA framework that overcomes these limitations by initializing timing graphs directly from reference STA tools (e.g., Synopsys PrimeTime).
INSTA's core engine utilizes two custom CUDA kernels: a forward kernel for statistical arrival time propagation, and a backward kernel for gradient backpropagation from timing endpoints, enabling two unprecedented capabilities: (1) high-fidelity, rapid timing analysis for incremental netlist updates (e.g., gate sizing), and (2) gradient-based, global timing optimization at scale (e.g., timing-driven placement).
Notably, INSTA demonstrates a near-perfect 0.999 correlation with PrimeTime on a 15-million-pin design in a commercial $3nm$ node with runtime under 0.1 seconds.
In the experiments, we showcase INSTA's power through three applications: (1) serving as a fast evaluator in a commercial gate sizing flow, achieving 25x faster incremental update_timing runtime with almost no accuracy loss; (2) INSTA-Size, a gradient-based gate sizer that achieves up to 15% better Total Negative Slack (TNS) than PrimeTime's default engine by sizing 68% fewer amount of cells; and (3) INSTA-Place, a differentiable timing-driven placer that outperforms the state-of-the-art net-weighting placer by up to 16% in Half-Perimeter Wirelegnth (HPWL) and 59.4\% in TNS.
We will open-source INSTA upon acceptance.
Event Type
Research Manuscript
TimeMonday, June 2310:30am - 10:45am PDT
Location3004, Level 3
EDA3: Timing Analysis and Optimization


