33 papers
cs ar
0/02026
8- JulORRAM: An OpenROAD-Integrated RAM Generator Using Standard Cellsno summary yetcs-ar2607.12244NYU0 citesJul 14, 2026
- JunEPIC: A System Framework for Efficient Egocentric Perception on Embodied AR Glassesno summary yetcs-ar2606.15859NYU0 citesJun 14, 2026
- JunGoogle's Training Supercomputers from TPU v2 to Ironwood: Architectural Stability, Scale, Resilience, Power Efficiency, and Sustainability Across Five Generationsno summary yetcs-ar2606.15870Google Research0 citesJun 14, 2026
- JunArchitecture Carbon Tool v3: Enabling Sustainability-aware Silicon System Design Explorationno summary yetcs-ar2606.16889Meta / FAIR0 citesJun 15, 2026
- JunThe Kernel's Write: Application Read-Only Memoryno summary yetcs-ar2606.26138Stanford0 citesJun 18, 2026
- JunCHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design researchno summary yetcs-ar2606.27350Berkeley0 citesJun 25, 2026
- JunAgentic Hardware Design as Repository-Level Code Evolutionno summary yetcs-ar2606.28279NVIDIA0 citesJun 26, 2026
- MarWattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modelingno summary yetcs-ar2603.26435NVIDIA0 citesMar 27, 2026
2024
3- SepA Hardware-Aware Gate Cutting Framework for Practical Quantum Circuit Knittingno summary yetcs-ar2409.03870Tencent8 citesSep 5, 2024
- SepLoopTree: Exploring the Fused-layer Dataflow Accelerator Design Spaceno summary yetcs-ar2409.13625Google Research3 citesSep 20, 2024
- AprTao: Re-Thinking DL-based Microarchitecture Simulationno summary yetcs-ar2404.10921Google Research6 citesApr 16, 2024
2023
4- SepTailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Capacityno summary yetcs-ar2310.00192NVIDIA9 citesSep 29, 2023
- MayHighLight: Efficient and Flexible DNN Acceleration with Hierarchical Structured Sparsityno summary yetcs-ar2305.12718NVIDIA26 citesMay 22, 2023
- AprTeAAL: A Declarative Framework for Modeling Sparse Tensor Acceleratorsno summary yetcs-ar2304.07931NVIDIA0 citesApr 17, 2023
- AprRAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!no summary yetcs-ar2304.07935NVIDIA42 citesApr 17, 2023
2022
12021
6- DecN3H-Core: Neuron-designed Neural Network Accelerator via FPGA-based Heterogeneous Computing Coresno summary yetcs-ar2112.08193Alibaba17 citesDec 15, 2021
- OctSiliFuzz: Fuzzing CPUs by proxyno summary yetcs-ar2110.11519Google Research0 citesOct 5, 2021
- SepGoogle Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecksno summary yetcs-ar2109.14320Google Research2 citesSep 29, 2021
- MayRecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and Performanceno summary yetcs-ar2105.08820Meta / FAIR1 citesMay 18, 2021
- FebFuzzing Hardware Like Softwareno summary yetcs-ar2102.02308Google Research30 citesFeb 3, 2021
- JanRecSSD: Near Data Processing for Solid State Drive Based Recommendation Inferenceno summary yetcs-ar2102.00075Meta / FAIR4 citesJan 29, 2021
2020
4- DecDeepNVM++: Cross-Layer Modeling and Optimization Framework of Non-Volatile Memories for Deep Learningno summary yetcs-ar2012.04559Apple17 citesDec 8, 2020
- NovUnderstanding Training Efficiency of Deep Learning Recommendation Models at Scaleno summary yetcs-ar2011.05497Meta / FAIR8 citesNov 11, 2020
- JunEnabling Compute-Communication Overlap in Distributed Deep Learning Training Platformsno summary yetcs-ar2007.00156Meta / FAIR37 citesJun 30, 2020
- JanAchieving Multi-Port Memory Performance on Single-Port Memory with Coding Techniquesno summary yetcs-ar2001.09599Google Research0 citesJan 27, 2020
2019
2- JunAra: A 1 GHz+ Scalable and Energy-Efficient RISC-V Vector Processor with Multi-Precision Floating Point Support in 22 nm FD-SOIno summary yetcs-ar1906.00478Google Research7 citesJun 2, 2019
- JunMixed-Signal Charge-Domain Acceleration of Deep Neural networks through Interleaved Bit-Partitioned Arithmeticno summary yetcs-ar1906.11915Microsoft Research4 citesJun 27, 2019