125 papers
cs dc
0/02026
9- JulHYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Cachingno summary yetcs-dc2607.01299Princeton0 citesJul 1, 2026
- JulMixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Modelsno summary yetcs-dc2607.01844Salesforce0 citesJul 2, 2026
- JulBounded-Memory Parallel Image Pulling for Large Container Imagesno summary yetcs-dc2607.05596Amazon0 citesJul 6, 2026
- JulDecomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUsno summary yetcs-dc2607.11368UW0 citesJul 13, 2026
- JunAdaptive Resource Management and Quality Control for Streaming Video Generationno summary yetcs-dc2606.15319Princeton0 citesJun 13, 2026
- JunCacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agentsno summary yetcs-dc2606.16824UW0 citesJun 15, 2026
- JunSwarmX: Agentic Scheduling for Low-Latency Agentic Systemsno summary yetcs-dc2606.21401Tencent0 citesJun 19, 2026
- JunConcordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inferenceno summary yetcs-dc2606.23521UW0 citesJun 22, 2026
- JanPractical One-Round-Trip BFT Replicationno summary yetcs-dc2601.03390NYU0 citesJan 6, 2026
2024
8- NovHPCAdvisor: A Tool for Assisting Users in Selecting HPC Resources in the Cloudno summary yetcs-dc2411.15448Microsoft Research0 citesNov 23, 2024
- OctWorkflows Community Summit 2024: Future Trends and Challenges in Scientific Workflowsno summary yetcs-dc2410.14943NVIDIA1 citesOct 19, 2024
- SepImitater: An Efficient Shared Mempool Protocol with Application to Byzantine Fault Toleranceno summary yetcs-dc2409.19286Baidu0 citesSep 28, 2024
- MayTowards Building Autonomous Data Services on Azureno summary yetcs-dc2405.01813Microsoft Research8 citesMay 3, 2024
- MayLarge-Scale Metric Computation in Online Controlled Experiment Platformno summary yetcs-dc2405.08411Tencent1 citesMay 14, 2024
- MarPolylog-Competitive Deterministic Local Routing and Schedulingno summary yetcs-dc2403.07410Google Research0 citesMar 12, 2024
- FebLogical Synchrony Networks: A formal model for deterministic distributionno summary yetcs-dc2402.07433Google Research24 citesFeb 12, 2024
- JanFast Kronecker Matrix-Matrix Multiplication on GPUsno summary yetcs-dc2401.10187Microsoft Research1 citesJan 18, 2024
2023
8- OctDxPU: Large Scale Disaggregated GPU Pools in the Datacenterno summary yetcs-dc2310.04648Alibaba7 citesOct 7, 2023
- OctSharkGraph: A Time Series Distributed Graph Systemno summary yetcs-dc2310.15762Tencent0 citesOct 24, 2023
- SepTowards General and Efficient Online Tuning for Sparkno summary yetcs-dc2309.01901Tencent19 citesSep 5, 2023
- SepOobleck: Resilient Distributed Training of Large Models Using Pipeline Templatesno summary yetcs-dc2309.08125Amazon27 citesSep 15, 2023
- MayAccelerating MPI Collectives with Process-in-Process-based Multi-object Techniquesno summary yetcs-dc2305.10612Meta / FAIR2 citesMay 17, 2023
- AprRuntime Variation in Big Data Analyticsno summary yetcs-dc2304.03424Microsoft Research5 citesApr 7, 2023
- JanHector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architecturesno summary yetcs-dc2301.06284NVIDIA1 citesJan 16, 2023
- JanDurable Algorithms for Writable LL/SC and CAS with Dynamic Joiningno summary yetcs-dc2302.00135Google Research3 citesJan 31, 2023
2022
3- SepApplication Experiences on a GPU-Accelerated Arm-based HPC Testbedno summary yetcs-dc2209.09731NVIDIA2 citesSep 20, 2022
- SepOptimizing DNN Compilation for Distributed Training with Joint OP and Tensor Fusionno summary yetcs-dc2209.12769Alibaba6 citesSep 26, 2022
- JanA Compiler Framework for Optimizing Dynamic Parallelism on GPUsno summary yetcs-dc2201.02789NVIDIA0 citesJan 8, 2022
2021
17- DecEfficient and Local Parallel Random Walksno summary yetcs-dc2112.00655Google Research0 citesDec 1, 2021
- DecMemory-efficient array redistribution through portable collective communicationno summary yetcs-dc2112.01075Google Research1 citesDec 2, 2021
- NovCloudRCA: A Root Cause Analysis Framework for Cloud Computing Platformsno summary yetcs-dc2111.03753Alibaba1 citesNov 5, 2021
- NovDoing More by Doing Less: How Structured Partial Backpropagation Improves Deep Learning Clustersno summary yetcs-dc2111.10672Amazon0 citesNov 20, 2021
- OctA Unified and Refined Convergence Analysis for Non-Convex Decentralized Learningno summary yetcs-dc2110.09993Alibaba44 citesOct 19, 2021
- SepRabia: Simplifying State-Machine Replication Through Randomizationno summary yetcs-dc2109.12616Google Research2 citesSep 26, 2021
- AugThe Case for Task Sampling based Learning for Cluster Job Schedulingno summary yetcs-dc2108.10464Google Research4 citesAug 24, 2021
- AugCompiler-Driven FPGA Virtualization with SYNERGYno summary yetcs-dc2109.02484Amazon22 citesAug 28, 2021
- MayDeployment Archetypes for Cloud Applicationsno summary yetcs-dc2105.00560Google Research0 citesMay 2, 2021
- MayDistributed In-memory Data Management for Workflow Executionsno summary yetcs-dc2105.04720Snap4 citesMay 11, 2021
- MayTowards Demystifying Serverless Machine Learning Trainingno summary yetcs-dc2105.07806Microsoft Research102 citesMay 17, 2021
- MayDorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threadsno summary yetcs-dc2105.11118Google Research19 citesMay 24, 2021
- AprFaa$T: A Transparent Auto-Scaling Cache for Serverless Applicationsno summary yetcs-dc2104.13869Microsoft Research11 citesApr 28, 2021
- AprFrom Distributed Machine Learning to Federated Learning: A Surveyno summary yetcs-dc2104.14362Baidu16 citesApr 29, 2021
- MarNear-zero Downtime Recovery from Transient-error-induced Crashesno summary yetcs-dc2103.05185Amazon1 citesMar 9, 2021
- MarPower Modeling for Effective Datacenter Planning and Compute Managementno summary yetcs-dc2103.13308Google Research2 citesMar 22, 2021
- FebLayer-based Composite Reputation Bootstrappingno summary yetcs-dc2102.09951Alibaba2 citesFeb 1, 2021
2020
31- DecWISE: A Computer System Performance Index Scoring Frameworkno summary yetcs-dc2012.07984Amazon0 citesDec 14, 2020
- DecCompilation Techniques for Graph Algorithms on GPUsno summary yetcs-dc2012.07990Adobe0 citesDec 14, 2020
- DecTEMPI: An Interposed MPI Library with a Canonical Representation of CUDA-aware Datatypesno summary yetcs-dc2012.14363NVIDIA1 citesDec 28, 2020
- Nov10 Years Later: Cloud Computing is Closing the Performance Gapno summary yetcs-dc2011.00656Google Research14 citesNov 2, 2020
- NovDeepTriage: Automated Transfer Assistance for Incidents in Cloud Servicesno summary yetcs-dc2012.03665Microsoft Research10 citesNov 25, 2020
- OctMove Fast and Meet Deadlines: Fine-grained Real-time Stream Processing with Cameono summary yetcs-dc2010.03035Microsoft Research1 citesOct 6, 2020
- OctTurboTransformers: An Efficient GPU Serving System For Transformer Modelsno summary yetcs-dc2010.05680Tencent9 citesOct 9, 2020
- Octmdspan in C++: A Case Study in the Integration of Performance Portable Features into International Language Standardsno summary yetcs-dc2010.06474NVIDIA13 citesOct 13, 2020
- OctIMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEadsno summary yetcs-dc2010.06574NVIDIA11 citesOct 13, 2020
- OctScheduling Opportunistic Links in Two-Tiered Reconfigurable Datacentersno summary yetcs-dc2010.07920Microsoft Research4 citesOct 15, 2020
- SepA Virtual Frame Buffer Abstraction for Parallel Rendering of Large Tiled Display Wallsno summary yetcs-dc2009.03368NVIDIA2 citesSep 7, 2020
- SepTime-Based Roofline for Deep Learning Performance Analysisno summary yetcs-dc2009.04598NVIDIA0 citesSep 9, 2020
- SepApplying the Roofline model for Deep Learning performance optimizationsno summary yetcs-dc2009.11224Baidu1 citesSep 23, 2020
- SepDistributed Many-to-Many Protein Sequence Alignment using Sparse Matricesno summary yetcs-dc2009.14467Microsoft Research1 citesSep 30, 2020
- JulAnalytics of Longitudinal System Monitoring Data for Performance Predictionno summary yetcs-dc2007.03451Google Research2 citesJul 7, 2020
- JulSelf-healing Dilemmas in Distributed Systems: Fault Correction vs. Fault Toleranceno summary yetcs-dc2007.05261Google Research1 citesJul 10, 2020
- JulFast Distributed Bandits for Online Recommendation Systemsno summary yetcs-dc2007.08061Adobe0 citesJul 16, 2020
- JunMaking Convolutions Resilient via Algorithm-Based Error Detection Techniquesno summary yetcs-dc2006.04984NVIDIA7 citesJun 8, 2020
- JunJAMPI: efficient matrix multiplication in Spark using Barrier Execution Modeno summary yetcs-dc2007.01811Google Research0 citesJun 27, 2020
- MayCommunication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networksno summary yetcs-dc2005.02426Tencent6 citesMay 5, 2020
- MayLearning to Accelerate Heuristic Searching for Large-Scale Maximum Weighted b-Matching Problems in Online Advertisingno summary yetcs-dc2005.04355Alibaba0 citesMay 9, 2020
- MayHeterogeneous CPU/GPU co-execution of CFD simulations on the POWER9 architecture: Application to airplane aerodynamicsno summary yetcs-dc2005.05899NVIDIA35 citesMay 12, 2020
- MayHyperLogLog Sketch Acceleration on FPGAno summary yetcs-dc2005.13332Microsoft Research4 citesMay 24, 2020
- MarDistributed Hierarchical GPU Parameter Server for Massive Scale Deep Learning Ads Systemsno summary yetcs-dc2003.05622Baidu10 citesMar 12, 2020
- MarScaling Strongly Consistent Replicationno summary yetcs-dc2003.07760Microsoft Research4 citesMar 17, 2020
- MarContainerStress: Autonomous Cloud-Node Scoping Framework for Big-Data ML Use Casesno summary yetcs-dc2003.08011NVIDIA0 citesMar 18, 2020
- MarA Hybrid MPI+Threads Approach to Particle Group Finding Using Union-Findno summary yetcs-dc2003.11468Google Research0 citesMar 25, 2020
- MarAiiDA 1.0, a scalable computational infrastructure for automated reproducible workflows and data provenanceno summary yetcs-dc2003.12476Microsoft Research283 citesMar 24, 2020
- FebSensitivity Analysis in the Dupire Local Volatility Model with Tensorflowno summary yetcs-dc2002.02481Google Research0 citesFeb 6, 2020
- FebReliable Distributed Clustering with Redundant Data Assignmentno summary yetcs-dc2002.08892Google Research1 citesFeb 20, 2020
- JanHigh Performance I/O For Large Scale Deep Learningno summary yetcs-dc2001.01858NVIDIA0 citesJan 7, 2020
2019
19- NovFast Dimensional Analysis for Root Cause Investigation in a Large-Scale Service Environmentno summary yetcs-dc1911.01225Meta / FAIR31 citesNov 1, 2019
- NovFederated Learning with Autotuned Communication-Efficient Secure Aggregationno summary yetcs-dc1912.00131Google Research15 citesNov 30, 2019
- OctPerformance Impact of Memory Channels on Sparse and Irregular Algorithmsno summary yetcs-dc1910.03679NVIDIA0 citesOct 8, 2019
- OctHyper: Distributed Cloud Processing for Large-Scale Deep Learning Tasksno summary yetcs-dc1910.07172Meta / FAIR1 citesOct 16, 2019
- OctSNF: Serverless Network Functionsno summary yetcs-dc1910.07700Google Research2 citesOct 17, 2019
- OctAutomatically Batching Control-Intensive Programs for Modern Acceleratorsno summary yetcs-dc1910.11141Google Research2 citesOct 23, 2019
- JulDeepPlace: Learning to Place Applications in Multi-Tenant Clustersno summary yetcs-dc1907.12916Adobe4 citesJul 30, 2019
- JunDistributed Weighted Matching via Randomized Composable Coresetsno summary yetcs-dc1906.01993Google Research2 citesJun 5, 2019
- JunA Performance Study of the 2D Ising Model on GPUsno summary yetcs-dc1906.06297NVIDIA33 citesJun 14, 2019
- JunNear Optimal Coflow Scheduling in Networksno summary yetcs-dc1906.06851Google Research22 citesJun 17, 2019
- JunA Static Analysis-based Cross-Architecture Performance Prediction Using Machine Learningno summary yetcs-dc1906.07840Baidu4 citesJun 18, 2019
- JunMediaPipe: A Framework for Building Perception Pipelinesno summary yetcs-dc1906.08172Google Research219 citesJun 14, 2019
- MayEfficient Inter-Datacenter Bulk Transfers with Mixed Completion Time Objectivesno summary yetcs-dc1905.01749Microsoft Research1 citesMay 5, 2019
- MayMassively Parallel Computation via Remote Memory Accessno summary yetcs-dc1905.07533Google Research8 citesMay 18, 2019
- MayDynamic Algorithms for the Massively Parallel Computation Modelno summary yetcs-dc1905.09175Google Research0 citesMay 22, 2019
- MarAuto-Vectorizing TensorFlow Graphs: Jacobians, Auto-Batching And Beyondno summary yetcs-dc1903.04243Google Research6 citesMar 8, 2019
- MarTonY: An Orchestrator for Distributed Machine Learning Jobsno summary yetcs-dc1904.01631LinkedIn2 citesMar 24, 2019
- JanCommunication cost of consensus for nodes with limited memoryno summary yetcs-dc1901.01665Microsoft Research2 citesJan 7, 2019
- JanAnalysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloadsno summary yetcs-dc1901.05758Microsoft Research67 citesJan 17, 2019
2018
7- OctsignSGD with Majority Vote is Communication Efficient And Fault Tolerantno summary yetcs-dc1810.05291Amazon25 citesOct 11, 2018
- SepInterstellar: Using Halide's Scheduling Language to Analyze DNN Acceleratorsno summary yetcs-dc1809.04070Google Research56 citesSep 10, 2018
- SepHDArray: Parallel Array Interface for Distributed Heterogeneous Devicesno summary yetcs-dc1809.05657Apple0 citesSep 15, 2018
- MayDynamic Control Flow in Large-Scale Machine Learningno summary yetcs-dc1805.01772DeepMind68 citesMay 4, 2018
- MayFork and Join Queueing Networks with Heavy Tails: Scaling Dimension and Throughput Limitno summary yetcs-dc1805.05197Amazon0 citesMay 14, 2018
- AprBigDL: A Distributed Deep Learning Framework for Big Datano summary yetcs-dc1804.05839Tencent138 citesApr 16, 2018
- MarCuLDA_CGS: Solving Large-scale LDA Problems on GPUsno summary yetcs-dc1803.04631Alibaba3 citesMar 13, 2018
2017
3- SepAdaptive Processing of Spatial-Keyword Data Over a Distributed Streaming Clusterno summary yetcs-dc1709.02533Google Research1 citesSep 8, 2017
- AugTensorFlow Estimators: Managing Simplicity vs. Flexibility in High-Level Machine Learning Frameworksno summary yetcs-dc1708.02637Google Research16 citesAug 8, 2017
- MarWPaxos: Wide Area Network Flexible Consensusno summary yetcs-dc1703.08905Microsoft Research2 citesMar 27, 2017
2016
5- OctHybrid-DCA: A Double Asynchronous Approach for Stochastic Dual Coordinate Ascentno summary yetcs-dc1610.07184Tencent0 citesOct 23, 2016
- OctStatic Analysis Using the Cloudno summary yetcs-dc1610.08198Microsoft Research1 citesOct 26, 2016
- AugDistributed or Monolithic? A Computational Architecture Decision Frameworkno summary yetcs-dc1608.00944Meta / FAIR37 citesAug 2, 2016
- JunLeveraging energy storage to optimize data center electricity cost in emerging power marketsno summary yetcs-dc1606.01536Microsoft Research8 citesJun 5, 2016
- JunTensor Contractions with Extended BLAS Kernels on CPU and GPUno summary yetcs-dc1606.05696NVIDIA50 citesJun 17, 2016
2015
5- DecDistributed Balanced Partitioning via Linear Embeddingno summary yetcs-dc1512.02727Google Research2 citesDec 9, 2015
- JunDynamic Service Migration in Mobile Edge Computing Based on Markov Decision Processno summary yetcs-dc1506.05261Amazon0 citesJun 17, 2015
- MayCompressing Communication in Distributed Protocolsno summary yetcs-dc1506.00290Microsoft Research3 citesMay 31, 2015
- MarDynamic Service Placement for Mobile Micro-Clouds with Predicted Future Costsno summary yetcs-dc1503.02735Amazon221 citesMar 9, 2015
- JanHybrid Update/Invalidate Schemes for Cache Coherence Protocolsno summary yetcs-dc1502.00101Microsoft Research0 citesJan 31, 2015
2014
22013
2- NovBreathe before Speaking: Efficient Information Dissemination Despite Noisy, Limited and Anonymous Communicationno summary yetcs-dc1311.3425Microsoft Research23 citesNov 14, 2013
- FebBenchmarking Usability and Performance of Multicore Languagesno summary yetcs-dc1302.2837Google Research36 citesFeb 12, 2013