Ali R. Butt
Professor, Computer Science & ECE (by courtesy) · Virginia Tech
Systems, Storage & I/O for ML/DL, cloud, HPC.
Ali is a Professor of Computer Science (and ECE by courtesy) and Associate Department Head for Faculty Development in CS@VT. He is an ACM Distinguished Member. He received his Ph.D. degree in Electrical and Computer Engineering from Purdue University in 2006. He is a recipient of an NSF CAREER Award (2008), IBM Faculty Awards (2008, 2015), a VT College of Engineering (COE) Dean's award for “Outstanding New Assistant Professor” (2009), an IBM Shared University Research Award (2009), and NetApp Faculty Fellowships (2011, 2015). He was named a VT COE Faculty Fellow in 2013. Ali was an Academic Visitor at IBM Almaden Research Center (Summer 2012) and a Visiting Research Fellow at Queen's University of Belfast (Summer 2013). He has served as the Associate Editor for IEEE Transactions on Cloud Computing (2018–present), ACM Transactions on Storage (2016–present), IEEE Transactions on Parallel and Distributed Systems (2013–2016), Cluster Computing: The Journal of Networks, Software Tools and Applications (2013–present), and Sustainable Computing: Informatics and Systems (2010–2015). He is an alumni of the National Academy of Engineering's US Frontiers of Engineering (FOE) Symposium (2009), US–Japan FOE (2012), and National Academy of Science's AA Symposium on Sensor Science (2015). He was also an organizer for the US FOE in 2010. Ali's research interests are in: cloud and high-performance computing systems; systems support for machine and deep learning applications; file, I/O, and storage systems; distributed systems; and large-scale experimental computer systems. At Virginia Tech he leads the Distributed Systems & Storage Laboratory (DSSL) and directs the stack@cs Center for Computer Systems.
Office: 220 Gilbert Place, Suite 4108, Blacksburg, VA 24060
Selected recent publications
2026
Eliminate Branches by Melding IR Instructions
Proceedings of the European Conference on Object-Oriented Programming (ECOOP), Brussels, Belgium, 26 pages, July 2026. *equal contribution.
Proceedings of the ACM International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS), Ann Arbor, MI, 28 pages, June 2026 (AR: 15.2%).
Assessing the Impact of Code Changes on the Fault Localizability of Large Language Models
Proceedings of the 19th IEEE International Conference on Software Testing, Verification and Validation (ICST), Daejeon, Republic of Korea, 12 pages, May 2026 (AR: 26.9%).
2025
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
Proceedings of the 16th ACM Symposium on Cloud Computing (SoCC), Online, USA, 12 pages, November 2025 (AR: 20%).
Proceedings of the 16th ACM Symposium on Cloud Computing (SoCC), Online, USA, 12 pages, November 2025 (AR: 20%).
IP-FL: Incentive-driven Personalization in Federated Learning
Proceedings of the 39th IEEE International Parallel & Distributed Processing Symposium (IPDPS 2025), Milan, Italy, 11 pages, June 2025.
FLStore: Efficient Federated Learning Storage for Non-training Workloads
Proceedings of the 8th Annual Conference on Machine Learning and Systems (MLSys), Santa Clara, CA, 11 pages, May 2025.
2024
TreeCNN and NILMTK Unite: Illuminating Energy Efficiency in Real-World Scenarios
Proceedings of the IEEE International Conference on Big Data (BigData), Washington, D.C., USA, pages 10, December 2024.
FedCaSe: Enhancing Federated Learning with Heterogeneity-aware Caching and Scheduling
Proceedings of the ACM Symposium on Cloud Computing (SoCC), Redmond, WA, 17 pages, November 2024. (AR: 30%).
Memory Allocation Under Hardware Compression
Proceedings of the 57th IEEE/ACM International Symposium on Microarchitecture (MICRO), Austin, TX, 17 pages, November 2024. (AR: 23%).
ACM Transactions on Storage, 20(3): 18:1-18:35, June 2024. [ Link to paper ]
Application-Aware Memory Management for Containerized HPC Workflows
Proceedings of the 38th IEEE International Parallel & Distributed Processing Symposium (IPDPS), San Francisco, CA, 14 pages, May 2024.
FLOAT: Federated Learning Optimizations with Automated Tuning
Proceedings of the 19th ACM European Conference on Computer Systems (EuroSys), Athens, Greece, 19 pages, April 2024. (AR: 16%).
Tarazu: An Adaptive End-to-End I/O Load-Balancing Framework for Large-Scale Parallel File Systems
ACM Transactions on Storage, 20(2): 11:1-11:42, April 2024. [ Link to paper ]
2023
Towards Cost-Effective and Resource-Aware Aggregation at Edge for Federated Learning
Proceedings of the 2023 IEEE International Conference on Big Data (BigData), Sorrento, Italy, 10 pages, December 2023. (AR: 17.49%).
SHADE: Enable Fundamental Cacheability for Distributed Deep Learning Training
Proceedings of the 21st USENIX Conference on File and Storage Technologies (FAST), Santa Clara, CA, 17 pages, February 2023. (AR: 22.8%).
2022
Heterogeneity-Aware Adaptive Federated Learning Scheduling
Proceedings of the IEEE International Conference on Big Data (BigData), Osaka, Japan, 10 pages, December 2022. (AR: 19.2%).
Translation-optimized Memory Compression for Capacity
Proceedings of the 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), Chicago, IL, 20 pages, October 2022. (AR: 23.9%).
TIFF: Tokenized Incentive for Federated Learning
Proceedings of the IEEE International Conference on Cloud Computing (CLOUD), Barcelona, Spain, 10 pages, July 2022. (AR: 22.4%).
SCHEDTUNE: A Heterogeneity-aware GPU Scheduler for Deep Learning
Proceedings of the 22nd IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), Taormina (Messina), Italy, 10 pages, May 2022.
Tokenized Incentive for Federated Learning
Proceedings of the AAAI International Workshop on Trustable, Verifiable and Auditable Federated Learning (FL-AAAI-22) in conjunction with AAAI 2022, Vancouver, BC, Canada, 9 pages, March 2022.
2020
Large-Scale Analysis of Docker Images and Performance Implications for Container Storage Systems
IEEE Transactions on Parallel and Distributed Systems, pages 12, 2020. [ Link to paper ]
Understanding HPC Application I/O Behavior Using System Level Statistics
Proceedings of the 27th IEEE International Conference on High Performance Computing, Data, and Analytics (HiPC) 2020, Pune, India, pages 10, December 2020.
ACM Transactions on Mathematical Software, 46(4), pages 20, November 2020. [ Link to paper ]
DupHunter: Flexible High-Performance Deduplication for Docker Registries
Proceedings of the USENIX Annual Technical Conference (ATC), Boston, MA, pages 14, July 2020. (AR: 18.6%).
An Integrated Indexing and Search Service for Distributed File Systems
IEEE Transactions on Parallel and Distributed Systems, 31(10):2375-2391, October 2020. [ Link to paper ]
Customizable Scale-Out Key-Value Stores
IEEE Transactions on Parallel and Distributed Systems, 31(9): 2081-2096, September 2020. [ Link to paper ]
On the Use of Containers in High Performance Computing Environments
Proceedings of the IEEE International Conference on Cloud Computing (CLOUD), Beijing, China, pages 9, October 2020. (AR: 17%).
MARBLE: A Multi-GPU Aware Job Scheduler for Deep Learning on HPC Systems
Proceedings of the IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), Melbourne, Victoria, Australia, pages 10, May 2020.
Efficient Metadata Indexing for HPC Storage Systems
Proceedings of the 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), Melbourne, Australia, pages 10, May 2020.
2019
Large-Scale Analysis of the Docker Hub Dataset
Proceedings of the IEEE International Conference on Cluster Computing (Cluster), Albuquerque, NM, pages 11, September 2019.
FSMonitor: Scalable File System Monitoring for Arbitrary Storage Systems
Proceedings of the IEEE International Conference on Cluster Computing (Cluster), Albuquerque, NM, pages 11, September 2019.
A Quantitative Study of Deep Learning Training on Heterogeneous Supercomputers
Proceedings of the IEEE International Conference on Cluster Computing (Cluster), Albuquerque, NM, pages 12, September 2019.
MOANA: Modeling and Analyzing I/O Variability in Parallel System Experimental Design
IEEE Transactions on Parallel and Distributed Systems, 30(8):1843-1856, August 2019. [ Link to paper ]
iez: Resource Contention Aware Load Balancing for Large-Scale Parallel File Systems
Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS), Rio de Janeiro, Brazil, pages 11, May 2019.
An Analysis Workflow-Aware Storage System for Multi-Core Active Flash Arrays
IEEE Transactions on Parallel and Distributed Systems, 30(2):271--285, February 2019. [ Link to paper ]
2018
BESPOKV: Application Tailored Scale-Out Key-Value Stores
Proceedings of The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), Dallas, TX, pages 16, November 2018. (AR:23.6%)
A Heterogeneity-Aware Task Scheduler for Spark
Proceedings of the IEEE International Conference on Cluster Computing (Cluster), Belfast, UK, pages 11, September 2018.
Finding and Counting Tree-Like Subgraphs using MapReduce
IEEE Transactions on Multi-Scale Computing Systems, 4(3):217–230, July 2018.
Improving Docker Registry Design based on Production Workload Analysis
Proceedings of the 16th USENIX Conference on File and Storage Technologies (FAST'18), Oakland, CA, pages 14, Feb 2018. (AR:16.5%)
Chameleon: An Adaptive Wear Balancer for Flash Clusters
Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS), Vancouver, Canada, pages 10, May 2018. (AR: 24.5%)
2017
TagIt: A Fast and Efficient Scientific Data Discovery Service
Proceedings of the 2017 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis (SC'17), Denver, CO, pages 10, November 2017. (AR:18.7%).
2016
CHOPPER: Optimizing Data Partitioning for In-Memory Data Analytics Frameworks
Proceedings of the IEEE International Conference on Cluster Computing (Cluster), Taipei, Taiwan, pages 10, September 2016. (AR: 24%).
ClusterOn: Building Highly Configurable and Reusable Clustered Data Services using Simple Data Nodes
Proceedings of the 8th USENIX Workshop on Hot Topics in Storage and File Systems (HotStorage), Denver, CO, pages 5, June 2016. (AR: 36.9%).
2015
MOS: Workload-aware Elasticity for Cloud Object Stores
Proceedings of the 25th ACM Symposium on High-Performance Parallel and Distributed Computing (HPDC), Kyoto, Japan, pages 12, May 2016. (AR: 15.5%). An initial design, Taming the Cloud Object Storage with MOS, appeared in Proceedings of the 10th Parallel Data Storage Workshop (PDSW), Austin, Texas, pages 6, November 2015. (AR: 36%). A related poster was presented in 6th ACM SIGOPS Asia-Pacific Workshop on Systems (APSys 2015), Tokyo, Japan, July 2015.
2016
MEMTUNE: Dynamic Memory Management for In-memory Data Analytic Platforms
Proceedings of the IEEE International Parallel & Distributed Processing Symposium (IPDPS), Chicago, IL, pages 10, May 2016. (AR: 23%).
2015
AnalyzeThis: An Analysis Workflow-Aware Storage System
Proceedings of the 2015 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis (SC'15), Austin, TX, pages 12, Nov. 2015. (AR: 22.1%).
Anatomy of Cloud Monitoring and Metering: A case study and open problems
Proceedings of the 6th ACM SIGOPS Asia-Pacific Workshop on Systems (APSys 2015), Tokyo, Japan, pages 7, July 2015. (AR: 29.4%).
Pricing Games for Hybrid Object Stores in the Cloud: Provider vs. Tenant
Proceedings of the the 7th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud), Santa Clara, CA, pages 7, July 2015. (AR: 32.8%).
CAST: Tiering Storage for Data Analytics in the Cloud
Proceedings of the International ACM Symposium on High-Performance Distributed Computing (HPDC), Portland, Oregon, pages 12, June 2015. (AR: 16.4%).
Proceedings of the 29th IEEE International Parallel and Distributed Processing Symposium (IPDPS), Hyderabad, India, pages 10, May 2015. (AR: 21.8%).
Proceedings of the IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid), Shenzhen, Guangdong, China, pages 10, May 2015. (AR: 25.7%).
An In-Memory Object Caching Framework with Adaptive Load Balancing
Proceedings of the ACM European Conference on Computer Systems (EuroSys), Bordeaux, France, pages 16, April 2015. (AR: 20.8%).
Privacy-Preserving Scanning of Big Content for Sensitive Data Exposure with MapReduce
Proceedings of the Fifth ACM Conference on Data and Application Security and Privacy (CODASPY), San Antonio, TX, pages 12, March 2015. (AR: 21.3%).
2014
On the Use of Microservers in Supporting Hadoop Applications
Proceedings of the IEEE International Conference on Cluster Computing (Cluster), Madrid, Spain, pages 66--74, September 2014. (AR: 23.8%).
mrOnline: MapReduce Online Performance Tuning
Proceedings of the ACM Symposium on High-Performance Parallel and Distributed Computing (HPDC), Vancouver, Canada, pages 12, June 2014. (AR: 16.2%).
hatS: A Heterogeneity-Aware Tiered Storage for Hadoop
Proceedings of the 14th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID), Chicago, IL, pages 10, May 2014. (AR: 19.1%).
2013
On Timely Staging of HPC Job Input Data
IEEE Transactions on Parallel and Distributed Systems, 24(9):1841--1851, September 2013. [ Link to paper ]
On Reducing Energy Management Delays in Disks
Journal of Parallel and Distributed Computing, 73(6):823--835, June 2013. [ Link to paper ]
2012
Cooperative Storage-Level De-Duplication for I/O Reduction in Virtualized Data Centers
Proceedings of the IEEE International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems (MASCOTS), Washington D.C., August 2012.
On the Use of GPUs in Realizing Cost-Effective Distributed RAID
Proceedings of the IEEE International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems (MASCOTS), Washington D.C., August 2012.
Proceedings of the ACM Symposium on High-Performance Parallel and Distributed Computing (HPDC), Delft, The Netherlands. June 2012. (AR: 16.1%).
SAHAD: Subgraph Analysis in Massive Networks Using Hadoop
Proceedings of the 26th IEEE International Parallel and Distributed Processing Symposium (IPDPS), Shanghai, China. May 2012. (AR: 20.7%).
2011
Symphony: A Scheduler for Client-Server Applications on Coprocessor-based Heterogeneous Clusters
Proceedings of the IEEE International Conference on Cluster Computing (Cluster 2011), Austin, TX. September 2011.
Timely Result-Data Offloading for Improved HPC Center Scratch Provisioning and Serviceability
IEEE Transactions on Parallel and Distributed Systems, 22(8):1307--1322, August 2011. [ Link to paper ]
Towards Synthesizing Realistic Workload Traces for Studying the Hadoop Ecosystem
Proceedings of the 19th Annual Meeting of the IEEE International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems (MASCOTS), Singapore, July 2011.
CATCH: A Cloud-based Adaptive Data Transfer Service for HPC
Proceedings of the 25th IEEE International Parallel and Distributed Processing Symposium (IPDPS), Anchorage, AK. May 2011. (AR: 19.6%).
2010
Functional Partitioning to Optimize End-to-End Performance on Many-core Architectures
Proceedings of the ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis (SC10), New Orleans, LA. November 2010. (AR: 20%).
Misc
Teaching
-
CS 3214Intro. to Computer Systems
-
CS 4284Systems & Networking Capstone
-
CS 5204Operating Systems
-
CS 6204Special Topics
-
CS 3204Operating Systems
Recent professional activities
- TPC Co-Chair, 26th USENIX Conference on File and Storage Technologies (FAST)2028
- General Co-Chair, IEEE International Conference on Cluster Computing (Cluster)2026
- TPC Area Co-Chair for Data Analytics, Visualization, & Storage, The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC'26)2026
- Work-in-Progress/Posters Co-Chair, 22nd USENIX Conference on File and Storage Technologies (FAST)2024
- General Chair, 32nd International ACM Symposium on High-Performance Parallel and Distributed Computing (HPDC)2023
- Treasurer, 14th ACM International Conference on Future Energy Systems (e-Energy)2023
- Finance Chair, International Conference on Supercomputing (ICS)2022
- TPC Co-Chair, IEEE International Conference on Networking, Architecture, and Storage (NAS)2022
- Tutorial Co-Chair, IEEE International Conference on Big Data (IEEE BigData)2021
- TPC Area Co-Chair for Clouds and System Software, The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC'20)2020
- TPC Area Chair for Cloud Computing and Data Centers, 40th IEEE International Conference on Distributed Computing Systems (ICDCS)2020
- Steering Committee, HPDC2019–present
- TPC Co-Chair, 28th ACM International Symposium on High-Performance Parallel and Distributed Computing (HPDC)2019
- Organizer, NSF Visioning Workshop on Data Storage Research 20252018
- Associate Editor, ACM Transactions on Storage (TOS)2016–present
- Steering Committee, Joint International Workshop on Parallel Data Storage & Data Intensive Scalable Computing Systems (PDSW-DISCS)2016–present
- TPC member, USENIX FAST'26, HPDC'26, SoCC'25, IEEE BigData'25, …
- Panelist: NSF, DOE, NSERC, QNRF, LXR, …