Adaptive Resource Management for Efficient, Resilient, and Sustainable Systems
My research is located at the intersection of distributed, operating, and database systems, where it aims to make data-intensive applications in heterogeneous and distributed computing environments more efficient, resilient, and sustainable. I primarily work on adaptive resource management and carbon-aware computing methods for data analytics, stream processing, scientific workflows, and AI/ML applications running on cloud and edge infrastructure. At the University of Glasgow, I lead the Carbon-Conscious Computing Lab and take an empirical and iterative systems research approach.
Distributed Processing Systems Run on Increasingly Diverse Computing Infrastructures
Many organizations increasingly rely on data-driven and AI-enabled applications, for example to recommend content to millions of users, detect fraud in business transactions, identify genetic disorders by comparing genomic data, and monitor environmental conditions using sensor networks. For this, businesses, research institutions, municipalities, and other organizations employ scalable, fault-tolerant distributed systems and computing infrastructures.
The underlying computing infrastructures are becoming increasingly diverse, distributed, and dynamic, spanning cloud data centers, edge servers, and resource-constrained IoT devices. This allows applications to run closer to data sources and users, enabling lower latencies, improved security and privacy, and reduced energy consumption for wide-area networking, but it also creates challenging computing environments. At the same time, data centers are becoming more heterogeneous, offering an increasingly large variety of machine architectures and generations.
Running Data-Intensive Applications Efficiently Is a Key Challenge
Efficiently managing today's diverse and dynamic distributed computational resources so that data-intensive systems provide the required performance and dependability remains a major challenge. Even for expert users, configuring distributed systems so they meet application requirements is often difficult: system behavior depends on many factors, configuration spaces are large, and both execution environments and workloads change dynamically over time.
As a result, data-intensive distributed applications deployed in practice commonly suffer from low resource utilization, severe failures, and limited energy efficiency. Meanwhile, computing's already substantial environmental footprint is projected to continue to rise sharply. So, as an increasing number of businesses, scientific organizations, municipalities, and government bodies develop and deploy data- and AI-driven applications, it is critical, both economically and environmentally, that computing infrastructures are used efficiently.
Methods for Adaptively Managing Applications and Resources
The main objective of my work is to enable the efficient and dependable use of heterogeneous, distributed computing infrastructures for data-intensive and AI applications. To this end, I develop methods that make the design and operation of resource-efficient and resilient systems easier, with a particular focus on adaptive resource management across cloud and edge environments. Ultimately, we aim to realize systems that continuously adapt to diverse computing infrastructures, dynamically changing workloads, failures, and the fluctuating availability of low-carbon energy.
Central to this approach are techniques for modeling and optimizing the performance, dependability, and efficiency of distributed systems. Specifically, we combine lightweight profiling and experimentation methods with predictive performance models and explicitly address uncertainty both at planning time and during execution.
In line with this, we pursue the following three closely connected strategies:
- Adaptive Resource Allocation: We envision systems that automatically select adequate resources (e.g. types of resources, scale-outs, usage of accelerators, communication channels) and adapt virtual infrastructures as needed to meet specific performance objectives and constraints.
- Dynamic Scaling & Scheduling: We envision systems that automatically adjust resource configurations at runtime in response to changes in computing environments, using monitoring data, load prediction, detected failures, and observed behavior in similar situations.
- Automatic System Tuning: We envision systems that automatically tune key system parameters (e.g. cache and buffer sizes, memory allocations, CPU/GPU shares, DVFS, failure tolerance strategies, and snapshotting frequencies) on the basis of previous executions of similar applications, dedicated profiling runs, and predictive performance models.
By grounding these decisions in predictive models of observed system behavior, a wide spectrum of data-intensive edge and cloud workloads — including data analytics, scientific workflows, machine learning, and stream processing — can adapt continuously to heterogeneous and dynamic environments, such as variable compute capacity, failures, and low-carbon energy availability. Adaptive resource management thereby provides the means for making a range of widely relevant systems more efficient, resilient, and sustainable.
Research Methodology
Systems research: I primarily focus on empirical systems research. Therefore, together with doctoral/postdoc researchers, collaborators, and partners, I evaluate new ideas by implementing prototypes in the context of relevant open-source systems (such as Kubernetes, Flink, Spark, Nextflow, Flower, and FreeRTOS) and by conducting experiments on actual hardware, with authentic applications, and using real-world data. For this, we use state-of-the-art infrastructures, including diverse clusters, private and public cloud resources, and IoT devices and sensors.
Iterative approach: I am convinced that iterative processes and short feedback cycles are vital for research, so we implement a multi-stage approach: We first present new ideas in workshops or work-in-progress tracks, then submit rigorously evaluated work to renowned international conferences, and finally publish more extensive findings in reputable journals. I also believe it is essential to contribute to interdisciplinary projects with industry and public sector partners. Such joint projects offer important opportunities for knowledge transfer and innovation and provide first-hand exposure to real problems, often leading to well-motivated new ideas for foundational research.
Research teams and training: I place great value on building a close-knit, cohesive research team with a shared, well-defined agenda, so that there are ample opportunities to involve different perspectives, learn from one another, receive feedback, and iterate. To support this, I facilitate both structured and informal opportunities for exchange (e.g. lab lunches, reading groups, and away days), and I ensure that new members are well integrated into both my team and the wider institutional environment. Furthermore, I take the time to provide detailed feedback, and I organize research seminars and opportunities for doctoral researchers to collaborate with other groups.
Open exchange: Believing in discourse and feedback, I regularly collaborate with researchers from other institutions and with industry, co-organize conferences and workshops, and review research papers and proposals. Moreover, we make our results available as widely and quickly as possible through open-access publications, open-source software, and research-led teaching.