Research

My research spans distributed systems, networking, and agentic AI systems. I come at AI from the systems side: my interest is less in the models themselves than in what it takes to make AI systems practical, reliable, and scalable — and in how this new class of workload is reshaping the infrastructure beneath it. Cloud computing and programmable networks have long been part of that story.

At KAUST, I lead the SANDS Lab.

Research Directions

Systems research has always been about the gap between what a machine can do in principle and what it does under real workloads, at real scale, with real failures. With AI, that gap has moved: an AI system is not a model but a distributed system around one, with its own state, control loops, failure modes, and cost structure — and we do not yet have good ways to observe it, reason about it, or hold it to a guarantee.

Our work follows four threads. They are distinct lines of inquiry, but they lean on one another: progress on one usually exposes the next question in another.

Systems foundations for agentic AI

Multi-agent LLM systems plan, call tools, and act in real environments, yet we still largely evaluate them through demonstrations. We are building benchmarks, evaluation frameworks, and realistic testbeds that make their behavior reproducible and observable, measuring not only whether a task succeeded but how the system got there and where it went wrong. A longer-horizon question we find compelling: as a growing share of infrastructure is consumed by machines rather than people, what should AI-native network and service interfaces look like?

Systems-aware machine learning

ML algorithms are often designed as though the infrastructure executing them were uniform and free. It is neither. We co-design the two, developing training and inference mechanisms that account explicitly for communication costs, latency hierarchies, and hardware heterogeneity, building on our work in programmable network acceleration and resource-aware optimization. In shared clusters, the same perspective leads to workload-aware scheduling and orchestration that remains resilient under interference.

Decentralized and collaborative machine learning

Much valuable data cannot move. Federated, peer-to-peer, and distributed optimization let participants learn together while data stays local, but doing so at scale strains communication and trust at the same time. We work on communication-efficient protocols, compression and synchronization schemes that preserve convergence guarantees, and mechanisms for trustworthy collaboration when participants are heterogeneous, unreliable, or not fully trusted.

Programming the cloud: performance, predictability, and security

Everything above runs on cloud infrastructure whose behavior is frequently opaque to the people depending on it. We study programming models and abstractions that deliver predictable performance and strong isolation without giving up efficiency: resource visibility and management, fair and constrained scheduling, and isolation mechanisms for untrusted tenants. Without transparency and control at this layer, the properties the other three threads aim for do not survive contact with a shared, multi-tenant environment.

Taken together, these threads serve one goal: making AI and distributed systems practical, reliable, and scalable at the scales that matter. We favor work that leaves behind something reusable — an abstraction, an open benchmark, a reproducible tool — over results that hold only in the setting that produced them.

Research Group

I am privileged to be working with these very talented individuals:

Students

  • Achref Rebai
  • Boris Radovic (co-advised with Veljko Pejovic)
  • Jihao Xin
  • Kaihua Liang
  • Mohammed K. Aljahdali
  • Mouheb Ben Nasr
  • Norah Alballa
  • Saverio Pasqualoni
  • Shayan Ali Hassan
  • Tongzhou Gu
  • Yangzhixin Luo
  • Yixi Chen (co-advised with Suhaib Fahmy)

Postdocs

Research Staff

Alumni

Check out our alumni!


© 2012-2026. All rights reserved.