The 20th USENIX Symposium on Operating Systems Design and Implementation met at The Westin Seattle on July 13–15, 2026, co-located as usual with USENIX ATC. Twenty editions is a reasonable moment to ask what the venue is now for, because the answer has shifted noticeably — and the shift tells cloud infrastructure teams something about where the next several years of platform work is heading.
A note on scope: OSDI’s full proceedings are the authoritative record, and this is not a comprehensive paper roundup. What follows draws on papers whose acceptance and authorship can be confirmed independently, chosen because they mark the boundaries of the field’s current attention rather than because they are the only work worth reading.
Two Tracks, Two Different Claims
The structural fact worth noticing first is that OSDI now sorts papers into a Research track and an Operational Systems track. That division is not cosmetic. A research paper argues that a technique is sound and novel; an operational systems paper argues that something was built, deployed, and survived contact with real workloads. Systems venues spent years struggling to evaluate the second kind of contribution against the criteria of the first, and the split is an admission that “we ran this in production and here is what broke” is a distinct form of knowledge.
For anyone running infrastructure rather than publishing about it, the Operational Systems track is usually the more useful read. It is where the failure modes appear.
LLM Serving Has Become the Gravitational Center
Of the confirmed OSDI ‘26 papers, two of three concern large language model systems, which is a fair proxy for where the field’s energy has gone.
OpenTela: Unifying Decentralized HPC Clusters for Heterogeneous LLM Serving — from Xiaozhe Yao, Imanol Schlag and Ana Klimovic at ETH Zurich, with Youhe Jiang and Eiko Yoneki at Cambridge, Ilia Badanin at EPFL, Qinghao Hu at MIT, and Binhang Yuan at HKUST — appears in the Operational Systems track, which is the more interesting placement. The problem it takes on is one most enterprises will recognize even if they would not phrase it in HPC terms: inference capacity that is real but fragmented. An organization accumulates GPUs across clusters purchased at different times, with different interconnects and different generations of accelerator, under different administrative control. Each pool individually is too small or too heterogeneous to serve a large model well. Treating them as one substrate is a scheduling and placement problem rather than a hardware problem.
This is the same structural issue that has been reshaping GPU scheduling inside Kubernetes, approached from the HPC side. The convergence is notable: two research communities that historically shared little vocabulary are now solving the same allocation problem, because the economics of accelerator scarcity apply identically to both.
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration attacks the cost of inference-time reasoning. Tree-of-thought methods improve model output by exploring multiple reasoning paths and scoring them, which means the compute cost scales with the breadth of exploration rather than with output length. Applying speculative execution to that exploration — committing to promising branches before their evaluation completes — is the same class of optimization that speculative decoding brought to token generation, moved up a level of abstraction.
The infrastructure implication is the part worth holding onto. Reasoning models shifted a large share of AI compute from training to inference, and inference is the part that runs continuously for as long as a product exists. Work that reduces the cost of a reasoning step changes serving capacity planning directly, in a way that training-side improvements do not. This is the mechanism by which inference economics keeps propagating into datacenter decisions.
The Hardware Boundary Is Still Moving
Against that concentration, RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICs — Maximilian Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu, Zhiyi Hu, Jonas Dann and Gustavo Alonso, all at ETH Zurich — is a reminder that the substrate underneath is not settled either. It appears in the Research track.
RDMA over Converged Ethernet lets one machine read another’s memory without involving the remote CPU, and it is foundational to both distributed training and disaggregated storage. The recurring difficulty is that production RDMA deployments need services the basic protocol does not provide — congestion handling, tenant isolation, encryption, telemetry — and adding them in software reintroduces the CPU involvement RDMA existed to remove. Pushing those services into the SmartNIC’s offload engine is an attempt to keep the bypass while getting the services.
That matters for cloud infrastructure specifically because RDMA’s awkward property has always been that it assumes a well-behaved network and a trusted tenant. Neither holds in multi-tenant cloud. Work that makes RDMA safe to expose across tenant boundaries expands what a cloud provider can offer as a primitive rather than reserve for internal use — and Alonso’s group has been a consistent source of the hardware-software boundary work that eventually shows up in commercial offerings.
What to Take From It
The reading for infrastructure planners is that systems research has largely stopped treating AI workloads as a special case to be accommodated and started treating them as the default workload to be designed for. Scheduling, memory management, network offload, and resource federation are all being reconsidered with distributed training and inference serving as the assumed use case rather than an exception.
The practical consequence is about defaults rather than features. Platform primitives whose behavior was tuned for stateless request-response workloads — scheduler heuristics, network QoS policies, storage tiering — are being redesigned around workloads with radically different characteristics. Teams whose clusters mix both should expect the tuning advice they inherited to age faster than usual over the next several releases, and should read the Operational Systems track when deciding which of that advice to keep.
