A single stalled flow inside a lossless RDMA fabric can stall a whole GPU training job, not just the flow that caused it. That is the tradeoff Priority-Based Flow Control was built around: without PFC, RDMA drops packets under load and throughput collapses; with PFC enabled, one congested queue can pause every other flow sharing its priority class, including flows nowhere near the actual bottleneck. Head-of-line blocking of this kind has been the central unresolved cost of RDMA congestion control since PFC became standard practice, and it gets worse, not better, as cluster sizes and per-GPU bandwidth grow.

Barre, a scheme presented at USENIX ATC 2025 by researchers from Fudan University and ByteDance, is a useful case study in how that problem is actually being addressed in 2025 and 2026 production clusters — not through new hardware, but through a small set of changes to a decades-old control-loop pattern.

The Baseline Barre Is Competing Against

Most production RDMA deployments still run some variant of DCQCN, the Data Center Quantized Congestion Notification scheme Microsoft and Mellanox published at SIGCOMM 2015. DCQCN pairs Explicit Congestion Notification marking with a rate-based additive-increase, multiplicative-decrease (AIMD) control loop, and it has been the working baseline for RDMA fabrics for a decade. It is also known to need careful, deployment-specific parameter tuning, and its response to sudden large-scale congestion events — the “incast” pattern typical of collective communication in distributed training — is not aggressive enough to prevent queue buildup before PFC pauses trigger.

The obvious research response is to design a more sophisticated scheme. The more interesting response, and the one Barre takes, is to ask what happens if the existing AIMD loop is kept and three narrowly scoped mechanisms are added around it, using only capabilities that already exist in commodity programmable NICs.

Three Changes to an Old Control Loop

Barre runs on programmable SmartNICs — the published deployment uses NVIDIA BlueField-3 — and works from three event types those NICs already expose: transmission events, congestion notification packets, and round-trip-time measurements. On top of that, it adds three mechanisms.

Fast Increase replaces a fixed additive-increase step with one that scales to the flow’s own recent rate history across multiple RTTs — larger increments for flows still ramping up from a low rate, smaller ones for flows already near line rate. This is a standard idea in AIMD tuning, but applying it per-flow at RTT granularity is what makes it responsive enough for training-job traffic patterns that shift rapidly between bursts and idle.

Dual-lock is the more unusual piece. Conventional AIMD schemes trigger a rate increase when either a byte-counter threshold or a timer threshold is reached — an “OR” condition. Barre changes that to an “AND” condition: both thresholds must be satisfied before the rate increases. That single logic change stops the scheme from ramping rate back up too quickly immediately after a congestion event, which is exactly when a premature increase does the most damage to fairness between competing flows.

Inflight Monitor tracks the volume of unacknowledged, in-flight bytes within a rolling RTT window and preemptively cuts the sending rate if that volume crosses a threshold — acting before an explicit congestion signal arrives rather than waiting for one. This is the mechanism doing the most work against incast: a burst of many-to-one traffic, the pattern that dominates the all-to-all and all-reduce collectives in distributed training, can overwhelm a switch buffer faster than ECN marking and RTT feedback can respond. Inflight Monitor gives the scheme a way to react to the buildup itself, not just its downstream signals.

None of the three requires custom ASICs, kernel bypass beyond what RDMA already uses, or a redesign of the switch fabric. That is a deliberate design constraint, and it is the reason the scheme is deployable on infrastructure operators already own.

What a Year in Production Showed

The paper’s strongest evidence is not a simulation. Barre ran for more than a year on a production 400 Gbps RoCEv2 cluster scaling to roughly 10,000 GPUs, and across that deployment it improved average AI training task throughput by 9.6%. In large-scale incast scenarios, it cut peak switch queue lengths by as much as 79.9%. In NCCL AlltoAll benchmarks — the collective-communication pattern that stresses congestion control hardest because every node briefly becomes both sender and receiver to every other node — it outperformed DCQCN by a wide margin while avoiding PFC pause triggers that DCQCN configurations frequently hit under the same load. The paper also reports throughput comparable to InfiniBand, which is the more expensive alternative fabric operators reach for specifically to avoid RoCE’s congestion-management headaches.

A 9.6% average throughput improvement sounds modest next to headline numbers from newer hardware generations, but it is a fleet-wide, year-long production result on a cluster that was already tuned and already running, not a lab benchmark measuring best-case conditions. Gains of that kind, sustained at scale without new capital expenditure, are a different and arguably more credible category of result than a conference demo.

A Problem the Field Is Actively Re-Litigating

Barre is one entry in what is now a genuinely active research area. The scale of that activity is itself notable: China’s Journal of Computer Research and Development published a survey specifically on RDMA congestion-control techniques for datacenter networks in its 2026 volume, covering the field’s evolution from early congestion-awareness and rate-regulation algorithms through more recent reinforcement-learning-based approaches — evidence that after two decades, RDMA congestion control has accumulated enough distinct technique families to need a taxonomy, not just new point solutions. The pattern across this literature is consistent: PFC’s head-of-line-blocking cost has not gone away, and every new congestion-control proposal is, in one form or another, an attempt to make the lossless-network assumption cheaper to keep.

What to Watch Before Generalizing the Result

The caution is proportional to the evidence quality, not opposed to it. Barre’s numbers come from one operator’s cluster, one NIC generation, and one collective-communication workload profile; a 9.6% gain on a different traffic mix, a different oversubscription ratio, or a different generation of switch silicon is not guaranteed to reproduce. The Dual-lock and Inflight Monitor mechanisms are also, by the paper’s own framing, tuned around the specific failure mode of GPU-cluster incast — they are not obviously general-purpose improvements for RDMA traffic that looks more like traditional storage or database replication. For anyone evaluating congestion-control changes on their own fabric, the more durable takeaway from Barre is not the specific 9.6% figure but the design principle underneath it: before reaching for new hardware, check whether the existing control loop’s trigger logic and monitoring granularity are actually doing what the workload needs.