When a platform engineer today debates whether a multi-cluster Kubernetes placement controller should schedule workloads by cost, by GPU availability, or by data-residency constraints, they are re-running an argument that distributed-systems researchers were already having twenty-five years ago — just with different vocabulary and a different substrate underneath. The vocabulary was “grid computing.” The substrate was a patchwork of university supercomputers, national labs, and research clusters, stitched together with middleware that had to solve resource discovery, job submission, and scheduling across administrative domains that trusted each other only partially. Most of that middleware is gone. The research questions it forced into the open are not — they resurface, largely unattributed, in the design documents of every modern container-orchestration platform.
This is a look at that lineage: where it started, what problem it was actually solving, why the grid-computing model gave way to the cloud, and which of its research questions Kubernetes and its multi-cluster extensions are now answering again.
The I-WAY Experiment and the Birth of Globus
The practical starting point for computational grid research is usually dated to the I-WAY experiment, demonstrated at the Supercomputing ‘95 conference in San Diego. Argonne National Laboratory computer scientists and networking engineers built a high-speed network linking supercomputing sites across the United States into a single shared resource for the duration of the conference, so that researchers could run applications that spanned multiple institutions’ hardware at once. The experiment worked, and it exposed a problem that had no existing solution: there was no common software layer for discovering what compute was available across institutions, authenticating a user across those institutions’ separate security domains, and submitting a job that could actually run on whichever site had free capacity.
The Globus Toolkit was built to be that layer. Developed with early DARPA funding following I-WAY and first released in 1997, Globus provided a security framework for cross-domain authentication, a resource-information service for discovering what was available, and job-submission and data-movement interfaces that middleware and applications could build on. It became the reference middleware for the era of grid computing that followed, deployed across research and national-lab computing sites worldwide through the 2000s.
Globus itself did not survive as an open-source project in that form. The Globus team ended maintenance and distribution of the open-source Globus Toolkit at the close of 2018, redirecting effort toward Globus.org — now a hosted, subscription-supported research-data-management and workflow service operated by the University of Chicago and Argonne, rather than a do-it-yourself middleware stack that institutions install and run themselves. The shift is itself a data point about the research question this article is tracing: middleware that requires every participating institution to install and operate the same complex stack has a much harder adoption path than a shared service model — a lesson the cloud era absorbed early and grid computing absorbed late.
DRMAA: A Standard for a Fragmented Scheduling Landscape
Globus solved cross-domain discovery and identity. It did not solve a narrower but equally stubborn problem: every site’s local batch scheduler — Sun Grid Engine, PBS, Condor, LSF, and others — exposed a different API for submitting and monitoring jobs. An application written to submit work to one site’s scheduler would not run unmodified against another’s. For a genuinely portable grid application, that was a serious obstacle.
The Distributed Resource Management Application API, standardized by the Open Grid Forum as GFD.22 and ratified to full recommendation status in 2007, existed to close that gap. DRMAA defined a common, language-bindable API for job submission, monitoring, and control, so that an application coded against DRMAA could run against Sun Grid Engine, Condor, or any other DRMAA-compliant scheduler without modification. It was an unglamorous piece of infrastructure — an abstraction layer over abstraction layers — but it is the direct conceptual ancestor of a design pattern that shows up constantly in cloud infrastructure today: a stable, provider-agnostic API sitting in front of heterogeneous, provider-specific backends.
GridWay and the Metascheduler Model
Sitting above both Globus and DRMAA-compliant local schedulers was a class of software called a metascheduler: a layer that took a job destined for “the grid” as a whole and decided, dynamically, which specific site and local scheduler should actually run it, based on current load, data locality, and policy. GridWay was one implementation of this model — a DRMAA-compliant metascheduler designed to sit on top of Globus-based grids and route jobs to whichever underlying site had capacity, without the application needing to know which site that would be.
Editorial note: GridWay originated from research work published by the Distributed Systems Architecture group formerly based at Universidad Complutense de Madrid — the group whose prior web presence occupied this domain before its 2026 acquisition. This article discusses GridWay as a documented piece of the grid-computing research record, citable through the published academic literature. It is not authored, reviewed, or endorsed by any current or former member of that group, and this publication is not affiliated with Universidad Complutense de Madrid, GridWay, or any successor project. See the site’s about page for the full non-affiliation disclosure.
The metascheduler model is worth naming precisely because it is the piece of grid-computing research that mapped most directly onto what came next. A metascheduler’s core problem — given a job and a set of heterogeneous, independently-administered resource pools, decide where to place it — is exactly the problem a modern multi-cluster Kubernetes placement controller solves, minus the cross-institutional trust boundary that made grid computing so much harder to operate in practice.
Why Grid Computing Gave Way to the Cloud
Grid computing’s core assumption was that valuable compute already existed, scattered across institutions with their own ownership and administration, and the job of the middleware was to federate access to it. Virtualization broke that assumption’s economic logic — a shift this site has traced elsewhere in the resurgence of open, vendor-neutral private-cloud platforms that inherited grid computing’s federation instincts without its cross-institutional trust burden. Once a cloud provider could offer effectively unlimited, uniformly-administered, on-demand compute from a single trust domain, the hard problems that grid middleware existed to solve — cross-domain identity federation, heterogeneous scheduler interoperability, discovery across independently-administered sites — mostly stopped mattering for the majority of workloads. It is far simpler to rent uniform compute from one provider than to federate access to non-uniform compute across many.
What grid computing got right, and what the cloud era spent roughly a decade re-learning, was that a stable API in front of a scheduling decision is more valuable than the scheduling decision itself. DRMAA’s core insight — write once against a common interface, let the interface route to heterogeneous backends — reappears nearly unchanged in the modern container-orchestration API model, just with Kubernetes’ declarative resource model as the common interface instead of a job-submission call.
Kubernetes as the New Metascheduler
Multi-cluster Kubernetes placement tooling is where the metascheduler pattern has most visibly resurfaced — a theme this site has also covered from the GPU-scheduling side in our commentary on Kubernetes as AI infrastructure. KubeFleet — the open-source project underlying Azure Kubernetes Fleet Manager, contributed to the CNCF Sandbox in 2025 — designates one cluster as a hub and propagates workloads to member clusters according to declarative placement policies that weigh cluster capacity, cost, GPU availability, and topology-spread constraints. CNCF’s KubeAdmiral project solves an adjacent version of the same problem: propagating and synchronizing resources across federated Kubernetes clusters from a single control plane. Functionally, both are metaschedulers: a layer that accepts a workload destined for “the fleet” as a whole and decides, dynamically, which member cluster should actually run it.
The differences from the grid era are instructive rather than incidental. A Kubernetes fleet’s member clusters are typically all administered by the same organization, or at minimum operate under a shared identity provider and API contract — the cross-institutional trust problem that consumed a large share of grid-computing research effort barely exists in this model. What has grown instead is the sophistication of the placement decision itself: GPU-aware bin-packing, topology-aware failure-domain spreading, and cost-based placement across cloud regions are all scheduling problems with more real-time signal and finer-grained resource typing than GridWay’s era had available.
What Carries Forward and What Doesn’t
Three research threads carry forward cleanly. First, the case for a stable placement API decoupled from the underlying scheduling implementation — DRMAA’s founding argument — is now uncontroversial cloud-infrastructure design practice. Second, the metascheduler pattern of dynamic, policy-driven placement across a pool of heterogeneous execution targets is precisely what KubeFleet and KubeAdmiral implement, under a different name and a much larger operational scale than GridWay ever ran at. Third, the underlying scheduling theory — bin-packing, fairness, preemption, gang-scheduling for tightly-coupled jobs — is continuous research territory that never stopped being relevant; grid-era schedulers and Kubernetes’ GPU-scheduling extensions are both instances of the same open scheduling problems, studied by an overlapping research community.
What does not carry forward is the trust model. Grid computing’s hardest unsolved problem — federating identity, policy, and resource access across institutions that only partially trust one another — has not been re-solved by Kubernetes; it has mostly been avoided, by consolidating fleets under single-organization administration. The multi-cloud and multi-tenant placement problems that come closest to the original grid-computing trust boundary — placing workloads across genuinely independent cloud providers with different security models — remain comparatively immature research territory relative to how well-developed intra-fleet, single-organization scheduling has become.
Open Research Questions Today
The gap that remains open is federation across genuinely independent administrative domains — the problem grid computing set out to solve and never fully closed before the cloud made it commercially unnecessary for most workloads. It resurfaces wherever compute must be placed across boundaries that are not under one organization’s control: cross-provider multicloud placement with differing security models, data-sovereignty-constrained scheduling where workloads must stay within a jurisdiction, and the identity-federation problem for AI agent systems that increasingly need to coordinate work across organizational boundaries the way research grids once did across university boundaries. The tooling has changed completely. The research question — how do you place work across resources you don’t uniformly control or trust — is the same one Globus and DRMAA were built to answer, and it is still not fully answered.
Frequently Asked Questions
What was the Globus Toolkit?
The Globus Toolkit was middleware for building computational grids, providing cross-institution security, resource discovery, and job-submission services. First released in 1997 following the I-WAY networking experiment, it became the reference software for grid computing through the 2000s. Its open-source distribution ended in 2018; Globus.org now operates as a hosted research-data service.
What is DRMAA and why did it matter?
DRMAA (Distributed Resource Management Application API) is an Open Grid Forum standard, ratified in 2007, defining a common API for submitting and monitoring jobs across different local batch schedulers. It let grid applications run against Sun Grid Engine, Condor, or other DRMAA-compliant systems without rewriting scheduler-specific code.
What was GridWay’s role in grid computing?
GridWay was a DRMAA-compliant metascheduler that ran on top of Globus-based grids, dynamically deciding which site should execute a submitted job based on load and availability. It exemplified the metascheduler pattern: a policy-driven placement layer sitting above heterogeneous, independently-scheduled resource pools.
How does Kubernetes multi-cluster scheduling relate to grid computing research?
Projects like KubeFleet and KubeAdmiral solve a structurally similar problem to grid metaschedulers: placing workloads across a pool of heterogeneous clusters based on policy, capacity, and cost. The key difference is trust — Kubernetes fleets are typically single-organization, sidestepping the cross-institutional identity federation problem that made grid computing’s version of this problem much harder.

