The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

Google AX and Agent Substrate explained: what the two layers actually do

Updated

Server racks in a data center corridor

Agent Substrate is the runtime that hosts AI agents cheaply and safely, AX is the Kubernetes-style control plane that manages them; this post explains each layer and how they fit together.

Developer posts on X spent days arguing about a new pair of names: Google AX and Agent Substrate. The excitement is easy to understand once you see the problem they target. Companies now run hundreds of AI agents for automation, and the old homes for software do not fit them. This post explains the two layers in plain terms: what each one does, how they connect, and what to check before you trust them with real work.

Why agents do not fit microservices or batch jobs

Classic cloud thinking offers two shapes. A microservice stays up all the time and answers requests. A batch job starts, runs to the end and exits. AI agents are a third shape entirely:

  • They hold state across long conversations and tool calls.
  • They write and run code by themselves, including commands nobody reviewed.
  • They wait a lot: for model responses and for human approvals, sometimes hours.
  • Left unwatched, they can loop and burn money with no result.

A microservice platform charges you while the agent waits. A batch platform has nowhere to keep the agent's memory between steps. So teams either pay for idle servers or build fragile glue code around both. That gap is the whole reason a purpose-built agent stack gets attention.

The lower layer: Agent Substrate, the runtime and data plane

Substrate is the ground the agents stand on. Three ideas carry most of its value.

Zero-idle snapshot and resume. When an agent waits for a model reply or a human approval, Substrate freezes its memory and state to disk, frees the CPU and RAM, and restores the agent when the next event arrives. Early coverage describes resume in under half a second and far higher agent density per server. Treat those figures as the project's reported claims, not independent measurements, but the mechanism itself is familiar: it is the same freeze and thaw idea behind serverless scale-to-zero, applied to agent state instead of HTTP handlers.

Hardware-grade sandboxing. Agents run generated code, so every run is partly untrusted. Substrate isolates each agent in micro virtual machines or gVisor-style sandboxes so a bad command stays inside its own box. If you have operated shared CI runners or multi-tenant notebooks, you already know this discipline: never let generated code touch the host directly.

A data plane, not a dashboard. Substrate moves bytes and schedules compute. It does not decide what the agent should do next. That separation matters because it lets the upper layer change policy without touching execution.

The upper layer: AX, the Kubernetes-style control plane

AX is where humans manage the fleet. It borrows the Kubernetes playbook: declare the desired state, let the controller converge reality toward it, inspect everything with short commands. Operators coming from kubectl will recognize verbs like apply, get, suspend, resume and ssh.

Four objects do nearly all the work:

Task. One isolated unit of agent work. Create it, watch it, pause it, resume it, open a shell into it when it misbehaves. If pods are the unit of Kubernetes, Tasks are the unit of AX.

Workspace. Everything the agent needs at boot: the Git repos checked out, the MCP servers connected, the tools installed. Declaring the workspace up front is what makes an agent run repeatable instead of a snowflake.

Gateway. The network gate. It lists which outside servers and APIs the agent may call and blocks the rest. This is the object security reviewers will read first, because an agent with open internet access and code execution is a breach waiting for a prompt.

Model. Which language model the agent calls and where the keys live. Central key management beats scattering API keys across agent configs, the same lesson secrets management taught a decade ago.

How the two layers connect

A typical flow reads like this. You declare a Task with a Workspace, a Gateway policy and a Model reference through the AX CLI. AX admits the Task and hands it to Substrate. Substrate starts the sandbox, wires the workspace, and runs the agent. When the agent waits, Substrate snapshots it and reclaims the machine. When the model replies or the human approves, Substrate restores the snapshot and the agent continues. AX keeps showing you one steady Task the whole time, while underneath the agent may have been frozen and thawed many times.

Worked example: a fictional night-shift triage agent

The context below is fictional. Fictional shop BrightCart (fictional) runs one agent that reads the error queue every night, drafts incident notes and waits for an on-call engineer to approve each note before filing it.

On classic servers the agent holds a full container for eight hours to do ninety minutes of real work. On a Substrate-style runtime the agent runs while drafting, freezes while waiting for approvals, and resumes when the engineer clicks approve. AX shows the operator one Task with a Gateway that allows only the ticket API and the docs server, and a Workspace with the runbook repo already checked out. The bill follows active compute instead of wall-clock hours, and the blast radius follows the Gateway policy instead of the agent's imagination.

Checklist: evaluate the pair honestly

  1. Reproduce the snapshot claim yourself: freeze an agent mid-conversation, resume it, and confirm no state was lost.
  2. Read the Gateway policy like an attacker: list every host the agent can reach and ask what each one allows.
  3. Test the CLI against your real workflow: can you suspend, inspect and resume a stuck agent faster than today?
  4. Measure one workload end to end before believing any density or cost claim.
  5. Keep one workload on the old stack as a control group for a month.

Straight answers

Frequently asked questions

What is Google Agent Substrate?

It is the lower layer: the runtime and data plane that AI agents run on, handling snapshot resume, sandboxing and resource packing, shared as open source according to early coverage.

What is Google AX?

It is the upper layer: a declarative, Kubernetes-inspired control plane with a kubectl-style CLI for managing agent workloads through Task, Workspace, Gateway and Model objects.

Is Agent Substrate the same as Kubernetes?

No. Kubernetes schedules containers that are usually always on or run to completion. Substrate schedules agents that hold state, pause for long waits and generate their own code, so it needs freeze and resume behavior Kubernetes does not have.

Why did AX and Substrate get so much attention on X?

Teams started running hundreds of agents and found classic server setups expensive while agents sat idle plus risky when agents ran untrusted code. A purpose-built runtime plus familiar management promised answers to both pains at once.

What is a Task in AX?

A unit of agent work running inside an isolated sandbox area, the smallest thing you create, inspect, suspend and resume.

What is a Workspace in AX?

The prepared environment an agent starts with: Git repos, MCP servers and tools set up before the first step runs.

What is a Gateway in AX?

The network policy layer that controls which outside servers and APIs an agent may reach.

What is a Model object in AX?

The configuration that selects which language model an agent uses and where its API keys are held centrally.

Should my team adopt AX today?

Read the primitives first, then test one non-critical workload in a sandbox with strict gateway rules and a cost cap. New platforms earn production trust slowly, one workload at a time.

Bu sayfanın Türkçesi

Turn reading into a credential

This post is a free field note. Exams run at dated sittings in 15-seat classes; one price covers one attempt. All lessons are free.