The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

Managing AI agents like Kubernetes pods: Tasks, Workspaces, Gateways and Models

Updated

Engineers discussing delivery work in front of a screen

AX borrows the Kubernetes playbook for agents: declare Tasks with prepared Workspaces, fenced Gateways and central Models, then manage the fleet with short CLI verbs.

Kubernetes won operations by turning servers into declarations. You stopped nursing machines and started describing the desired state, then trusted controllers to close the gap. AX applies the same trick to a harder fleet: agents that think, wait, write code and hold state. The objects differ, but the muscle memory transfers.

The four objects, slowly

Task: the unit of work. A Task names one agent run: which Workspace it boots from, which Gateway fences it, which Model it calls, and what limits bind it. Operators list Tasks the way they list pods: which are running, which are suspended, which are failing, and for how long.

Workspace: repeatability. Every irreproducible agent incident starts with it worked on my machine. The Workspace declaration ends that excuse: exact repo revisions, exact MCP servers, exact tool versions. Two Tasks from the same Workspace revision should behave the same, and when they do not, the difference is evidence instead of mystery.

Gateway: the fence. Covered in depth in the sandboxing post, but operationally it is the object you will edit most. New tool needed? Add one destination with a reason. Incident review? Read the denied calls first. The gateway log is the firewall log of the agent era.

Model: the brain contract. Which model, which limits, whose keys. Separating the Model from the Task means you can move a fleet to a cheaper or calmer model without rewriting every agent, the same way changing an image tag redeploys a fleet without rewriting the deployment.

A day in the life of an agent operator

Morning starts with a fleet listing: how many Tasks ran overnight, how many suspended on approval gates, how many hit their step budgets. A stuck triage Task gets a shell session: the operator reads its workspace files, checks the gateway denials, finds a changed API response the agent could not parse, and resumes after fixing the seed file. Afternoon brings a rollout: the new Workspace revision goes to five Tasks first. Error rates hold, so the change widens to fifty. Evening ends with budgets: one research Task burned twice its model allowance chasing a vague question, so its owner tightens the question and the cap together.

None of this requires believing in autonomy hype. It is ticket queue work with new nouns.

Worked example: a fictional rollout with guardrails

The context below is fictional. Fictional team ParcelTrack (fictional) runs sixty documentation Tasks that rewrite runbooks from incident logs. The new Workspace revision adds a stricter style guide.

The operator declares the new revision on six Tasks and compares: approval rates, style violations, model spend per merged page. Five of six improve on every axis; the sixth loops on one malformed log. The malformed log is quarantined as a fixture, the revision widens to all sixty, and the old revision stays declared but scaled to zero for a week. If regressions appear, the rollback is one declaration change, not sixty panicked edits.

Decision table: which object to touch first

SymptomFirst objectWhy
Agent behaves differently across runsWorkspaceUndeclared inputs differ; pin them
Agent reaches somewhere it should notGatewayThe fence has a hole; close it
Costs spike without more outputModel and Task limitsThe brain or the budget is unbounded
Agent stuck mid-taskTask suspend and inspectFreeze first, debug second, evidence preserved
Whole fleet degradesWorkspace revisionOne shared input changed; diff it

Checklist: fleet habits worth copying from Kubernetes

  1. Every Task declares owner, budget and revision. No anonymous agents.
  2. Workspace revisions are immutable and tagged, never edited in place.
  3. The suspend, inspect, resume and destroy path is written and rehearsed.
  4. Gateway denials are reviewed weekly like firewall logs.
  5. Rollouts go to a canary slice first, with rollback as one declaration.

Straight answers

Frequently asked questions

How is managing agents like managing pods?

You declare desired state, controllers converge reality, and you inspect with short commands. Tasks play the role pods play: the unit you create, watch and restart.

What does the AX CLI do?

It applies declarations and inspects live objects with familiar verbs such as apply, get, suspend, resume and ssh, according to early descriptions.

What goes into a Workspace declaration?

Git repos and revisions, MCP servers and credentials references, tool versions and seed files, everything the agent needs before step one.

What goes into a Gateway declaration?

Allowed destination hosts, rate limits where supported, and logging rules, with everything else denied by default.

What goes into a Model declaration?

The model identifier, spending or rate limits, and a reference to centrally held keys, so prompts and files stay key-free.

How do you debug a stuck agent Task?

Get its status and recent events, open a shell into the sandbox, read the workspace files and the gateway log, then suspend or resume instead of deleting evidence.

How do you roll out a new agent version?

Same as services: declare the new Workspace revision on a few Tasks first, compare outcomes, then widen. Keep the old revision declared until the new one proves itself.

Do I still need runbooks for agent fleets?

More than before. Agents fail in new ways, so the suspend, inspect, resume and destroy paths must be written down and rehearsed.

Where do costs show up in this model?

Per Task and per Model: active compute seconds, model call counts and snapshot storage. Tag every Task with an owner and a budget.

Bu sayfanın Türkçesi

Turn reading into a credential

This post is a free field note. Exams run at dated sittings in 15-seat classes; one price covers one attempt. All lessons are free.