Google AX (Agent eXecution) is an open-source framework that brings Kubernetes-style orchestration to LLM agents. You define agents in YAML manifests. AX handles scheduling, scaling, health monitoring, state persistence, and rolling upgrades. It runs on any Kubernetes cluster and supports Gemini, GPT, Claude, Llama, or any OpenAI-compatible endpoint. 8 Custom Resource Definitions. Apache 2.0. If you've deployed microservices, you already know how to deploy agents.
# The Kubernetes Analogy
AX maps every Kubernetes concept you already know to an agent equivalent. The mental model is immediate if you've ever written a Deployment YAML. Here is the full CRD mapping.
| Kubernetes | Google AX | What it does |
|---|---|---|
| Pod | AgentInstance | One running agent with its model, tools, and context window |
| Deployment | AgentDeployment | Desired state: replicas, model version, rollout strategy |
| Service | AgentEndpoint | Stable gRPC/HTTP URL to reach an agent pool |
| HPA | AgentScaler | Scale on queue depth, latency, or token throughput |
| ConfigMap | AgentConfig | System prompts, tool definitions, guardrails, model params |
| PVC | AgentMemory | Persistent conversation history, vector embeddings |
| Job | AgentTask | One-shot agent execution with completion tracking |
| NetworkPolicy | AgentPolicy | Which tools an agent can access, rate limits, cost caps |
# Deploy Your First Agent
Here is a complete agent deployment manifest. If you can read a Kubernetes Deployment, you can read this.
apiVersion: ax.google.com/v1alpha1 kind: AgentDeployment metadata: name: coding-agent namespace: production spec: replicas: 3 model: provider: openai-compatible endpoint: http://vllm-atria:8000/v1 model: mistralai/Atria-Dawn-FP8 maxTokens: 32768 config: systemPrompt: | You are a senior software engineer. Review code, fix bugs, and write tests. Always explain changes. temperature: 0.1 tools: - name: file_read - name: file_write - name: shell_exec policy: sandboxed - name: github_api memory: type: persistent storageClass: ssd-retain size: 10Gi health: livenessProbe: promptCheck: "Respond with OK" periodSeconds: 30 readinessProbe: toolCheck: [file_read, shell_exec] periodSeconds: 60 rollout: strategy: canary canaryPercent: 20
# Deploy the agent kubectl apply -f coding-agent.yaml # Watch it come up kubectl get agentinstances -n production -w NAME MODEL READY TASKS coding-agent-0 Atria-Dawn-FP8 True 0 coding-agent-1 Atria-Dawn-FP8 True 2 coding-agent-2 Atria-Dawn-FP8 True 1 # Scale based on queue depth kubectl apply -f - <<EOF apiVersion: ax.google.com/v1alpha1 kind: AgentScaler spec: target: coding-agent minReplicas: 1 maxReplicas: 20 metrics: - type: queueDepth target: 5 # scale up when >5 tasks queued EOF # Rolling model upgrade (zero downtime) kubectl set model agentdeployment/coding-agent \ --model=mistralai/Atria-Dawn-v2-FP8
# Architecture
YAML files declare what the agent is, what tools it has, what model it uses, and how it should behave. kubectl apply -f agent.yaml deploys it. GitOps-native from day one.
Standard Kubernetes reconciliation loop. Watches CRDs and ensures actual state matches desired state. Handles scaling decisions, health check responses, canary deployments, and model version rollouts. Runs as a single Deployment in the ax-system namespace.
Lightweight Go sidecar injected into every AgentInstance pod. Manages the agent's event loop, tool execution sandbox, memory persistence, token streaming, and cost tracking. Exposes Prometheus metrics (ax_tokens_total, ax_task_duration_seconds) and OpenTelemetry traces for every agent turn.
GKE, EKS, AKS, k3s, or bare-metal. AX uses standard Kubernetes primitives. No GCP lock-in. Works with Istio, Linkerd, or no service mesh. Tested on K8s 1.28+.
# Why This Is Different
Every team building agents today has reinvented some version of "deploy agent, keep it alive, scale it, update the model." AX standardizes all of it. Here are the features no other framework offers.
Swap models mid-conversation. AX serializes the agent's context to AgentMemory, spins up a new instance with the upgraded model, hydrates context, and resumes. Zero downtime. Like a K8s rolling deployment but for LLM sessions.
Liveness probes send a simple prompt and check for a coherent response. Readiness probes verify tool connections work. If an agent loops or hallucinates repeatedly (detected via repetition score), the controller kills and restarts it with a fresh context.
AgentScaler watches a task queue (Redis, SQS, Pub/Sub, or built-in). More pending tasks → more agent replicas. Scales to zero when idle. You pay for compute only when agents are actually reasoning. Cooldown periods prevent thrashing.
Set per-agent or per-namespace token budgets. When an agent hits its cost cap, AX pauses the instance and fires a Kubernetes Event. No surprise $10K bills from a runaway coding agent loop.
- ● v1alpha1. API is unstable. Expect breaking changes before v1beta1 (targeted Q1 2027).
- ● No multi-agent coordination. AX manages individual agents. Agent-to-agent communication requires external plumbing (gRPC, message queues).
- ● Memory migration is lossy. Rolling upgrades between different model families (GPT → Claude) may lose context fidelity during serialization.
Google AX is the missing infrastructure layer for production agents. Every team building agents today has reinvented some subset of what AX provides. AX standardizes all of it using primitives every infrastructure engineer already knows. kubectl apply -f agent.yaml is the new docker run. It is v1alpha1, it has rough edges, and the multi-agent story is not there yet. But the pattern is right. Agents just got boring infrastructure. That's the highest compliment.
Zero-Token Confidence: Your Model Knows It's Wrong
A new technique extracts calibrated confidence scores from any LLM without generating a single extra token. Just read the hidden states.
Enjoyed this?
New episodes Mon, Wed, Sat.