NVIDIA shipped a free, open-source router that connects every RTX GPU, DGX Spark, and Apple M4 Mac on your network into one Ollama/LM Studio endpoint. Up to 18 devices.
Nothing leaves the LAN. Apache 2.0.
# What PAIR Actually Does (And Doesn't)
PAIR stands for Personal AI Router. Launched Sep 3, 2026 at IFA. Open-sourced on GitHub under NVIDIA/Personal-AI-Router (v0.1.1 beta).
The name tells you exactly what it is. A router. Not a GPU pool.
Each request runs whole on a single node. PAIR picks which node runs it.
Adding machines means more parallel requests, not faster individual ones. Think load balancer, not supercomputer.
# Supported Hardware
The hardware list is broader than expected. NVIDIA routing requests to Apple silicon was nobody's bet. But here it is.
NVIDIA is routing inference requests to Apple silicon. Read that again. The GPU company is treating competitor hardware as a first-class inference node.
M4 Macs on macOS Tahoe are fully supported. The typical setup is now 1 MacBook + 1 gaming PC. 24GB+ VRAM recommended per node.
# How It Works
The architecture is intentionally simple. Discovery, pairing, encryption, routing.
No distributed training, no tensor parallelism, no fancy orchestration. Just a router that knows which machines are available.
// engines/vllm.json { "name": "vllm", "protocol": "openai", "command": "vllm serve {"model"} --port {"port"}", "health_check": "/health", "models_endpoint": "/v1/models" }
# The Numbers
NVIDIA tested PAIR with up to 18 devices in a single cluster. The performance gains come from kernel optimizations and speculative decoding, not from combining GPUs.
Up to 1.5x speedup from kernel optimizations, speculative decoding, and faster prefill. These gains apply to each node individually. PAIR just makes them available to your whole network.
vLLM 1.x on RTX PRO 6000 Blackwell shows up to 1.4x throughput on two-DGX-Spark clusters. Not because GPUs merge. Because 2 machines serve 2 requests simultaneously.
Every prompt, every file, every response stays on the local network. Zero data leaves your LAN.
The 6-digit pairing code prevents unauthorized devices from joining. mTLS encrypts all traffic between nodes. No cloud relay, no phone-home, no telemetry endpoint.
NVIDIA just made every idle GPU in your house a production inference node. If you have 2 computers and a LAN cable, you have an AI cluster. The Mac support is the surprise move.
NVIDIA routing to Apple silicon was nobody's bet. Nothing leaves the network. Apache 2.0, free forever.
Enjoyed this?
New episodes Mon, Wed, Sat.