cd ~/series/demystifying-ai
NVIDIA PAIR - Turn Your LAN Into an Inference Cluster
EP 30 NEW RELEASE GitHub Sep 13, 2026 5 min

NVIDIA PAIR: Turn Your LAN Into an Inference Cluster

NVIDIA just shipped a free, open-source tool that discovers every compatible GPU on your local network and exposes them as a single inference endpoint. Gaming PCs, workstations, Macs. Apache 2.0.

Share:
// TL;DR

NVIDIA shipped a free, open-source router that connects every RTX GPU, DGX Spark, and Apple M4 Mac on your network into one Ollama/LM Studio endpoint. Up to 18 devices.

Nothing leaves the LAN. Apache 2.0.

# What PAIR Actually Does (And Doesn't)

PAIR stands for Personal AI Router. Launched Sep 3, 2026 at IFA. Open-sourced on GitHub under NVIDIA/Personal-AI-Router (v0.1.1 beta).

The name tells you exactly what it is. A router. Not a GPU pool.

What PAIR IS
✓ A request router that schedules whole requests to available nodes
✓ Concurrency scaling via more machines handling parallel requests
✓ Single endpoint speaking Ollama + OpenAI protocols
✓ LAN-only with mTLS encryption and 6-digit pairing
What PAIR ISN'T
✗ Not a GPU pool that merges hardware resources
✗ Not a VRAM combiner that spans memory across devices
✗ Not a model splitter that shards layers across nodes
✗ Not faster single requests from adding more machines
KEY DISTINCTION

Each request runs whole on a single node. PAIR picks which node runs it.

Adding machines means more parallel requests, not faster individual ones. Think load balancer, not supercomputer.

# Supported Hardware

The hardware list is broader than expected. NVIDIA routing requests to Apple silicon was nobody's bet. But here it is.

RTX 20+
Gaming GPUs
GeForce RTX 20-series and up
RTX PRO
Workstation
Professional GPU line
DGX Spark
Mini Datacenter
Desktop AI server
Apple M4+
The Surprise
macOS Tahoe required
THE HEADLINE

NVIDIA is routing inference requests to Apple silicon. Read that again. The GPU company is treating competitor hardware as a first-class inference node.

M4 Macs on macOS Tahoe are fully supported. The typical setup is now 1 MacBook + 1 gaming PC. 24GB+ VRAM recommended per node.

# How It Works

The architecture is intentionally simple. Discovery, pairing, encryption, routing.

No distributed training, no tensor parallelism, no fancy orchestration. Just a router that knows which machines are available.

ARCHITECTURE
1
LAN Discovery PAIR scans your local network for compatible devices
2
6-Digit Pairing Code Each device pairs with a one-time code, no passwords stored
3
mTLS Encrypted Channel All traffic between nodes is encrypted with mutual TLS
4
Single Endpoint Exposes one URL speaking Ollama + OpenAI API protocols
5
Route to Available Node Each request dispatched whole to the best available machine
Bundled Engines
✓ Ollama (built-in)
✓ LM Studio (built-in)
Custom Engines
+ vLLM (via JSON manifest)
+ Any engine via custom manifest
ENGINE MANIFEST EXAMPLE
// engines/vllm.json
{
  "name": "vllm",
  "protocol": "openai",
  "command": "vllm serve {"model"} --port {"port"}",
  "health_check": "/health",
  "models_endpoint": "/v1/models"
}

# The Numbers

NVIDIA tested PAIR with up to 18 devices in a single cluster. The performance gains come from kernel optimizations and speculative decoding, not from combining GPUs.

18
Max Devices Tested
1.5x
RTX 5090 Speedup
1.4x
Two-DGX Cluster
0
Data Leaving LAN
5090
RTX 5090 Gains Per-node Performance

Up to 1.5x speedup from kernel optimizations, speculative decoding, and faster prefill. These gains apply to each node individually. PAIR just makes them available to your whole network.

DGX
DGX Spark Clusters Aggregate Throughput

vLLM 1.x on RTX PRO 6000 Blackwell shows up to 1.4x throughput on two-DGX-Spark clusters. Not because GPUs merge. Because 2 machines serve 2 requests simultaneously.

SECURITY MODEL

Every prompt, every file, every response stays on the local network. Zero data leaves your LAN.

The 6-digit pairing code prevents unauthorized devices from joining. mTLS encrypts all traffic between nodes. No cloud relay, no phone-home, no telemetry endpoint.

// Bottom Line

NVIDIA just made every idle GPU in your house a production inference node. If you have 2 computers and a LAN cable, you have an AI cluster. The Mac support is the surprise move.

NVIDIA routing to Apple silicon was nobody's bet. Nothing leaves the network. Apache 2.0, free forever.

Enjoyed this?

New episodes Mon, Wed, Sat.