cd ~/series/demystifying-ai
Einsummable: Every Kernel Is a Join
EP 31 Hot Topic arXiv: 2609.03905 Sep 15, 2026 5 min

Einsummable: Every Kernel Is a Join

Write PyTorch. Get optimal multi-GPU sharding. Zero annotations. Zero device assignments. Zero communication code. Einsummable beats hand-tuned experts by modeling every GPU operation as a relational join.

Share:
// TL;DR

Write PyTorch. Get optimal multi-GPU sharding. Zero annotations. Zero device assignments. Zero communication code. Einsummable models every GPU operation as a relational join and finds decomposition plans that hand-tuned experts miss. 8.97 ms vs 13.80 ms (PyTorch) on LLaMA blocks.

# The Problem It Kills

Multi-GPU parallelism today requires PhD-level expertise. You need to understand tensor parallelism, pipeline parallelism, data parallelism, FSDP configurations, and the subtle ways they interact. Most teams spend weeks tuning sharding strategies that a compiler should handle.

Before Einsummable
Manual device placement
Sharding annotations per tensor
AllReduce / AllGather calls
FSDP / DeepSpeed configs
Weeks of tuning
Requires: distributed systems PhD
With Einsummable
Write PyTorch Standard model code. Nothing special.
↓
Einsummable Compiles Automatic sharding + distribution
↓
Multi-GPU Execution 35% faster than hand-tuned
Requires: nothing
KEY INSIGHT

Current auto-parallelizers (Megatron, Alpa, Unity) choose from a menu of named strategies: tensor parallelism, pipeline parallelism, data parallelism. Einsummable does not pick from a menu. It searches the full space of possible decompositions and discovers plans those tools structurally cannot find.

# How It Works: Joins, Not Strategies

The core insight is radical. Every tensor operation is a relational join followed by an aggregation. Matrix multiply, convolution, attention, normalization. All of them. Once you model operations this way, the parallelism problem becomes a database query optimization problem.

01
PyTorch Input Standard Model Description

Accept a PyTorch-like description of the AI computation. No annotations, no device maps, no sharding specs. Just the math your model needs to compute.

↓
02
Join-Agg Specs Relational Decomposition

Each operation is modeled as a relational join + aggregation over tensor relations. Tuples contain sub-tensors. Each operation exposes all possible decompositions through "join-agg specs." This is where named strategies become irrelevant.

↓
03
Optimizer Searches All Decompositions

The optimizer selects decompositions across the whole computation graph to minimize a communication-cost proxy. It does not choose between TP, PP, and DP. It searches the complete space and finds hybrid plans that mesh-based auto-parallelizers cannot express.

↓
04
Exchange Programs Topology-Aware Communication

Instead of calling AllReduce or AllGather, Einsummable synthesizes custom exchange programs at compile time. These are topology-aware generalizations of Volcano's exchange operator. Every byte of communication is purpose-built for the specific computation and hardware layout.

↓
Multi-GPU Execution Automatically Distributed

Faster than hand-tuned code. The compiled plan runs across all GPUs with zero programmer intervention. On 8x A100: 8.97 ms geometric mean on LLaMA transformer blocks.

# The Benchmarks

Tested on an 8-GPU A100 server running LLaMA transformer blocks. A fully automatic system beats the humans who wrote the manual code.

8x A100 | LLaMA TRANSFORMER BLOCKS | GEOMETRIC MEAN
Einsummable 8.97 ms
FASTEST
Hand-tuned PyTorch 13.80 ms
+35% slower
vLLM 15.90 ms
+44% slower
35%
Faster Than PyTorch
44%
Faster Than vLLM
0
Lines of Sharding Code
READ THAT AGAIN

A fully automatic system, with zero human guidance, produced a faster execution plan than expert engineers who manually tuned the parallelism. The compiler did not match human performance. It exceeded it. By 35%.

# What Makes It Different

Einsummable is not an incremental improvement on existing auto-parallelizers. It rethinks the problem from first principles. 4 architectural choices separate it from everything else.

No Canned Collectives

Einsummable invokes zero standard collectives. No AllReduce. No AllGather. No ReduceScatter. All communication is special-purpose, derived at compile time for the specific computation graph.

Topology-Aware

Exchange programs adapt to the actual GPU interconnect topology. NVLink bandwidth, PCIe bottlenecks, NUMA boundaries. The communication plan is physically optimal, not just logically correct.

Relational Model

Tensors are treated as relations. Operations become joins. This reframing exposes decomposition possibilities that tensor-centric frameworks cannot see. Decades of relational query optimization apply directly.

Zero Annotations

No device maps. No sharding specs. No @parallelize decorators. No mesh definitions. The programmer writes single-device PyTorch. Everything else is the compiler's problem.

Einsummable
Alldecompositions
Synthcollectives
Noneannotations
Megatron-LM
TP+PP+DPsearch
Cannedcollectives
Manualannotations
Alpa
Intra+Intersearch
Cannedcollectives
Minimalannotations
FSDP / DeepSpeed
DP+Stagessearch
Cannedcollectives
Configannotations
THE PATTERN

The database community solved query optimization decades ago by searching decomposition spaces instead of choosing from named strategies. Einsummable applies that same principle to GPU parallelism. The result: plans that named-strategy tools cannot even represent, let alone discover.

// Bottom Line

Einsummable is the strongest evidence yet that manual multi-GPU parallelism is a solvable compiler problem, not an engineering art. It beat hand-tuned PyTorch by 35% with zero human guidance. It beat vLLM by 44%. The sharding annotations you wrote last month? A compiler can do better. The AllReduce calls you carefully placed? A compiler can synthesize something faster. The 10 authors behind this paper (Zhimin Ding + 9 others) have shown that the right abstraction eliminates the problem entirely.

Enjoyed this?

New episodes Mon, Wed, Sat.