PyTorch 2.14 ships 3 things that change daily ML work: NVGEMM auto-tuned CUTLASS kernels, fault-tolerant process groups where nodes crash and rejoin without restarting, and Apple Silicon native linear algebra. Released Sep 2, 2026. Upgrade now.
# NVGEMM: The Kernel Backend You Didn't Know You Needed
TorchInductor now has a new GEMM backend called NVGEMM. It generates CUTLASS kernels using CuTeDSL, then auto-tunes them alongside Triton and ATen.
The winner is not picked manually. The compiler does it.
One config flag - max_autotune_gemm_backends - enables it. The compiler profiles every candidate and picks the fastest kernel per operation. No manual kernel selection, no performance guessing.
import torch
# Enable all three GEMM backends for auto-tuning
torch._inductor.config.max_autotune_gemm_backends = "CUTLASS,ATen,Triton"
# Compile your model - NVGEMM auto-tunes automatically
model = torch.compile(model)
# Requires: nvidia-cutlass-dsl 4.6.0
# NVFP4 paths require Blackwell GPUs Fuses post-GEMM operations directly into the kernel. Fewer memory round-trips. Less overhead per matmul.
Native scaled GEMM support. NVFP4 paths on Blackwell GPUs. Grouped-reduction epilogues for advanced quantization workflows.
Unsupported epilogues? Falls back to Triton automatically. No crashes, no manual intervention - the compiler handles the edge cases.
Kernels are generated via CuTeDSL, not hand-written. Requires nvidia-cutlass-dsl 4.6.0. Auto-tuned against every other backend.
Before NVGEMM, you either used Triton (good default) or hand-tuned CUTLASS (expert-only). Now the compiler profiles all 3 backends and picks the winner per operation. Faster kernels, zero effort.
# Node Fails? Training Continues.
This is the headline feature. PyTorch 2.14 introduces fault-tolerant process groups as a first-class c10d concept - not a hack, not a wrapper.
A node can crash and rejoin without restarting the entire training job.
The enable_reconfigure flag in init_process_group turns it on. When a node dies, surviving nodes detect the failure and reconfigure around it. The dead node can come back and re-join.
import torch.distributed as dist
# Initialize with fault tolerance enabled
dist.init_process_group(
backend="nccl2",
enable_reconfigure=True
)
# Node crashes... other nodes detect failure
# Surviving nodes reconfigure automatically
dist._reconfigure(new_peers)
# Crashed node can rejoin without full restart
# ReconfigureOptions: uuid, peer handles, timeout, hints In-place process-group reconfiguration. The ReconfigureOptions struct carries uuid, ordered/unordered peer handles, timeout, and hints. Helper functions: _supports_reconfigure, _get_reconfigure_handle, _reconfigure.
Flight Recorder used to be NCCL-only. Now it works for ANY distributed backend. Debug distributed failures regardless of your communication layer.
The nccl2 backend from torchcomms landed in-tree. One-sided RMA (Remote Memory Access) windows let nodes access remote memory directly. Gloo has native reconfigure support today.
Large training jobs on hundreds of nodes lose machines regularly. Before 2.14, a single node failure meant restarting the entire job from the last checkpoint.
Now the cluster reconfigures and keeps training. Hours saved per failure. No human intervention needed.
# Apple Silicon Gets Real
PyTorch on Apple Silicon used to mean "attention kernels work." That is no longer the story.
2.14 ships native linear algebra across the MPS backend. The Mac is becoming a first-class ML development platform.
Moving beyond attention kernels to full native linalg means most training and inference workloads now run natively on Apple Silicon. M-series Macs with unified memory become legitimate ML development machines, not just "good enough for prototyping."
# What Else Shipped
The big 3 features get the headlines. But 2.14 shipped a set of smaller changes that matter for anyone using torch.compile and torch.export in production.
Shapes become declarative rather than traced. The @dynamic_spec decorator works with compile, export, and tracing. No more shape surprises in production.
A new control flow primitive for branching inside compiled graphs. Cleaner than nested torch.cond chains. Makes compiled models with conditional logic more readable.
CUDA graphs now capture torch.while_loop. Iterative algorithms that were previously excluded from graph capture can now run at full CUDA graph speed.
Complex-valued tensors now work with torch.compile. Signal processing, quantum computing, and physics simulation workloads get compilation support at last.
PyTorch 2.14 is the infrastructure release. NVGEMM gives you faster kernels for free. Fault tolerance means training jobs survive hardware failures without human intervention.
Apple Silicon gets promoted from experiment to supported platform. This is the release to upgrade for, not because of a new feature, but because the foundation got stronger.
Enjoyed this?
New episodes Mon, Wed, Sat.