Culture

PyTorch Conference 2026 Showcases Ray Integrations

The upcoming PyTorch Conference North America 2026 in San Jose will highlight how developers are using the Ray distributed compute framework to scale massive artificial intelligence workloads.

PyTorch Blog1 day agoCulture
Image: PyTorch Blog

The PyTorch Foundation has announced its lineup for the PyTorch Conference North America 2026 in San Jose, emphasizing the growing role of the Ray distributed compute framework. As organizations build out secure AI infrastructure, Ray has emerged as a critical tool for scaling workloads across the entire development lifecycle. Keynote speaker Ion Stoica of Anyscale, Databricks, and UC Berkeley will discuss the co-evolution of Ray and Kubernetes alongside Google's Jago Macleod, exploring how these technologies are merging into a unified open-source AI stack.

Several major tech companies will share production case studies demonstrating Ray's performance benefits. LinkedIn engineers Tommy Li and Tao Huang will present a three-layer Ray-based training stack that utilizes FSDP and HSDP. By routing data loading through dedicated CPU nodes, LinkedIn cut dataloader memory by 50% to 70% and reclaimed 24% of its step time. Meanwhile, Uber will detail how migrating its Eats recommendation stack from TensorFlow and Horovod to PyTorch and Ray Data yielded a 5x speedup, a 90% reduction in data transformation memory, and a 20x boost in training throughput.

Other sessions will focus on specialized AI architectures and post-training platforms. Pinterest will introduce PTEnv, a post-training lifecycle harness that supports reinforcement learning via Ray and supervised fine-tuning with MS-Swift using PyTorch FSDP or DDP. For agentic workflows, Anyscale will showcase SkyRL, a modular reinforcement learning library that enables asynchronous training on mixture-of-experts (MoE) models with over 350 billion parameters using Megatron and vLLM. Additionally, the University of Minnesota will present PC-RF, a physics-constrained generative model for climate downscaling served via ONNX, Ray, and vLLM.

To address data bottlenecks, Google researchers will present Rapid Storage, a high-throughput gRPC-based protocol via fsspec designed to keep GPUs active. Finally, a session on Ray Core will detail how the technology scales scheduling across 10,000-node clusters and utilizes Ray Direct Transport to transfer PyTorch tensors directly between GPUs. Together, these presentations demonstrate Ray's transition from a simple scaling library into a foundational infrastructure layer for modern AI.

This is our own summary of reporting by PyTorch Blog

More in Culture