Research

MIT's Alex Zhang Drives Shift to Recursive Language Models

MIT PhD candidate Alex Zhang is pioneering recursive language models to bypass the limits of standard autoregressive decoders and unlock hidden capabilities in frontier AI.

Latent Space6 hrs agoResearch
Image: Latent Space

MIT researcher Alex Zhang is leading a paradigm shift toward Recursive Language Models (RLMs), a system design that allows models to treat their own prompts as external objects. This architectural evolution aims to solve the capability overhang of frontier models by replacing primitive scaffolding with self-improving harnesses. An RLM-based harness recently became the first to solve the challenging ARC-AGI-3 benchmark, beating OpenAI's Astra. This transition mirrors the 2025 shift from standard language models to reasoning models, positioning 2026 as the year of recursive systems that manage their own context and execute programmatic subagents.

Zhang's work also intersects with GPU Mode, a community he joined in 2023 focused on automating GPU kernel development. While frontier models like GPT-5.6 have successfully written highly efficient kernels—reducing costs for platforms like Terra and Luna by 80 percent—Zhang emphasizes that human expertise remains vital. On benchmarks like KernelBench, AI-generated solutions often suffer from verification and stability issues in end-to-end systems. Expert developers can still find elegant optimizations that replace enormous amounts of brute-force token search, bypassing the need to burn massive computational budgets.

This tension between brute-force scale and elegant design is highly visible in massive multi-agent experiments. For instance, OpenAI conducted a 10,000-agent experiment that consumed 130 billion output tokens to achieve approximately 40 million dollars in equivalent problem-solving. However, Zhang suggests that much of this swarm-based search is wasted. Instead of relying on massive, expensive runs, practitioners can use structured harnesses like Prime Agent to enable persistent agent-to-agent communication.

For practitioners, these developments redefine system design. Rather than treating models like Claude Code, Codex, or Pi as simple text-to-text decoders, developers must learn to design continuous harnesses. This academic path has historically minted industry leaders: past featured researchers include Shunyu Yao in 2024, who built Operator at OpenAI and became Chief AI Scientist of Tencent, and Jack Morris in 2025, who cofounded Engram at a 600 million dollar valuation. By leveraging benchmarks like SWE-bench, Quiet-STaR, and ReAct, developers can build invisible agent swarms operating beneath simple interfaces. This approach allows academics to take unconventional research bets, driving the industry toward self-directed Context Language Models (CLMs) that learn optimal policies directly within their weights.

This is our own summary of reporting by Latent Space

More in Research