Hardware

d-Matrix buys Wallaroo to orchestrate chip inference

Chipmaker d-Matrix has acquired AI inference orchestration startup Wallaroo.ai to simplify the deployment of its Corsair accelerators alongside traditional GPUs in mixed-silicon data centers.

Unite.AI3 Aug 2026Hardware
Image: Unite.AI

Silicon startup d-Matrix announced on August 3, 2026, its acquisition of Wallaroo.ai, a developer of AI inference orchestration software. This acquisition, the chipmaker's second in four months, integrates Wallaroo's platform and engineering staff into d-Matrix. The move addresses a software challenge for d-Matrix's Corsair accelerators, which entered production on June 9, 2026. Manufactured with Alchip on TSMC's N6 node, Corsair uses SRAM-based in-memory compute chiplets and LP-DDR5 memory to run air-cooled. Instead of replacing GPUs, Corsair runs beside them, handling the memory-bound decode phase of language-model requests while GPUs manage the compute-heavy prefill. Wallaroo's control plane, supporting x86, Arm, GPU, vLLM, and SGLang, will orchestrate this split.

This acquisition reflects an industry trend where hardware vendors buy software layers; examples include Qualcomm's July 2026 purchase of Modular, Nscale's acquisition of Anyscale, and Nebius's $643 million deal for Eigen AI. To fund its expansion, d-Matrix relies on a $275 million Series C closed on November 12, 2025, at a $2 billion valuation, co-led by BullhoundCapital, Triatomic Capital, and Temasek. The company previously acquired GigaIO's data center business in April 2026, absorbing its SuperNODE system and FabreX PCIe-based memory fabric, and has partnered with Infineon.

For practitioners, managing mixed-silicon clusters is complex. Wallaroo founder Vid Jain noted on March 16, 2026, that agentic traffic is highly variable: 84% of requests are short (around 2,000 tokens), 15% are medium (8,000 to 64,000 tokens), and under 1% exceed 128,000 tokens. Their cache-aware routing cut worst-case time-to-first-token by 75% in tests. Furthermore, a March 11, 2026, study by Gimlet Labs—co-authored by d-Matrix CTO Sudeep Bhoja—demonstrated the power of this split. Running gpt-oss-120b with a 1.6-billion-parameter draft model, moving speculative decoding to Corsair's 2GB of on-chip SRAM and 150 TB/s bandwidth while keeping prefill on GPUs yielded 2x to 10x faster end-to-end requests. Already, Parasail is deploying Corsair alongside Nvidia Hopper and Blackwell fleets across 40 data centers in 15 countries. By owning Wallaroo's orchestration layer, d-Matrix CEO Sid Sheth can now deliver these performance gains without forcing developers to build custom routing software themselves.

This is our own summary of reporting by Unite.AI

More in Hardware