MLCommons Launches DLRMv4 Recommendation Benchmark
MLCommons has released DLRMv4, a new training benchmark that replaces traditional recommendation models with sequence-based architectures to better reflect modern hyperscale workloads.

MLCommons has introduced DLRMv4, the latest iteration of its machine learning training benchmark for recommendation systems. The update replaces the older DLRMv2 encode-then-interaction stack with a Hierarchical Sequential Transducer Unit (HSTU) architecture. This shift aligns the training suite with the HSTU-based DLRMv3 inference benchmark, reflecting how modern hyperscale operators process user histories as chronological token sequences rather than aggregated dense features.
To support this sequence-based approach, DLRMv4 transitions from the Criteo 1TB dataset to Yambda-5B, a public dataset from Yandex Music containing 4.79 billion interactions across 1 million users, 9.39 million items, and 5 behavior types. By adding cross-product feature tables, the benchmark expands the embedding table footprint to approximately 560 GB at fp32 precision, up from 100 GB in DLRMv2. This massive footprint forces practitioners to shard embedding tables across multiple accelerators, mirroring real-world production challenges.
The reference implementation uses a 3-layer HSTU stack with 4 attention heads, a model dimension of 512, and a maximum sequence length of 4,096. Training runs in bf16 mixed precision using Adam for dense parameters and row-wise Adagrad for sparse embeddings. The total computational cost is 323.26 TFLOPs per GPU step, which breaks down into 79.16 TFLOPs for UVQK GEMMs, 59.37 TFLOPs for output projection GEMMs, and 184.72 TFLOPs for causal HSTU attention.
For evaluation, the benchmark sets a convergence target of 0.75 AUC on held-out data. Testing on AMD MI350 GPUs across different global batch sizes demonstrated tight convergence consistency. At a batch size of 8,192 on 8 GPUs, the model converged in a mean of 69.3 million samples (ranging from 61.9M to 80.3M) over 2 hours and 39 minutes. A batch size of 16,384 on 16 GPUs required a mean of 87.7 million samples (ranging from 78.0M to 100.9M) over 1 hour and 52 minutes. Scaling to a batch size of 32,768 across 32 GPUs required a mean of 113.0 million samples (ranging from 98.6M to 128.5M), achieving convergence in 1 hour and 23 minutes. This standardized setup allows practitioners to reliably measure hardware and software efficiency on modern sequential recommendation workloads.
This is our own summary of reporting by ML Commons



