Hardware

NVIDIA Dynamo-Triton Speeds Up HSTU Recommender Inference

NVIDIA launched an end-to-end inference workflow for HSTU generative recommenders using Dynamo-Triton, slashing latency to help companies serve highly personalized suggestions faster.

NVIDIA Developer Blog12 hrs agoHardware
Image: NVIDIA Developer Blog

NVIDIA has integrated support for Hierarchical Sequential Transduction Unit (HSTU) generative recommender models into its Dynamo-Triton inference server. Available through the NVIDIA recsys-examples repository, the new workflow addresses the high computational demands of sequential recommendation systems. By treating user interactions as continuous token streams, generative recommenders offer superior personalization but typically suffer from high latency when processing long histories.

To overcome these serving bottlenecks, the workflow combines PyTorch Ahead-of-Time Inductor (AOTI) compilation with FlexKV-powered key-value (KV) caching. PyTorch AOTI translates HSTU ranking models into native C++ binaries to bypass Python runtime overhead. Meanwhile, the FlexKV system manages a 20 GB GPU KV cache alongside host-side storage, saving reusable attention states so the model does not have to recompute a user's entire historical sequence on every new request.

In benchmark tests using the KuaiRand-1K ranking configuration on an NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU, the new pipeline demonstrated substantial performance gains. The evaluated models featured a hidden size of 512, four attention heads, BF16 weights, and a BF16 KV cache. The setup handled a maximum history sequence length of 8,192 tokens, comprising 4,096 item and action pairs, alongside six contextual features and a maximum candidate sequence length of 100. This yielded an effective sequence length of 8,298 tokens before alignment and 8,320 tokens after.

At a dynamic batch size of 8, the three-layer HSTU model achieved a 4.47x speedup with a 100 percent GPU KV-cache hit rate compared to the same AOTI configuration without caching, recording a latency of just 0.423 milliseconds per logical request. Under the same conditions, the deeper eight-layer HSTU model achieved a 5.93x speedup, bringing latency down to 0.678 milliseconds. The workflow also utilizes NV Embedding Cache to keep popular embeddings in GPU memory while storing the rest of the table in CPU memory, further optimizing hardware utilization.

This is our own summary of reporting by NVIDIA Developer Blog

More in Hardware