NVIDIA Accelerates BioNeMo MoE Training by 2.2x
NVIDIA has optimized its BioNeMo training recipe for biological foundation models, achieving a 2.21x throughput boost on B200 GPUs to make large-scale genomic modeling highly efficient.

NVIDIA has released an optimized training recipe within its BioNeMo framework designed to accelerate Mixture-of-Experts (MoE) biological foundation models. In a training benchmark utilizing eight NVIDIA B200 Tensor Core GPUs, the new recipe achieved up to 2.21 times the throughput of a standard Hugging Face baseline when training a Mixtral-8x7B model. This performance leap addresses the severe computational and memory bottlenecks that typically arise when scaling up sparse MoE architectures for complex biological and genomic datasets.
To achieve these speeds, the recipe leverages the NVIDIA Transformer Engine to replace slow, iterative Python loops with grouped expert execution. Instead of launching individual PyTorch operations for each expert, the engine's GroupedLinear feature batches multiple expert matrix multiplications into a single grouped operation. Additionally, the recipe introduces MXFP8 block-scaled 8-bit precision, which is hardware-accelerated on NVIDIA Blackwell GPUs. This format slashes memory consumption compared to traditional 16-bit BF16 precision, allowing practitioners to handle the extremely long sequences common in genomics without running out of activation memory.
The system further optimizes training by using the Transformer Engine Sequential API to fuse multiple steps. It merges GroupedLinear, ScaledSwiGLU, and routing-weight scaling into a single specialized kernel, preventing the need to materialize intermediate data. For AI practitioners, these advancements mean they can scale up the parameter capacity of biological models without suffering from fragmented GPU utilization or excessive quantization overhead. Developers can validate the setup using a two-GPU sanity configuration before scaling to a full eight-GPU Mixtral-8x7B run using PyTorch's FSDP2 and expert parallelism.
This is our own summary of reporting by NVIDIA Developer Blog



