Models

LG EXAONE 3.5 Model Nears 1 Million Downloads

LG AI Research's quantized EXAONE 3.5 7.8B model is approaching one million monthly downloads on Hugging Face, proving the high demand for efficient, localized bilingual AI models.

AlphaSignal4 days agoModels
Image: AlphaSignal

LG AI Research is seeing massive developer interest in its quantized EXAONE 3.5 7.8B Instruct model, which has surpassed 849,000 monthly downloads on Hugging Face. The bilingual English and Korean model utilizes activation-aware weight quantization (AWQ) to shrink its footprint, making it highly attractive for local deployment. This specific version uses W4A16 group-wise quantization with groups of 128 weights, compressing the model down to approximately 5GB of VRAM.

Under the hood, the decoder-only Transformer features 6.98 billion non-embedding parameters across 32 layers. It employs grouped-query attention with 32 query heads and 8 key/value heads, alongside a 102,400-token byte-level BPE vocabulary. While the quantized weights fit into 5GB, running inference at its maximum 32,768-token context window requires additional memory. For instance, a full 32K sequence at a batch size of one adds about 4 GiB of overhead for the 16-bit key-value cache alone, meaning practitioners must budget for extra GPU memory beyond the base weight size.

To achieve its long-context capabilities, LG extended the sequence length from the 4,096 tokens of EXAONE 3.0 to 32,768 tokens using a rotary-position-embedding theta of 1,000,000. The model was trained on 9 trillion tokens using two-stage pre-training, taxonomy-based supervised fine-tuning (SFT), and staged DPO/SimPO alignment. This rigorous training allows EXAONE 3.5 to score 70.7 on real-world benchmarks, outperforming rival models like Qwen 2.5 7B, which scored 52.7, and Llama 3.1 8B, which scored 48.6. Currently, the model is released under a research-only EXAONE 1.1 NC license, meaning commercial applications require direct licensing from LG AI Research.

This is our own summary of reporting by AlphaSignal

More in Models