LoongForge

Train LLMs, VLMs, Diffusion & Embodied models — faster.

Ready-to-run configs for 35+ model families — built on Megatron-LM and torch-native backends.

DreamZero training run compared side by side: LoongForge against the baseline Play the demo — 83 seconds, has background music
🎬 4.38× faster on this run — loss curves stay aligned with the baseline 📊 Benchmark
~5×
Max training speedup
35+
Model families supported
GPU·XPU
NVIDIA & Kunlun backends
Apache 2.0
Open source license
🧩
Easy
One framework, broad coverage
Full coverage of mainstream open-source LLMs, VLMs, MoE, diffusion, and VLA models. Ready-to-run configs and launch scripts included.
⚡
Efficient
Up to ~5× training speedup
Deep performance optimizations — fused kernels, adaptive FP8, MoE A2A overlap, and multimodal pipeline scheduling.
💎
Multi-chip
NVIDIA GPU & Kunlun XPU
Native heterogeneous hardware support — one framework, minimal migration between GPU and XPU.

📊 Benchmark

Training throughput speedup over mainstream open-source baselines — each model benchmarked on the same machine type with the same training hyperparameters

Pi0.5 VLA
2.80×
GR00T N1.6 VLA
2.31×
GR00T N1.7 VLA
1.79×
X-VLA VLA
1.79×
DreamZero WAM
4.38×
FastWAM WAM
2.25×
LingBot VA WAM
2.20×
Qwen3-VL-30B-A3B VLM
1.45×
DeepSeek-V3.2 Lite LLM
5.04×
1.0× baseline

DeepSeek-V3.2 Lite reflects DSA operator-level optimizations and was validated on a reduced-layer configuration due to test-bed scale limits.

Numbers were measured at a point in time and may evolve as implementations change on both sides.

Need a model LoongForge doesn't cover yet? Open an issue

🏗️ Architecture

One unified stack — from model composition down to GPU / XPU silicon.

LoongForge architecture

✨ Key Features

A quick tour of what sets LoongForge apart

🤖 Embodied
Embodied Training
Dedicated torch-native DDP/FSDP subsystem for VLA & WAM models, with DDP / ZeRO-1 / FSDP / HSDP.
🧠 MoE
Tri-Stream Overlap
MoE EP comm × compute × offload in parallel — higher throughput than upstream.
🎛️ Multimodal
Heterogeneous & Disaggregated
Independent TP / PP / DP per component + decoupled ViT / LLM scheduling that kills pipeline bubbles.
⚡ Performance
Adaptive FP8
Per-operator FP8 decisions by GEMM shape.
🔧 Performance
Fused Operators
FusedDSA / Sparse MLA kernels for end-to-end speedup.
🧵 Performance
ChunkPipe
Chunked long-sequence pipelining toward million-length contexts.
⚖️ Multimodal
DP Load Balancing
Fixes packing-induced imbalance at cluster scale.
🎯 Training
Pretrain + SFT + LoRA
One codebase covers key training stages.
🔁 Usability
HF ↔ Megatron
Bidirectional checkpoint conversion + online HF load/save.

💎 Hardware Compatibility

One codebase, two silicon stacks — production-ready on NVIDIA GPU and Baidu Kunlun XPU

NV

NVIDIA GPU

Built on the community Megatron + TransformerEngine ecosystem, with LoongForge optimizations layered on top.

昆

Kunlun XPU

XPU Plugin mechanism shields the upper stack from adaptation differences, while integrating an XPU-specific optimization toolchain.

🏛️ Supported Models

From compact SLMs to large-scale MoE giants — all batteries-included

DeepSeek-V2

v2-litev2

DeepSeek-V3

v3

DeepSeek-V3.2

v3.2

DeepSeek-V4

v4-flashv4-pro

LLaMA2

7B13B70B

LLaMA3

8B70B

LLaMA3.1

8B70B405B

Qwen

1.8B → 72B

Qwen1.5

0.5B → 72B

Qwen2

0.5B → 72B

Qwen2.5

0.5B → 72B

Qwen3

0.6B → 480B-A35BCoder-30B-A3B

Qwen3-Next

80B-A3B

MiniMax

m2.1m2.5m2.7

MIMO

mimo-7b

GLM

glm5

🚀 Quick Start

From install to launch — jump straight to the tutorial for your model type

1

Install

Recommended: one Docker image bundles the CUDA/XPU toolchains, patched Megatron, and TransformerEngine — so every node trains from the same environment. Source build is also supported.

Show Docker build commands
$ git clone --recurse-submodules \
    https://github.com/baidu-baige/LoongForge.git
$ docker build --build-arg COMPILE_ENV=hopper --build-arg ENABLE_LEROBOT=false \
    -t loongforge:latest \
    -f ./LoongForge/docker/Dockerfile .
$ docker build --build-arg ENABLE_LEROBOT=false \
    --build-arg BASE_IMAGE=loongforge/loongforge_kunlun:py310_torch25 \
    -t loongforge-kunlun:latest \
    -f LoongForge/docker/Dockerfile.xpu .
3

Explore & launch

Browse ready-to-run configs and example scripts to launch your first run.

🌟 Powered by LoongForge

Open-source models trained with LoongForge or its predecessor AIAK-Training-LLM

🤝 Community

Built in the open — join discussions, report issues, and contribute

👥 Contributors · 14+

LoongForge is built in the open by these developers — your name could be next.

🛠️ Become a contributor