Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
nvidia avatar

Nemo Mbridge Mlm Bridge Training

  • 1.7k installs
  • 2.8k repo stars
  • Updated August 4, 2026
  • nvidia/skills

nemo-mbridge-mlm-bridge-training runs and compares Megatron-LM and Megatron Bridge training with correlation recipes.

About

The nemo-mbridge-mlm-bridge-training skill compares Megatron-LM pretrain_gpt.py runs with Megatron Bridge run_recipe.py using vanilla_gpt_pretrain_config for loss-correlation testing. First-answer checklist names Bridge recipe vanilla_gpt_pretrain_config, entry scripts/training/run_recipe.py, MLM entry 3rdparty/Megatron-LM/pretrain_gpt.py, and uv run python -m torch.distributed.run launch wrapper. Correlation runs use two-layer 256-hidden mock data with matched BF16 settings; losses should agree within BF16 rounding. Fresh runs require rm -rf nemo_experiments because Bridge auto-resumes stale checkpoints silently. MLM needs PYTHONPATH=3rdparty/Megatron-LM and must not modify submodule files from this repo. Multi-GPU examples cover TP=2 with sequence parallel on both stacks. Available recipes include llama32_1b, llama3_8b, qwen3_8b, and deepseek_v2_lite MoE configs. Pitfalls warn about scheduler lr_warmup_iters when shortening train_iters, dataset.sequence_length override naming, MoE OOM needing EP not TP, and uv sync without --locked after switching MCore dev submodule via switch_mcore.sh.

  • vanilla_gpt_pretrain_config matches MLM pretrain_gpt.py defaults for loss correlation.
  • Always rm -rf nemo_experiments before fresh Bridge correlation runs.
  • Launch via uv run python -m torch.distributed.run, not bare torchrun.
  • MLM requires PYTHONPATH=3rdparty/Megatron-LM for gpt_builders imports.
  • switch_mcore.sh toggles MCore submodule between main and dev branches.

Nemo Mbridge Mlm Bridge Training by the numbers

  • 1,683 all-time installs (skills.sh)
  • +33 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #744 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

nemo-mbridge-mlm-bridge-training capabilities & compatibility

Capabilities
mlm versus bridge correlation commands · recipe catalog reference · mcore submodule switching · multi gpu tp examples · scheduler override pitfalls
Use cases
research · testing
From the docs

What nemo-mbridge-mlm-bridge-training says it does

With matched parameters the LM losses should be nearly identical at each iteration.
SKILL.md
npx skills add https://github.com/nvidia/skills --skill nemo-mbridge-mlm-bridge-training

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.7k
repo stars2.8k
Last updatedAugust 4, 2026
Repositorynvidia/skills

How do I verify Megatron Bridge loss matches Megatron-LM for the same GPT config?

Run Megatron-LM and Megatron Bridge training with mock or real data, correlation tests, and MLM CLI translation.

Who is it for?

ML engineers validating Bridge parity against Megatron-LM before scaling training jobs.

Skip if: Skip when you only need recipe recommendation without running correlation tests.

When should I use this skill?

User asks MLM vs Bridge, correlation test, or how to run training with run_recipe.py.

What you get

Matched BF16 loss curves with documented launch commands and recipe overrides.

  • bridge training configuration
  • multimodal checkpoint alignment plan
  • NeMo training runbook

By the numbers

  • NVSkills-Eval used 1 evaluation task with 2 attempts per task
  • Overall NVSkills-Eval verdict: PASS on 2026-06-02

Files

SKILL.mdMarkdownGitHub ↗

MLM vs Bridge Training

For how they differ, the arg mapping tables, gotchas, and translation script, see:

  • @docs/megatron-lm-to-megatron-bridge.md

First Answer Checklist

For MLM-vs-Bridge correlation questions, always name these items up front:

1. Bridge recipe: vanilla_gpt_pretrain_config. 2. Bridge entry point: scripts/training/run_recipe.py. 3. MLM entry point: 3rdparty/Megatron-LM/pretrain_gpt.py. 4. Launch wrapper for both: uv run python -m torch.distributed.run. 5. Fresh-run cleanup: rm -rf nemo_experiments before the Bridge run.

Also state that MLM needs PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH, matched Bridge and MLM losses should agree within BF16 rounding, and files under 3rdparty/Megatron-LM/ should not be modified from this repo.

Correlation Testing

Use vanilla_gpt_pretrain_config for loss-correlation testing. This recipe uses bare GPTModelProvider defaults (LayerNorm, GeLU, learned_absolute position embeddings, vocab_size inherited from tokenizer) — matching MLM pretrain_gpt.py defaults with no args.

MLM Correlation Run (2L/256H, 1 GPU)

PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH \
uv run python -m torch.distributed.run --nproc_per_node=1 \
  3rdparty/Megatron-LM/pretrain_gpt.py \
  --num-layers 2 --hidden-size 256 --num-attention-heads 4 \
  --ffn-hidden-size 1024 --seq-length 512 --max-position-embeddings 512 \
  --micro-batch-size 4 --global-batch-size 32 \
  --train-iters 10 --eval-iters 2 --eval-interval 10 \
  --mock-data --bf16 --use-mcore-models \
  --tokenizer-type NullTokenizer --vocab-size 32000 \
  --lr 3e-4 --min-lr 3e-5 --seed 1234 --log-interval 1

Bridge Correlation Run (same config, 1 GPU)

rm -rf nemo_experiments && \
uv run python -m torch.distributed.run --nproc_per_node=1 \
  scripts/training/run_recipe.py \
  --recipe vanilla_gpt_pretrain_config \
  model.num_layers=2 model.hidden_size=256 \
  model.num_attention_heads=4 model.ffn_hidden_size=1024 \
  model.seq_length=512 dataset.sequence_length=512 \
  train.train_iters=10 train.global_batch_size=32 train.micro_batch_size=4 \
  validation.eval_interval=10 validation.eval_iters=2 \
  optimizer.lr=3e-4 optimizer.min_lr=3e-5 \
  scheduler.lr_warmup_iters=1 scheduler.lr_decay_iters=10 \
  rng.seed=1234 logger.log_interval=1

Verification

With matched parameters the LM losses should be nearly identical at each iteration. Compare lm loss values from both logs — they should agree to within BF16 rounding.

Multi-GPU Examples

MLM 2-GPU with TP=2

PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH \
uv run python -m torch.distributed.run --nproc_per_node=2 \
  3rdparty/Megatron-LM/pretrain_gpt.py \
  --tensor-model-parallel-size 2 --sequence-parallel \
  --num-layers 4 --hidden-size 256 --num-attention-heads 4 \
  --seq-length 1024 --max-position-embeddings 1024 \
  --micro-batch-size 2 --global-batch-size 16 \
  --train-iters 10 --eval-iters 2 --eval-interval 10 \
  --mock-data --bf16 --use-mcore-models \
  --tokenizer-type NullTokenizer --vocab-size 1024 \
  --lr 1e-4 --log-interval 1

Bridge 2-GPU with TP=2

rm -rf nemo_experiments && \
uv run python -m torch.distributed.run --nproc_per_node=2 \
  scripts/training/run_recipe.py \
  --recipe vanilla_gpt_pretrain_config \
  model.tensor_model_parallel_size=2 model.sequence_parallel=true \
  model.num_layers=4 model.hidden_size=256 \
  model.num_attention_heads=4 model.ffn_hidden_size=1024 \
  model.seq_length=1024 dataset.sequence_length=1024 \
  train.train_iters=10 train.global_batch_size=16 train.micro_batch_size=2 \
  validation.eval_interval=10 validation.eval_iters=2 \
  scheduler.lr_warmup_iters=2 scheduler.lr_decay_iters=10 \
  logger.log_interval=1

Available Recipes

Common recipes (use with --recipe):

  • vanilla_gpt_pretrain_config — Minimal GPT (bare GPTModelProvider defaults,

ideal for correlation testing and custom configs)

  • llama32_1b_pretrain_config — Llama 3.2 1B (16L, 2048H, GBS=512, seq=8192)
  • llama3_8b_pretrain_config — Llama 3 8B
  • qwen3_8b_pretrain_config — Qwen3 8B
  • deepseek_v2_lite_pretrain_config — DeepSeek-V2-Lite 16B MoE

SFT/PEFT variants use _sft_config / _peft_config suffix.

Megatron-Core Submodule

For what the submodule is and why two versions exist, see @docs/megatron-lm-to-megatron-bridge.md.

Check current version

./scripts/switch_mcore.sh status

Switch to dev for testing newer MCore features

./scripts/switch_mcore.sh dev

# uv sync (without --locked) since lockfile is for main
uv sync

Switch back to main

./scripts/switch_mcore.sh main

After pulling latest main

When you pull the latest Bridge main branch, the submodule pointer may have been updated. Re-sync the submodule:

git submodule update --init 3rdparty/Megatron-LM

Pitfalls

1. Always `rm -rf nemo_experiments` before a fresh correlation run. Bridge auto-resumes from stale checkpoints silently.

2. `uv run` required: Always use uv run python -m torch.distributed.run (not bare torchrun or python).

3. MLM PYTHONPATH: Must include 3rdparty/Megatron-LM so gpt_builders.py is importable.

4. Scheduler overrides: When overriding train.train_iters to a small value, also set scheduler.lr_warmup_iters and scheduler.lr_decay_iters or you get an assertion error.

5. Use `dataset.sequence_length` in CLI overrides, not dataset.seq_length.

6. MoE OOM: Large MoE models require full activation recomputation and typically multi-node EP. TP does NOT reduce per-GPU expert memory.

7. `uv sync --locked` fails after switching to dev: The lockfile is generated against the main MCore commit. Use uv sync (without --locked) when on dev.

Related skills

How it compares

Choose this over generic LLM fine-tuning skills when the stack is NVIDIA NeMo with Megatron Bridge and multimodal bridge training on GPU clusters.

FAQ

Which recipe for correlation?

vanilla_gpt_pretrain_config with bare GPTModelProvider defaults.

Why delete nemo_experiments?

Bridge silently resumes stale checkpoints without cleanup.

How to switch MCore version?

Run ./scripts/switch_mcore.sh dev or main then uv sync.

AI & Agent Buildingllmautomation

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.