-
Notifications
You must be signed in to change notification settings - Fork 456
All issues
Issue creation is restricted in this repository
- #3754 · anwithk opened
on May 8, 2026
Issues
is:issue state:open
is:issue state:open
Search results
[bug] nonlearnable params exist on different device than learnable parameters
area:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationbugSomething isn't workingSomething isn't workingwaiting-on-customerWaiting on the original author to respondWaiting on the original author to respondStatus: Open.#5574 In NVIDIA-NeMo/Megatron-Bridge;[support] Is it possible to convert only adapter weights when PEFT is in use?
area:peftParameter-efficient fine-tuning (LoRA, adapters)Parameter-efficient fine-tuning (LoRA, adapters)needs-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#5560 In NVIDIA-NeMo/Megatron-Bridge;PoR: [model] tracking: GLM-5.2 production training qualification
area:ckptCheckpoint conversion, loading, export, and save pathsCheckpoint conversion, loading, export, and save pathsarea:modelModel implementations and HF bridge logicModel implementations and HF bridge logicarea:perfPerformance optimizations and benchmarkingPerformance optimizations and benchmarkingarea:recipeTraining recipes and launch configsTraining recipes and launch configsarea:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationfeatureNew capabilities, enhancements, or enablement workNew capabilities, enhancements, or enablement workPoRPlan of record item for roadmap and release trackingPlan of record item for roadmap and release trackingtrackingTracking issue for an ongoing project with smaller stepsTracking issue for an ongoing project with smaller stepsStatus: Open.#5476 In NVIDIA-NeMo/Megatron-Bridge;[bug] MiniMax-M3 EP32 inference produces step-0 NaN logits with all-to-all
area:modelModel implementations and HF bridge logicModel implementations and HF bridge logicblockedWork cannot move forward until an external dependency is clearedWork cannot move forward until an external dependency is clearedbugSomething isn't workingSomething isn't workingStatus: Open.#5462 In NVIDIA-NeMo/Megatron-Bridge;[feature] Integrate Megatron-Core Generalized Tensor Parallelism (GTP)
area:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationfeatureNew capabilities, enhancements, or enablement workNew capabilities, enhancements, or enablement workneeds-triageNew item needs classification and ownershipNew item needs classification and ownershipStatus: Open.#5434 In NVIDIA-NeMo/Megatron-Bridge;[bug] BF16 sequence parallelism causes significant gradient norm drift with TP=4 on H200
area:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationbugSomething isn't workingSomething isn't workingwaiting-on-maintainersWaiting on maintainers to respondWaiting on maintainers to respondStatus: Open.#5411 In NVIDIA-NeMo/Megatron-Bridge;[bug] Warmup LR is one step behind: OptimizerParamScheduler not advanced before the first optimizer step
area:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationbugSomething isn't workingSomething isn't workingwaiting-on-maintainersWaiting on maintainers to respondWaiting on maintainers to respondStatus: Open.#5399 In NVIDIA-NeMo/Megatron-Bridge;[Tracking] Simplify Bridge conversion/modeling for standalone installation and RL integrations
area:ckptCheckpoint conversion, loading, export, and save pathsCheckpoint conversion, loading, export, and save pathsfeatureNew capabilities, enhancements, or enablement workNew capabilities, enhancements, or enablement worktrackingTracking issue for an ongoing project with smaller stepsTracking issue for an ongoing project with smaller stepsStatus: Open.#5371 In NVIDIA-NeMo/Megatron-Bridge;[bug] ddp.grad_reduce_in_fp32 CLI override is silently overwritten by mixed_precision.setup()
area:trainingTraining loop, callbacks, and runtime integrationTraining loop, callbacks, and runtime integrationbugSomething isn't workingSomething isn't workingwaiting-on-maintainersWaiting on maintainers to respondWaiting on maintainers to respondStatus: Open.#5365 In NVIDIA-NeMo/Megatron-Bridge;[support] Qwen3.5-VL-35B-A3B MFU is less than 10% on B300.
area:perfPerformance optimizations and benchmarkingPerformance optimizations and benchmarkingwaiting-on-maintainersWaiting on maintainers to respondWaiting on maintainers to respondStatus: Open.#5318 In NVIDIA-NeMo/Megatron-Bridge;[bug] DeepSeek V4 Pro FP8-MX pretraining gives gradient norm 'nan' on GB300
area:recipeTraining recipes and launch configsTraining recipes and launch configsbugSomething isn't workingSomething isn't workingwaiting-on-maintainersWaiting on maintainers to respondWaiting on maintainers to respondStatus: Open.#5306 In NVIDIA-NeMo/Megatron-Bridge;[bug] DeepSeek MLA with q_lora_rank=null builds a trainable Q-norm the HF architecture does not define
area:modelModel implementations and HF bridge logicModel implementations and HF bridge logicbugSomething isn't workingSomething isn't workingwaiting-on-maintainersWaiting on maintainers to respondWaiting on maintainers to respondStatus: Open.#5261 In NVIDIA-NeMo/Megatron-Bridge;