-
Notifications
You must be signed in to change notification settings - Fork 1.6k
Pull requests: modelscope/ms-swift
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Handle missing AgentFlan query for string suffixes
#9949
opened Aug 19, 2026 by
taking-lying-flat
Contributor
Loading…
Fix zero multimodal learning rates in Megatron
#9948
opened Aug 19, 2026 by
taking-lying-flat
Contributor
Loading…
Fix Megatron GRPO vocab-parallel log-prob gradients
#9947
opened Aug 19, 2026 by
taking-lying-flat
Contributor
Loading…
Fix LoRA IndexError when syncing weights to vLLM for Qwen models in GRPO colocate mode
#9946
opened Aug 19, 2026 by
jinchenyu
Loading…
1 of 4 tasks
The ViT gradient checkpoint was incorrectly turned off during multimodal model training.
#9945
opened Aug 19, 2026 by
z0o0ey
Collaborator
Loading…
1 of 4 tasks
Add --dataloader_multiprocessing_context to work around Python 3.14 incompatibility
#9941
opened Aug 18, 2026 by
sliedes
Loading…
1 of 4 tasks
[Megatron] Preserve RNG state across checkpoint resume
#9935
opened Aug 17, 2026 by
taking-lying-flat
Contributor
Loading…
[Megatron] Fix reference adapter loading for LoRA RLHF
#9934
opened Aug 17, 2026 by
taking-lying-flat
Contributor
Loading…
Fix reward model margin broadcasting and alignment
#9927
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
Fix DPO IPO log-prob normalization
#9925
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
docs: fix Reward Model links in RLHF guides
#9910
opened Aug 14, 2026 by
tutao0123
Loading…
1 of 4 tasks
Fix BatchSamplerShard tail sampling
#9907
opened Aug 14, 2026 by
taking-lying-flat
Contributor
Loading…
[Train] Reduce padding-free embedding output memory
#9893
opened Aug 11, 2026 by
taking-lying-flat
Contributor
Loading…
fix(deploy): handle reasoning-only stream chunks
#9881
opened Aug 11, 2026 by
taking-lying-flat
Contributor
Loading…
1 of 4 tasks
Fix IterablePackingDataset workers to use spawn
#9876
opened Aug 9, 2026 by
zupengwang
Loading…
1 of 4 tasks
[Feature] Add --streaming_shard: split the streaming dataset across data-parallel ranks
#9860
opened Aug 5, 2026 by
Rapisurazurite
Loading…
1 of 4 tasks
Fix provider message preprocessing regressions
#9840
opened Aug 4, 2026 by
taking-lying-flat
Contributor
Loading…
Fix Qwen3.5 multimodal packing kwargs compatibility
#9838
opened Aug 3, 2026 by
taking-lying-flat
Contributor
Loading…
Support model initialization without loading pretrained weights
#9821
opened Jul 30, 2026 by
LiXinYuECNU
Loading…
feat(sft):Add optional fused linear ce for qwen2/qwen3 on npu
#9784
opened Jul 22, 2026 by
yangguang-zhang
Loading…
1 of 4 tasks
Previous Next
ProTip!
Exclude everything labeled
bug with -label:bug.