feat(laguna): support packed THD context parallelism - #3640
Conversation
|
/ok to test ea7baf2 |
|
🌿 Preview your docs: https://nvidia-preview-preview-0ee3a52b69d9.docs.buildwithfern.com/nemo/automodel |
2a72786 to
a0c66d6
Compare
|
cw-dfw matched loss parity completed on Slurm job |
|
/ok to test a0c66d6 |
46491d7 to
e39ee24
Compare
|
Conflict repair + validation update Rebased the two THD+CP commits onto current The matched cw-dfw validation remains:
Post-rebase local checks: 48 Laguna/block-diagonal CP tests passed; Ruff passed; recipe/model-coverage tests passed (5 passed, 1 skipped). |
|
/ok to test 10e7090 |
jgerh
left a comment
There was a problem hiding this comment.
Completed tech pubs review of docs/model-coverage/llm/poolside/laguna.mdx and provided a few copyedits
Co-authored-by: jgerh <163925524+jgerh@users.noreply.github.com>
|
/ok to test 0ee3a52 |
|
/ok to test 0ee3a52 |
What does this PR do ?
Add native packed THD plus context-parallel training support for Laguna, including full and sliding-window attention semantics.
This is stacked on the Laguna XS 2.1 recipe PR so this diff contains only the THD/CP implementation and validation recipe.
Changelog
Before your PR is "Ready for review"
Pre checks:
Validation
16473638: 100-step, eight-GPU EP8 + CP2 run completed.Controlled CP1/CP2 loss parity
16481738: both matched 100-step runs completed on the same eight-GPU allocation.6bf1bf3f0a969ad6a7120e980d0a11c452be89814396c25386962fca6d973e56.0.00946(0.44%); maximum:0.02258(0.92%, step 0); RMSE:0.01039; Pearson r:0.999795.2.95205, CP22.97463; step 99: CP12.10886, CP22.12573.39.02 GiB; CP242.90 GiB.