Repository navigation
Keep Dream/LLaDA2 rope shim out of the shared ROPE_INIT_FUNCTIONS - #1810
Conversation
The Dream and LLaDA2-MoE adapters restored transformers v4's "default"
rope init by inserting it into the global ROPE_INIT_FUNCTIONS dict.
From transformers 5.17, PreTrainedModel._init_weights builds its rope
lookup as {"default": module.compute_default_rope_parameters,
**ROPE_INIT_FUNCTIONS}, so the global entry overrides every native
model's own default rope init. After loading either model, constructing
e.g. a LlamaForCausalLM in the same process raises
AttributeError: 'LlamaConfig' object has no attribute 'rope_theta'.
Patch only the remote modeling modules instead, as the Ouro adapter
already does, via a shared restore_default_rope_init helper.
| return inv_freq, 1.0 | ||
|
|
||
|
|
||
| def restore_default_rope_init(model_name: str, dotted_ref: str, rotary_class_name: str) -> None: |
There was a problem hiding this comment.
This patches only the modeling copies already imported, and it force-imports the default revision. boot(..., revision=...) loads a different copy afterwards, which now fails with KeyError: 'default'. The old global entry covered that case. Can the caller's revision (from model_kwargs) be forwarded to the force-import here?
There was a problem hiding this comment.
Good catch, thanks. restore_default_rope_init now takes a revision and forwards it to force_import_remote_class, and both Dream and LLaDA2 pass model_kwargs.get("revision") from prepare_loading, so the pinned revision's own modeling copy is imported and patched before loading. Added test_patches_the_requested_revision, where the revision's module copy only appears once that revision is imported (it fails without the forward). Pushed in 83bf0ff.
Side note, not changed here: Dream's other two force-imports (DreamGenerationConfig, DreamAttention) also use the default revision, so with a pinned revision their patches land on the default copy only. That predates this PR; happy to forward the revision there too, either in this PR or a follow-up, whichever you prefer.
There was a problem hiding this comment.
@YHC66 A good idea, we should address those additional locations. I am going to merge this so it can be included in the 4.1.0 release that is going out today, if you'd be willing to address those additional changes in a follow up for the next version, that would be greatly appreciated. Thank you!
boot(..., revision=...) imports a separate copy of the remote modeling file; force-importing only the default revision left that copy without the 'default' rope entry (KeyError: 'default').
* Forward the requested revision to every remote-code force-import prepare_loading force-imports remote modeling classes so their modules land in sys.modules to patch. Each Hub revision is its own module copy, and only the default revision was imported, so a pinned revision=... loaded by from_pretrained afterwards got an unpatched copy. #1810 fixed this for restore_default_rope_init; do the same for Dream's DreamGenerationConfig and DreamAttention imports and for the BD3LM, GIDD, InternLM2, OpenELM, Ouro, Raven and RWKV-7 adapters. * Forward the requested revision in Baichuan's force-import
Description
The Dream and LLaDA2-MoE adapters restore transformers v4's
"default"rope init (which their remote code looks up) by inserting it into the globaltransformers.modeling_rope_utils.ROPE_INIT_FUNCTIONSdict.Starting with transformers 5.17.0,
PreTrainedModel._init_weightsre-initializes rotary buffers withso the global
"default"entry now overrides every native model's owncompute_default_rope_parameters(5.13–5.16 only consultedROPE_INIT_FUNCTIONSfor non-default rope types). After a Dream or LLaDA2 model has been prepared, building any native model in the same process breaks:Models whose config still carries
rope_thetawould instead silently get the v4 inv_freq in place of their own rope init. With transformers 5.17 this also shows up as order-dependent failures intests/unitunder xdist (test_gemma2_embed_hook_out_magnitude_matches_sqrt_d_model_scaling,test_config_flag_assignment.py), which pass when run alone.This PR patches only the remote modeling modules, the way the Ouro adapter already does: a shared
restore_default_rope_init(model_name, dotted_ref, rotary_class_name)helper in_remote_code_compat.pyforce-imports the remote module, rebinds its module-levelROPE_INIT_FUNCTIONSto a copy with"default"restored, and attachescompute_default_rope_parametersto its rotary class for v5's_init_weights. Dream and LLaDA2-MoE use it;_register_default_rope_initis removed. Ouro is left unchanged.Type of change
Checklist:
Testing (Python 3.12, macOS):
prepare_loadingleaves the shared registry without"default"(they fail before this change on both transformers 5.13 and 5.17). NewTestRestoreDefaultRopeInittests cover the helper with a stand-in remote module, including building a native Llama afterwards.DreamModelbuilds and runs on transformers 5.13 and 5.17, and itsinv_freqmatches the v4 formula.tests/integration/model_bridge/test_llada2_moe_adapter.pypasses on both.tests/unit -m "not slow": transformers 5.13 (lock), 6722 passed, 0 failed. transformers 5.17: 11 failures/errors ondevdrop to 5; the remaining 5 (NemotronH cache mocks, vLLM worker extension,test_resolve_state_dict_key_dense_mlp_fallback) are unrelated and also fail ondev.