Skip to content

Keep Dream/LLaDA2 rope shim out of the shared ROPE_INIT_FUNCTIONS - #1810

Merged
jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
YHC66:fix-dream-rope-global-registry
Sep 25, 2026
Merged

jlarson4 merged 2 commits into
TransformerLensOrg:devfrom
YHC66:fix-dream-rope-global-registry

Conversation

@YHC66

@YHC66 YHC66 commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Description

The Dream and LLaDA2-MoE adapters restore transformers v4's "default" rope init (which their remote code looks up) by inserting it into the global transformers.modeling_rope_utils.ROPE_INIT_FUNCTIONS dict.

Starting with transformers 5.17.0, PreTrainedModel._init_weights re-initializes rotary buffers with

rope_init_fn_with_self = {
    "axial": getattr(module, "compute_axial_rope_parameters", None),
    "default": getattr(module, "compute_default_rope_parameters", None),
    **ROPE_INIT_FUNCTIONS,
}

so the global "default" entry now overrides every native model's own compute_default_rope_parameters (5.13–5.16 only consulted ROPE_INIT_FUNCTIONS for non-default rope types). After a Dream or LLaDA2 model has been prepared, building any native model in the same process breaks:

from transformers import LlamaConfig, LlamaForCausalLM
from transformer_lens.model_bridge.supported_architectures.dream import DreamArchitectureAdapter

cfg = LlamaConfig(hidden_size=16, intermediate_size=32, num_hidden_layers=1,
                  num_attention_heads=2, vocab_size=32)
LlamaForCausalLM(cfg)  # ok
DreamArchitectureAdapter(dream_cfg).prepare_loading("Dream-org/Dream-v0-Instruct-7B", {})
LlamaForCausalLM(cfg)  # AttributeError: 'LlamaConfig' object has no attribute 'rope_theta'

Models whose config still carries rope_theta would instead silently get the v4 inv_freq in place of their own rope init. With transformers 5.17 this also shows up as order-dependent failures in tests/unit under xdist (test_gemma2_embed_hook_out_magnitude_matches_sqrt_d_model_scaling, test_config_flag_assignment.py), which pass when run alone.

This PR patches only the remote modeling modules, the way the Ouro adapter already does: a shared restore_default_rope_init(model_name, dotted_ref, rotary_class_name) helper in _remote_code_compat.py force-imports the remote module, rebinds its module-level ROPE_INIT_FUNCTIONS to a copy with "default" restored, and attaches compute_default_rope_parameters to its rotary class for v5's _init_weights. Dream and LLaDA2-MoE use it; _register_default_rope_init is removed. Ouro is left unchanged.

Type of change

  • Bug fix (non-breaking change which fixes an issue)

Checklist:

  • I have commented my code, particularly in hard-to-understand areas
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Testing (Python 3.12, macOS):

  • The Dream/LLaDA2 adapter tests now assert that prepare_loading leaves the shared registry without "default" (they fail before this change on both transformers 5.13 and 5.17). New TestRestoreDefaultRopeInit tests cover the helper with a stand-in remote module, including building a native Llama afterwards.
  • With the real Dream remote code, a tiny DreamModel builds and runs on transformers 5.13 and 5.17, and its inv_freq matches the v4 formula. tests/integration/model_bridge/test_llada2_moe_adapter.py passes on both.
  • tests/unit -m "not slow": transformers 5.13 (lock), 6722 passed, 0 failed. transformers 5.17: 11 failures/errors on dev drop to 5; the remaining 5 (NemotronH cache mocks, vLLM worker extension, test_resolve_state_dict_key_dense_mlp_fallback) are unrelated and also fail on dev.

The Dream and LLaDA2-MoE adapters restored transformers v4's "default"
rope init by inserting it into the global ROPE_INIT_FUNCTIONS dict.
From transformers 5.17, PreTrainedModel._init_weights builds its rope
lookup as {"default": module.compute_default_rope_parameters,
**ROPE_INIT_FUNCTIONS}, so the global entry overrides every native
model's own default rope init. After loading either model, constructing
e.g. a LlamaForCausalLM in the same process raises
AttributeError: 'LlamaConfig' object has no attribute 'rope_theta'.

Patch only the remote modeling modules instead, as the Ouro adapter
already does, via a shared restore_default_rope_init helper.

@jlarson4 jlarson4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for finding and resolving this @YHC66! The 5.17 _init_weights diagnosis is good and moving to the Ouro-style module-local patch is a sound solution. One small comment to resolve before I can merge

return inv_freq, 1.0


def restore_default_rope_init(model_name: str, dotted_ref: str, rotary_class_name: str) -> None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This patches only the modeling copies already imported, and it force-imports the default revision. boot(..., revision=...) loads a different copy afterwards, which now fails with KeyError: 'default'. The old global entry covered that case. Can the caller's revision (from model_kwargs) be forwarded to the force-import here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, thanks. restore_default_rope_init now takes a revision and forwards it to force_import_remote_class, and both Dream and LLaDA2 pass model_kwargs.get("revision") from prepare_loading, so the pinned revision's own modeling copy is imported and patched before loading. Added test_patches_the_requested_revision, where the revision's module copy only appears once that revision is imported (it fails without the forward). Pushed in 83bf0ff.

Side note, not changed here: Dream's other two force-imports (DreamGenerationConfig, DreamAttention) also use the default revision, so with a pinned revision their patches land on the default copy only. That predates this PR; happy to forward the revision there too, either in this PR or a follow-up, whichever you prefer.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@YHC66 A good idea, we should address those additional locations. I am going to merge this so it can be included in the 4.1.0 release that is going out today, if you'd be willing to address those additional changes in a follow up for the next version, that would be greatly appreciated. Thank you!

boot(..., revision=...) imports a separate copy of the remote modeling
file; force-importing only the default revision left that copy without
the 'default' rope entry (KeyError: 'default').
@jlarson4
jlarson4 merged commit 1122461 into TransformerLensOrg:dev Sep 25, 2026
27 checks passed
jlarson4 pushed a commit that referenced this pull request Sep 28, 2026
* Forward the requested revision to every remote-code force-import

prepare_loading force-imports remote modeling classes so their modules
land in sys.modules to patch. Each Hub revision is its own module copy,
and only the default revision was imported, so a pinned
revision=... loaded by from_pretrained afterwards got an unpatched copy.
#1810 fixed this for restore_default_rope_init; do the same for Dream's
DreamGenerationConfig and DreamAttention imports and for the BD3LM,
GIDD, InternLM2, OpenELM, Ouro, Raven and RWKV-7 adapters.

* Forward the requested revision in Baichuan's force-import
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants