Skip to content

Add Muse Glimmer architecture adapter - #1822

Merged
jlarson4 merged 3 commits into
TransformerLensOrg:devfrom
almogtavor:add-muse-glimmer-adapter
Sep 29, 2026
Merged

jlarson4 merged 3 commits into
TransformerLensOrg:devfrom
almogtavor:add-muse-glimmer-adapter

Conversation

@almogtavor

Copy link
Copy Markdown
Contributor

Description

Adds a TransformerBridge adapter for Muse Glimmer (MuseGlimmerForConditionalGeneration). The text path is bridged and the vision tower and projector are opaque, as in the Gemma 3 / Llama 4 multimodal adapters. Bridge logits and zero-ablations match HF exactly on a tiny random config and on Muse-Glimmer-30B (bf16). The model needs transformers >= 5.15, so the new tests skip under the current lock (5.13).

Type of change

  • New feature (non-breaking change which adds functionality)

Checklist:

  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Bridge MuseGlimmerForConditionalGeneration (text path through model.language_model and lm_head, vision tower and projector as opaque components), register it in the factory, model registry, multimodal list and HF model_type map, and pass output_multiplier through to the bridge config. The unit test compares bridge logits and attention patterns against HF on a tiny random config and skips when transformers lacks muse_glimmer.

@jlarson4 jlarson4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for adding this adapter @almogtavor! It is well built and the parity checks are solid. A couple small comments below before I can merge, if you'd be willing to put in a quick commit to wrap them up.

"unembed": UnembeddingBridge(name="lm_head", config=self.cfg),
}

def apply_output_logits_transform(self, logits: torch.Tensor) -> torch.Tensor:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bridge forward returns HF's own logits, so the parity test still passes with the multiplier or the softcap deleted from this method. Only JacobianLens calls it. The tiny model's logits are also too small for the cap to change anything at 1e-4. The Granite and Falcon-H1 cases in test_output_logits_contract.py cover this kind of override, and they run under the locked transformers too. Could you add a Muse Glimmer case there?


self.component_mapping = {
"vision_encoder": GeneralizedComponent(name="model.vision_tower"),
"vision_projector": GeneralizedComponent(name="model.vision_projection"),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

HF runs vision_adapter before this Linear and perception_emb_norm after it. So vision_projector.hook_out is an intermediate value, rescaled before it reaches the text stream. In the other VLM adapters this hook is the embedding that actually gets merged in. gemma3_multimodal.py is a good reference: its projector's norm runs inside the wrapped module, so hook_out is exactly what HF scatters into the text embeddings. Can the mapping expose that point here too?

The tiny Muse Glimmer model's logits are too small for the tanh cap to change
anything, so the forward-parity test cannot catch a missing post-unembedding
transform. Add a contract case alongside Granite and Falcon-H1:

- nested text_config: final_logit_softcapping and output_multiplier both reach
  the bridge config;
- apply_output_logits_transform applies the multiplier before the softcap
  (20 * tanh(0.5 * logits / 20)), with logits large enough that dropping the
  factor fails the comparison.
HF runs vision_adapter -> vision_projection -> perception_emb_norm and scatters
only the last value into the text embeddings, so the vision_projector bridge now
wraps perception_emb_norm instead of the raw vision_projection Linear. Its
hook_out is then the embedding HF actually merges, matching the Gemma 3 and LLaVA
adapters; hook_in is the projection output.

Add a structural assertion to the adapter test.
@almogtavor

Copy link
Copy Markdown
Contributor Author

Both addressed in the follow-up commits: test_output_logits_contract.py now has a Muse Glimmer case pinning 20 * tanh(0.5 * logits / 20). It fails if the multiplier is dropped from the method. vision_projector now wraps perception_emb_norm, so hook_out is the embedding HF actually merges (as in the Gemma 3 adapter).

@jlarson4

Copy link
Copy Markdown
Collaborator

Approved @almogtavor, thank you for the updates! merging now

@jlarson4
jlarson4 merged commit 687e3f5 into TransformerLensOrg:dev Sep 29, 2026
27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants