Add Muse Glimmer architecture adapter - #1822
Conversation
Bridge MuseGlimmerForConditionalGeneration (text path through model.language_model and lm_head, vision tower and projector as opaque components), register it in the factory, model registry, multimodal list and HF model_type map, and pass output_multiplier through to the bridge config. The unit test compares bridge logits and attention patterns against HF on a tiny random config and skips when transformers lacks muse_glimmer.
jlarson4
left a comment
There was a problem hiding this comment.
Thanks for adding this adapter @almogtavor! It is well built and the parity checks are solid. A couple small comments below before I can merge, if you'd be willing to put in a quick commit to wrap them up.
| "unembed": UnembeddingBridge(name="lm_head", config=self.cfg), | ||
| } | ||
|
|
||
| def apply_output_logits_transform(self, logits: torch.Tensor) -> torch.Tensor: |
There was a problem hiding this comment.
Bridge forward returns HF's own logits, so the parity test still passes with the multiplier or the softcap deleted from this method. Only JacobianLens calls it. The tiny model's logits are also too small for the cap to change anything at 1e-4. The Granite and Falcon-H1 cases in test_output_logits_contract.py cover this kind of override, and they run under the locked transformers too. Could you add a Muse Glimmer case there?
|
|
||
| self.component_mapping = { | ||
| "vision_encoder": GeneralizedComponent(name="model.vision_tower"), | ||
| "vision_projector": GeneralizedComponent(name="model.vision_projection"), |
There was a problem hiding this comment.
HF runs vision_adapter before this Linear and perception_emb_norm after it. So vision_projector.hook_out is an intermediate value, rescaled before it reaches the text stream. In the other VLM adapters this hook is the embedding that actually gets merged in. gemma3_multimodal.py is a good reference: its projector's norm runs inside the wrapped module, so hook_out is exactly what HF scatters into the text embeddings. Can the mapping expose that point here too?
The tiny Muse Glimmer model's logits are too small for the tanh cap to change anything, so the forward-parity test cannot catch a missing post-unembedding transform. Add a contract case alongside Granite and Falcon-H1: - nested text_config: final_logit_softcapping and output_multiplier both reach the bridge config; - apply_output_logits_transform applies the multiplier before the softcap (20 * tanh(0.5 * logits / 20)), with logits large enough that dropping the factor fails the comparison.
HF runs vision_adapter -> vision_projection -> perception_emb_norm and scatters only the last value into the text embeddings, so the vision_projector bridge now wraps perception_emb_norm instead of the raw vision_projection Linear. Its hook_out is then the embedding HF actually merges, matching the Gemma 3 and LLaVA adapters; hook_in is the projection output. Add a structural assertion to the adapter test.
|
Both addressed in the follow-up commits: |
|
Approved @almogtavor, thank you for the updates! merging now |
Description
Adds a TransformerBridge adapter for Muse Glimmer (
MuseGlimmerForConditionalGeneration). The text path is bridged and the vision tower and projector are opaque, as in the Gemma 3 / Llama 4 multimodal adapters. Bridge logits and zero-ablations match HF exactly on a tiny random config and on Muse-Glimmer-30B (bf16). The model needs transformers >= 5.15, so the new tests skip under the current lock (5.13).Type of change
Checklist: