Scope-aware online editing · Cross-modal generalization · Constant per-edit overhead
Overview | Quick Start | Data | Models | Evaluation | Acknowledgement | Citation
ScopeEdit is a scope-aware online editor for multimodal large language models. Instead of only asking whether an edit succeeds on the original request, ScopeEdit controls where the edited knowledge is allowed to propagate and where it should remain invisible.
ScopeEdit reframes online multimodal editing as Edit-Scoped Generalization: each edit should be absorbed reliably, propagate only to semantically supported image-text variants, and remain inactive on unrelated inputs.
The method implements this principle through three coupled mechanisms:
- Modality-local absorption: an always-active local branch writes the current correction into modality-local edit geometry, supporting reliable absorption while preserving out-of-scope locality.
- Evidence-gated shared propagation: a shared branch is activated only when textual and visual evidence are both directionally aligned and comparably supported, enabling warranted cross-modal transfer while suppressing leakage.
- Scope-separated recursive geometry: local and shared branches use fixed orthogonal low-rank write coordinates and maintain separate history-only preconditioners, updated with Sherman--Morrison recursions for constant per-edit overhead.
Our pilot study shows that reliable edits are not necessarily scope-correct. Among already successful edits, only 62.20% achieve proper generalization; 28.60% under-generalize to in-scope variants, 7.20% over-generalize to out-of-scope inputs, and 2.00% suffer from entangled failures. Layer-wise analysis further shows that cross-modal edit responses concentrate in deeper semantic layers, motivating evidence-gated shared propagation in ScopeEdit.
Install the environment:
python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r easyedit_pip.txtThis repository focuses on two core configurations:
| Config | Backbone | Default Tasks |
|---|---|---|
hparams/TRAINING/MORE/llava15_scopeedit.yaml |
LLaVA-v1.5-7B | E-IC / E-VQA |
hparams/TRAINING/MORE/blip2_scopeedit.yaml |
BLIP-2 OPT-2.7B | E-IC / E-VQA |
This repository uses the multimodal editing data from MMEdit. The required download links are listed below. For the original data description and detailed locality setup, see MMEdit.md.
| Data | Link | Usage |
|---|---|---|
| E-IC | Google Drive | Image captioning editing |
| E-VQA | Google Drive | Visual question answering editing |
| Images | Google Drive | Shared images for E-IC and E-VQA |
Place data and images under:
/root/autodl-tmp/data
├── editing-data
└── <image folders>
If your data or image root is different, update coco_image and rephrase_image in the corresponding YAML config.
The default paths below match the provided YAML configs. You may use different local paths as long as the YAML files are updated accordingly.
llava15_scopeedit.yaml expects a local Transformers-format LLaVA-v1.5-7B checkpoint:
/root/autodl-tmp/models/llava-v1.5-7b
Download:
mkdir -p /root/autodl-tmp/models
huggingface-cli download llava-hf/llava-1.5-7b-hf \
--local-dir /root/autodl-tmp/models/llava-v1.5-7bModel page: https://huggingface.co/llava-hf/llava-1.5-7b-hf
blip2_scopeedit.yaml expects:
/root/autodl-tmp/models/opt-2.7b
/root/blip_models/blip2_pretrained_opt2.7b.pth
/root/blip_models/eva_vit_g.pth
Download OPT-2.7B:
mkdir -p /root/autodl-tmp/models
huggingface-cli download facebook/opt-2.7b \
--local-dir /root/autodl-tmp/models/opt-2.7bModel page: https://huggingface.co/facebook/opt-2.7b
Download BLIP-2 vision-side checkpoints:
mkdir -p /root/blip_models
wget -O /root/blip_models/blip2_pretrained_opt2.7b.pth \
https://storage.googleapis.com/sfr-vision-language-research/LAVIS/models/BLIP2/blip2_pretrained_opt2.7b.pth
wget -O /root/blip_models/eva_vit_g.pth \
https://storage.googleapis.com/sfr-vision-language-research/LAVIS/models/BLIP2/eva_vit_g.pthMiniGPT-4 checkpoints are not part of the main ScopeEdit reproduction path. For the original MMEdit setup, see MMEdit.md.
Before evaluation, verify the local paths in the two core configs.
# hparams/TRAINING/MORE/llava15_scopeedit.yaml
model_name: /root/autodl-tmp/models/llava-v1.5-7b
tokenizer_name: /root/autodl-tmp/models/llava-v1.5-7b
coco_image: /root/autodl-tmp/data
rephrase_image: /root/autodl-tmp/data# hparams/TRAINING/MORE/blip2_scopeedit.yaml
name: /root/autodl-tmp/models/opt-2.7b
tokenizer_name: /root/autodl-tmp/models/opt-2.7b
qformer_checkpoint: /root/blip_models/blip2_pretrained_opt2.7b.pth
state_dict_file: /root/blip_models/eva_vit_g.pth
coco_image: /root/autodl-tmp/data
rephrase_image: /root/autodl-tmp/dataRun ScopeEdit with one of the following commands.
python test_scopeedit_multisteps.py \
hparams/TRAINING/MORE/llava15_scopeedit.yaml \
--train-ds /root/autodl-tmp/data/editing-data/caption/caption_train_edit.json \
--test-ds /root/autodl-tmp/data/editing-data/caption/caption_eval_edit.json \
--warmup-iters 10 \
--n-edits 100 \
--edit-steps 1,10,20,30,100 \
--more-steps 5python test_scopeedit_multisteps.py \
hparams/TRAINING/MORE/blip2_scopeedit.yaml \
--train-ds /root/autodl-tmp/data/editing-data/caption/caption_train_edit.json \
--test-ds /root/autodl-tmp/data/editing-data/caption/caption_eval_edit.json \
--warmup-iters 10 \
--n-edits 100 \
--edit-steps 1,10,20,30,100 \
--more-steps 5E-VQA does not require a separate config. Use one of the core YAML files above and replace the dataset paths. The script automatically selects the dataset class when the path contains vqa.
python test_scopeedit_multisteps.py \
hparams/TRAINING/MORE/llava15_scopeedit.yaml \
--train-ds /root/autodl-tmp/data/editing-data/vqa/vqa_train.json \
--test-ds /root/autodl-tmp/data/editing-data/vqa/vqa_eval.json \
--warmup-iters 10 \
--n-edits 100 \
--edit-steps 1,10,20,30,100 \
--more-steps 5For BLIP-2, replace the config with hparams/TRAINING/MORE/blip2_scopeedit.yaml.
The script prints warmup status, online sequential editing results, and frozen-parameter evaluation results. Default output directories:
/root/autodl-tmp/results/SCOPEEDIT_Eval_LLaVA15
/root/autodl-tmp/results/SCOPEEDIT_Eval_BLIP2_OPT
This implementation is built on top of EasyEdit. We thank the EasyEdit authors and community for providing the model-editing infrastructure.
For multimodal data preparation and the original MMEdit notes, see MMEdit.md. It preserves the MMEdit dataset links, image files, and checkpoint preparation details used for reproduction checks.
@misc{li2026multimodal,
title={Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing},
author={Li, Siyuan and Zhang, Youyuan and Liu, Ruitong and Wang, Junxi and Li, Jing},
year={2026},
note={Manuscript under review}
}
