Skip to content
lab-klcPublic

About

Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing.

Resources

Stars

5 stars

Watchers

0 watching

Forks

Repository files navigation

ScopeEdit

Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing

Scope-aware online editing · Cross-modal generalization · Constant per-edit overhead

Python PyTorch Transformers Task License

Overview | Quick Start | Data | Models | Evaluation | Acknowledgement | Citation

ScopeEdit is a scope-aware online editor for multimodal large language models. Instead of only asking whether an edit succeeds on the original request, ScopeEdit controls where the edited knowledge is allowed to propagate and where it should remain invisible.

ScopeEdit motivation

Overview

ScopeEdit reframes online multimodal editing as Edit-Scoped Generalization: each edit should be absorbed reliably, propagate only to semantically supported image-text variants, and remain inactive on unrelated inputs.

The method implements this principle through three coupled mechanisms:

  • Modality-local absorption: an always-active local branch writes the current correction into modality-local edit geometry, supporting reliable absorption while preserving out-of-scope locality.
  • Evidence-gated shared propagation: a shared branch is activated only when textual and visual evidence are both directionally aligned and comparably supported, enabling warranted cross-modal transfer while suppressing leakage.
  • Scope-separated recursive geometry: local and shared branches use fixed orthogonal low-rank write coordinates and maintain separate history-only preconditioners, updated with Sherman--Morrison recursions for constant per-edit overhead.

Pilot Observation

Pilot analysis

Our pilot study shows that reliable edits are not necessarily scope-correct. Among already successful edits, only 62.20% achieve proper generalization; 28.60% under-generalize to in-scope variants, 7.20% over-generalize to out-of-scope inputs, and 2.00% suffer from entangled failures. Layer-wise analysis further shows that cross-modal edit responses concentrate in deeper semantic layers, motivating evidence-gated shared propagation in ScopeEdit.

Quick Start

Install the environment:

python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r easyedit_pip.txt

This repository focuses on two core configurations:

Config Backbone Default Tasks
hparams/TRAINING/MORE/llava15_scopeedit.yaml LLaVA-v1.5-7B E-IC / E-VQA
hparams/TRAINING/MORE/blip2_scopeedit.yaml BLIP-2 OPT-2.7B E-IC / E-VQA

Data Preparation

This repository uses the multimodal editing data from MMEdit. The required download links are listed below. For the original data description and detailed locality setup, see MMEdit.md.

Data Link Usage
E-IC Google Drive Image captioning editing
E-VQA Google Drive Visual question answering editing
Images Google Drive Shared images for E-IC and E-VQA

Place data and images under:

/root/autodl-tmp/data
├── editing-data
└── <image folders>

If your data or image root is different, update coco_image and rephrase_image in the corresponding YAML config.

Model Preparation

The default paths below match the provided YAML configs. You may use different local paths as long as the YAML files are updated accordingly.

LLaVA-v1.5-7B

llava15_scopeedit.yaml expects a local Transformers-format LLaVA-v1.5-7B checkpoint:

/root/autodl-tmp/models/llava-v1.5-7b

Download:

mkdir -p /root/autodl-tmp/models
huggingface-cli download llava-hf/llava-1.5-7b-hf \
  --local-dir /root/autodl-tmp/models/llava-v1.5-7b

Model page: https://huggingface.co/llava-hf/llava-1.5-7b-hf

BLIP-2 OPT-2.7B

blip2_scopeedit.yaml expects:

/root/autodl-tmp/models/opt-2.7b
/root/blip_models/blip2_pretrained_opt2.7b.pth
/root/blip_models/eva_vit_g.pth

Download OPT-2.7B:

mkdir -p /root/autodl-tmp/models
huggingface-cli download facebook/opt-2.7b \
  --local-dir /root/autodl-tmp/models/opt-2.7b

Model page: https://huggingface.co/facebook/opt-2.7b

Download BLIP-2 vision-side checkpoints:

mkdir -p /root/blip_models
wget -O /root/blip_models/blip2_pretrained_opt2.7b.pth \
  https://storage.googleapis.com/sfr-vision-language-research/LAVIS/models/BLIP2/blip2_pretrained_opt2.7b.pth
wget -O /root/blip_models/eva_vit_g.pth \
  https://storage.googleapis.com/sfr-vision-language-research/LAVIS/models/BLIP2/eva_vit_g.pth

MiniGPT-4 checkpoints are not part of the main ScopeEdit reproduction path. For the original MMEdit setup, see MMEdit.md.

Path Check

Before evaluation, verify the local paths in the two core configs.

# hparams/TRAINING/MORE/llava15_scopeedit.yaml
model_name: /root/autodl-tmp/models/llava-v1.5-7b
tokenizer_name: /root/autodl-tmp/models/llava-v1.5-7b
coco_image: /root/autodl-tmp/data
rephrase_image: /root/autodl-tmp/data
# hparams/TRAINING/MORE/blip2_scopeedit.yaml
name: /root/autodl-tmp/models/opt-2.7b
tokenizer_name: /root/autodl-tmp/models/opt-2.7b
qformer_checkpoint: /root/blip_models/blip2_pretrained_opt2.7b.pth
state_dict_file: /root/blip_models/eva_vit_g.pth
coco_image: /root/autodl-tmp/data
rephrase_image: /root/autodl-tmp/data

Evaluation

Run ScopeEdit with one of the following commands.

LLaVA-v1.5 on E-IC

python test_scopeedit_multisteps.py \
  hparams/TRAINING/MORE/llava15_scopeedit.yaml \
  --train-ds /root/autodl-tmp/data/editing-data/caption/caption_train_edit.json \
  --test-ds /root/autodl-tmp/data/editing-data/caption/caption_eval_edit.json \
  --warmup-iters 10 \
  --n-edits 100 \
  --edit-steps 1,10,20,30,100 \
  --more-steps 5

BLIP-2 OPT on E-IC

python test_scopeedit_multisteps.py \
  hparams/TRAINING/MORE/blip2_scopeedit.yaml \
  --train-ds /root/autodl-tmp/data/editing-data/caption/caption_train_edit.json \
  --test-ds /root/autodl-tmp/data/editing-data/caption/caption_eval_edit.json \
  --warmup-iters 10 \
  --n-edits 100 \
  --edit-steps 1,10,20,30,100 \
  --more-steps 5

E-VQA

E-VQA does not require a separate config. Use one of the core YAML files above and replace the dataset paths. The script automatically selects the dataset class when the path contains vqa.

python test_scopeedit_multisteps.py \
  hparams/TRAINING/MORE/llava15_scopeedit.yaml \
  --train-ds /root/autodl-tmp/data/editing-data/vqa/vqa_train.json \
  --test-ds /root/autodl-tmp/data/editing-data/vqa/vqa_eval.json \
  --warmup-iters 10 \
  --n-edits 100 \
  --edit-steps 1,10,20,30,100 \
  --more-steps 5

For BLIP-2, replace the config with hparams/TRAINING/MORE/blip2_scopeedit.yaml.

Output

The script prints warmup status, online sequential editing results, and frozen-parameter evaluation results. Default output directories:

/root/autodl-tmp/results/SCOPEEDIT_Eval_LLaVA15
/root/autodl-tmp/results/SCOPEEDIT_Eval_BLIP2_OPT

Acknowledgement

This implementation is built on top of EasyEdit. We thank the EasyEdit authors and community for providing the model-editing infrastructure.

For multimodal data preparation and the original MMEdit notes, see MMEdit.md. It preserves the MMEdit dataset links, image files, and checkpoint preparation details used for reproduction checks.

Citation

@misc{li2026multimodal,
  title={Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing},
  author={Li, Siyuan and Zhang, Youyuan and Liu, Ruitong and Wang, Junxi and Li, Jing},
  year={2026},
  note={Manuscript under review}
}

About

Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing.

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages