Blaž Rolih, Matic Fučka, Filip Wolf, Luka Čehovin Zajc
University of Ljubljana, Faculty of Computer and Information Science
- Data-adaptive training: synthesise changes directly from target-domain feature statistics.
- Strong benchmark results: best results on four of five evaluated datasets, with a 14.1 percentage-point improvement in average F1 over the prior methods compared in the paper.
- Ready to explore: a browser demo, pretrained checkpoints, and training and evaluation code.
Quickstart | Checkpoints | Training | Data | How it works | Citation | Questions
git clone https://github.com/blaz-r/mason_cd.git
cd mason_cd
conda create -n mason_env python=3.12
conda activate mason_env
pip install -r requirements.txtThe default configuration loads DINOv3 ViT-L/16. Request access on its model page and authenticate with an account that has access:
hf auth login
Frozen DINOv3 weights are loaded separately from Meta's repository; they are omitted from MaSoN checkpoints saved by this code. Backbone access is therefore required for local training and evaluation.
python train.py --config configs/mason.yaml \
--data.dataset oscd96 \
--eval_only --ckpt_path blaz-r/mason_oscd96This downloads the OSCD dataset and MaSoN checkpoint, runs benchmark evaluation, and writes metrics to res.csv under ./results. The first run also downloads the backbone.
This command evaluates a dataset, use the browser demo to try an individual image pair:
Choose the checkpoint for your use case:
| Use case | Checkpoint |
|---|---|
| Zero-shot Sentinel-2 RGB inference | MaSoN trained on SSL4EO · Demo |
| Paper benchmark evaluation | Dataset-specific checkpoints below |
F1 values below are percentages. Released checkpoints are individual runs; the paper reports averages across seeds.
| Dataset argument | Hugging Face checkpoint | Released checkpoint F1 | Paper mean F1 |
|---|---|---|---|
oscd96 |
mason_oscd96 | 44.05 | 43.5 |
levir |
mason_levir | 39.37 | 38.9 |
sysu |
mason_sysu | 62.14 | 60.6 |
clcd |
mason_clcd | 56.05 | 54.3 |
gvlm |
mason_gvlm | 58.26 | 55.4 |
Set both --data.dataset and --ckpt_path when switching benchmarks. For example:
python train.py --config configs/mason.yaml \
--data.dataset clcd \
--eval_only --ckpt_path blaz-r/mason_clcdCheckpoints are also available on Google Drive. To evaluate a downloaded checkpoint:
python train.py --config configs/mason.yaml \
--data.dataset clcd \
--eval_only --ckpt_path checkpoints/clcd.ckptTrain and then evaluate on one of the integrated datasets:
python train.py --config configs/mason.yaml --data.dataset oscd96The default configuration trains for 1,000 steps, keeps the encoder frozen, and saves checkpoints under ./checkpoints. Synthetic masks supervise training; real change masks are used for evaluation.
See configs/mason.yaml for the encoder, decoder, noise settings and optimiser. Override configuration values from the command line:
python train.py --config configs/mason.yaml \
--data.dataset clcd --data.batch_size 4 --seed 42List available arguments with python train.py -h.
Integrated datasets (oscd96, levir, sysu, clcd, gvlm) are downloaded from Hugging Face and cached under ./datasets by default.
The default pipeline expects RGB values in [0, 255]. It resizes images to 256 × 256, scales them to [0, 1], and applies ImageNet normalisation with mean (0.485, 0.456, 0.406) and standard deviation (0.229, 0.224, 0.225). Supply unnormalised images to this pipeline to avoid applying normalisation twice.
Image pairs should cover the same area and be spatially aligned. Binary label images use 0 for unchanged and 255 for changed; the loader converts them to [0, 1].
Set --data.use_hf false, choose a dataset name, and point --data.data_path to its parent directory:
python train.py --config configs/mason.yaml \
--data.use_hf false --data.dataset my_dataset --data.data_path ./datasetsThe current dataset loader expects paired images and label files, including during training, although the training objective uses synthetic change masks.
Directory and HDF5 formats
Directory layout (matching PNG filenames in A, B and label):
data_root/
dataset_name/
train/
A/
B/
label/
test/
A/
B/
label/
val/
A/
B/
label/
By default, images are read from disk on demand. Set --data.load_in_mem direct to load the directory dataset into RAM.
For HDF5, set --data.load_in_mem hdf5 and use:
data_root/
dataset_name/
train.h5
test.h5
val.h5
Each split must contain imageA, imageB, label, and img_idx. Image arrays use N × H × W × C; label arrays use N × H × W. HDF5 splits are loaded into RAM.
- Extract features with a frozen encoder.
- During training, add data-adaptive latent perturbations (Gaussian noise) to simulate relevant changes and irrelevant variations, with synthetic masks as supervision.
- Train a decoder to detect changes from feature differences. At inference, compare the features of the actual image pair without adding perturbations.
See the paper and project page for the perturbation analysis, ablations, and qualitative results. Performance depends on the encoder's representations; transferring to a new domain or modality may require adaptation.
If you found this work useful, consider citing our paper and giving this repo a ⭐ 😃
@article{rolih2026mason,
title={Make Some Noise: Unsupervised Remote Sensing Change Detection Using Latent Space Perturbations},
author={Rolih, Blaž and Fučka, Matic and Wolf, Filip and Čehovin Zajc, Luka},
journal={IEEE Transactions on Geoscience and Remote Sensing},
year={2026},
doi={10.1109/TGRS.2026.3730181},
}For issues or questions, please open a GitHub issue or email me.