Skip to content

About

Early and mid fusion implementation using the YOLOv8n model from Ultralytics.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

RGBT-Tiny Object Detection: RGB+IR Fusion Study

Aerial object detection on the RGBT-Tiny dataset using YOLOv8n with three modality strategies: RGB baseline, IR baseline, early fusion (4-channel single backbone), and mid-level fusion (dual-backbone with MLP fusion modules). See evaluation_report.md for full results and analysis.


Requirements

pip install ultralytics opencv-python numpy matplotlib

Python 3.10+, PyTorch with CUDA recommended. All scripts were developed with Ultralytics 8.x.


Dataset Setup

The dataset must be arranged as two separate YOLO-format trees (one RGB, one IR) with identical filenames and label files:

Datasets/RGBT-Tiny_YOLO_rgb_ir/
    rgb/
        train/images/   rgb_frame_0001.jpg ...
        train/labels/   rgb_frame_0001.txt ...
        val/images/
        val/labels/
    ir/
        train/images/   rgb_frame_0001.jpg   (same filename as RGB)
        train/labels/   rgb_frame_0001.txt   (shared labels)
        val/images/
        val/labels/

To convert from the original COCO-format RGBT-Tiny dataset:

python Datasets/coco_to_yolo.py

This script reads Datasets/RGBT-Tiny/annotations_coco/ and produces the RGB/IR YOLO directory structure above.

To create day/night validation splits:

python Datasets/create_day_night_split.py

This classifies video sequences by mean RGB brightness (threshold 90.5) and copies the corresponding val images into:

Datasets/val_day_full/        (6986 day val images, RGB)
Datasets/val_night_full/      (5604 night val images, RGB)
Datasets/val_day_ir_full/     (6986 day val images, IR)
Datasets/val_night_ir_full/   (5604 night val images, IR)

The matching dataset YAML files in Datasets/ reference these directories.


Project Structure

final_project/
    README.md
    evaluation_report.md            Full results, analysis, and recommendations
    evaluate_full.py                General evaluation utility (size-based metrics)
    generate_result_tables.py       Render evaluation_report.txt files as PNG tables

    baselines/
        train_yolov8.py             Train RGB or IR single-modality baseline
        evaluate_yolov8.py          Evaluate a trained baseline model
        inference_yolov8.py         Single-image inference for baselines

    fusion/                         Early fusion (4-channel single backbone)
        train_fusion_ultralytics.py Training script (Ultralytics pipeline)
        evaluate_fusion.py          Evaluation script
        fusion_dataset.py           Custom 4-channel dataloader
        fusion_dataset.yaml         Dataset config (points to rgb+ir roots)
        yolov8n-fusion.yaml         Modified YOLOv8n arch (ch=4)
        check_ir_channels.py        Diagnostic: verify IR image channel structure
        __init__.py

    mid_fusion/                     Mid-level fusion (dual-backbone)
        train_mid_fusion_ultralytics.py  Training script (Ultralytics pipeline)
        eval_mid_fusion.py               Evaluation script
        mid_fusion_model.py              Dual-backbone model architecture
        mid_fusion_dataset.py            Custom 6-channel dataloader
        fusion_modules.py                Fusion block implementations (MLP, attention)
        inference_mid_fusion.py          Single-image inference
        mid_fusion_dataset.yaml          Dataset config
        yolov8_mid_fusion.yaml           Architecture reference

    Datasets/
        RGBT-Tiny_YOLO_rgb_ir/      Main dataset (rgb/ and ir/ subtrees)
        RGBT-Tiny_YOLO/             Original mixed YOLO format (not used for training)
        RGBT-Tiny/                  Original COCO-format annotations
        coco_to_yolo.py             Dataset conversion script
        create_day_night_split.py   Day/night split creation
        val_day_full_dataset.yaml   RGB day split config (used for evaluation)
        val_night_full_dataset.yaml RGB night split config
        val_day_ir_full_dataset.yaml    IR day split config
        val_night_ir_full_dataset.yaml  IR night split config
        val_day_full_fusion_dataset.yaml    Fusion day split config
        val_night_full_fusion_dataset.yaml  Fusion night split config
        (+ smaller 50-image subsets for quick testing)

    runs/                           Training outputs (weights, logs, plots)
        yolov8n_rgb_scratch_v2/     RGB baseline weights
        yolov8n_ir_scratch/         IR baseline weights
        yolov8n_fusion_ultralytics3/    Early fusion (original training)
        yolov8n_fusion_ultralytics4/    Early fusion (corrected training)
        mid_fusion/
            mlp_20260208_172237_phase1/ Mid-fusion phase 1 checkpoint
            mlp_20260210_113817_phase2/ Mid-fusion final evaluated model

    evaluation_results/             Structured evaluation outputs (metrics.json, plots)
    results_tables/                 PNG table images generated from evaluation results
    test_images/                    Sample RGB and IR images for quick testing
    archive/                        Superseded runs, scripts, and evaluation results

Training

1. RGB Baseline

python baselines/train_yolov8.py \
    --data Datasets/RGBT-Tiny_YOLO_rgb_ir/rgb_dataset.yaml \
    --epochs 250 \
    --batch 16 \
    --name yolov8n_rgb_scratch_v2

Key arguments:

Argument Default Description
--data rgb_dataset.yaml Dataset YAML path
--epochs 250 Training epochs
--batch 16 Batch size
--device 0 GPU index or cpu
--name yolov8n_rgb Run name under runs/
--pretrained (flag) Start from COCO pretrained YOLOv8n

2. IR Baseline

Identical to RGB -- just point to the IR dataset:

python baselines/train_yolov8.py \
    --data Datasets/RGBT-Tiny_YOLO_rgb_ir/ir_dataset.yaml \
    --epochs 250 \
    --batch 16 \
    --name yolov8n_ir_scratch

3. Early Fusion (4-channel RGB+IR)

The early fusion model concatenates RGB (3-ch) and IR (1-ch grayscale) into a 4-channel input. The first conv layer is initialized with zero weights for the IR channel so the model starts as a near-pure RGB model and learns IR contribution from scratch.

Augmentation note: The ChannelAwareHSV transform applies full HSV jitter only to the RGB channels (0:2) and brightness-only variation to the IR channel (3+), avoiding undefined behavior from OpenCV's BGR2HSV on non-3-channel images.

python fusion/train_fusion_ultralytics.py \
    --data fusion/fusion_dataset.yaml \
    --ir-root Datasets/RGBT-Tiny_YOLO_rgb_ir/ir \
    --epochs 250 \
    --patience 50 \
    --batch 16 \
    --name yolov8n_fusion_new

Key arguments:

Argument Default Description
--data fusion/fusion_dataset.yaml Dataset YAML (RGB paths)
--ir-root Datasets/.../ir Root of IR image tree
--epochs 250 Max training epochs
--patience 50 Early stopping patience
--batch 16 Batch size
--yaml fusion/yolov8n-fusion.yaml Model architecture YAML
--lr0 0.001 Initial learning rate
--name yolov8n_fusion_ultralytics Run name under runs/

4. Mid-Level Fusion (dual-backbone with MLP fusion)

The mid-fusion model uses two independent YOLOv8n backbones (one for RGB, one for IR), fuses feature maps at P3/P4/P5, then passes through a shared PAFPN neck and detection head. Training uses two phases: Phase 1 freezes the backbones and trains only the fusion head; Phase 2 fine-tunes all layers.

Both backbones are initialized from their respective trained baselines (RGB and IR). These baseline weights must exist before running mid-fusion training.

python mid_fusion/train_mid_fusion_ultralytics.py \
    --fusion-type mlp \
    --phase1-epochs 30 \
    --phase2-epochs 200 \
    --patience 50

Key arguments:

Argument Default Description
--fusion-type mlp Fusion module type: mlp, channel_attention, concat
--rgb-weights runs/yolov8n_rgb_scratch_v2/weights/best.pt RGB backbone init
--ir-weights runs/yolov8n_ir_scratch/weights/best.pt IR backbone init
--data fusion/fusion_dataset.yaml Dataset YAML (6-ch RGB+IR)
--ir-root Datasets/.../ir Root of IR image tree
--phase1-epochs 30 Phase 1 epochs (frozen backbones)
--phase1-lr 0.01 Phase 1 learning rate
--phase2-epochs 200 Phase 2 epochs (full fine-tuning)
--phase2-lr 0.001 Phase 2 learning rate
--patience 50 Early stopping patience (applies to phase 2)
--batch 8 Batch size
--device 0 GPU index or cpu
--name mlp_{timestamp} Run name under runs/mid_fusion/

Phase 1 saves to runs/mid_fusion/{name}_phase1/ and Phase 2 to runs/mid_fusion/{name}_phase2/. The evaluated model is the Phase 2 best.pt.

To skip Phase 1 and load a previously trained Phase 1 checkpoint:

python mid_fusion/train_mid_fusion_ultralytics.py \
    --fusion-type mlp \
    --skip-phase1 \
    --phase1-weights runs/mid_fusion/mlp_20260208_172237_phase1/weights/best.pt \
    --phase2-epochs 200 \
    --phase2-lr 0.001 \
    --patience 50

Evaluation

All evaluations use conf=0.001 (full PR curve sweep for correct mAP) and iou=0.6. Each script saves structured results (metrics.json, confusion matrices, PR curves) to evaluation_results/{model}_{modality}_{timestamp}/.

Baselines

# Full validation set
python baselines/evaluate_yolov8.py \
    --weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
    --data Datasets/RGBT-Tiny_YOLO_rgb_ir/rgb_dataset.yaml

# Day split
python baselines/evaluate_yolov8.py \
    --weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
    --data Datasets/val_day_full_dataset.yaml

# Night split
python baselines/evaluate_yolov8.py \
    --weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
    --data Datasets/val_night_full_dataset.yaml

# IR baseline (use IR dataset YAMLs)
python baselines/evaluate_yolov8.py \
    --weights runs/yolov8n_ir_scratch/weights/best.pt \
    --data Datasets/RGBT-Tiny_YOLO_rgb_ir/ir_dataset.yaml

Early Fusion

# Full validation set
python fusion/evaluate_fusion.py \
    --weights runs/yolov8n_fusion_ultralytics4/weights/best.pt

# Day split
python fusion/evaluate_fusion.py \
    --weights runs/yolov8n_fusion_ultralytics4/weights/best.pt \
    --data Datasets/val_day_full_fusion_dataset.yaml \
    --ir-root Datasets/val_day_ir_full

# Night split
python fusion/evaluate_fusion.py \
    --weights runs/yolov8n_fusion_ultralytics4/weights/best.pt \
    --data Datasets/val_night_full_fusion_dataset.yaml \
    --ir-root Datasets/val_night_ir_full

Key arguments for evaluate_fusion.py:

Argument Default Description
--weights (required) Path to best.pt
--data fusion/fusion_dataset.yaml Dataset YAML
--ir-root Datasets/.../ir IR images root (overrides YAML default)
--conf 0.001 Confidence threshold
--iou 0.6 NMS IoU threshold
--batch 16 Batch size

Mid-Level Fusion

# Full validation set
python mid_fusion/eval_mid_fusion.py \
    --weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
    --fusion-type mlp \
    --data fusion/fusion_dataset.yaml \
    --ir-root Datasets/RGBT-Tiny_YOLO_rgb_ir/ir \
    --modality full_val

# Day split
python mid_fusion/eval_mid_fusion.py \
    --weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
    --fusion-type mlp \
    --data Datasets/val_day_full_fusion_dataset.yaml \
    --ir-root Datasets/val_day_ir_full \
    --modality val_day_full

# Night split
python mid_fusion/eval_mid_fusion.py \
    --weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
    --fusion-type mlp \
    --data Datasets/val_night_full_fusion_dataset.yaml \
    --ir-root Datasets/val_night_ir_full \
    --modality val_night_full

Key arguments for eval_mid_fusion.py:

Argument Default Description
--weights (required) Path to best.pt
--fusion-type mlp Must match training fusion type
--data fusion/fusion_dataset.yaml Dataset YAML
--ir-root Datasets/.../ir IR images root
--modality full_val Label added to output directory name
--conf 0.001 Confidence threshold
--iou 0.6 NMS IoU threshold
--batch 8 Batch size

Single-Image Inference

# Baseline
python baselines/inference_yolov8.py \
    --weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
    --source test_images/rgb/sample.jpg

# Mid-fusion
python mid_fusion/inference_mid_fusion.py \
    --weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
    --fusion-type mlp \
    --rgb test_images/rgb/sample.jpg \
    --ir test_images/ir/sample.jpg

Results Summary

Full results, per-class breakdowns, day/night analysis, and improvement recommendations are in evaluation_report.md. Key mAP@0.5 numbers:

Model Full Val Day Night
RGB Baseline 0.406 0.497 0.236
IR Baseline 0.412 0.448 0.372
Early Fusion (corrected) 0.370 0.485 0.200
Mid Fusion (MLP) 0.390 0.481 0.239

The IR baseline outperforms all fusion models overall. The primary limitation of both fusion architectures is fixed-weight combination at inference time, which prevents the model from suppressing degraded RGB channels at night. See Section 7 and 8 of evaluation_report.md for architectural analysis and recommendations.

About

Early and mid fusion implementation using the YOLOv8n model from Ultralytics.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages