Aerial object detection on the RGBT-Tiny dataset using YOLOv8n with three modality
strategies: RGB baseline, IR baseline, early fusion (4-channel single backbone), and
mid-level fusion (dual-backbone with MLP fusion modules). See evaluation_report.md
for full results and analysis.
pip install ultralytics opencv-python numpy matplotlib
Python 3.10+, PyTorch with CUDA recommended. All scripts were developed with Ultralytics 8.x.
The dataset must be arranged as two separate YOLO-format trees (one RGB, one IR) with identical filenames and label files:
Datasets/RGBT-Tiny_YOLO_rgb_ir/
rgb/
train/images/ rgb_frame_0001.jpg ...
train/labels/ rgb_frame_0001.txt ...
val/images/
val/labels/
ir/
train/images/ rgb_frame_0001.jpg (same filename as RGB)
train/labels/ rgb_frame_0001.txt (shared labels)
val/images/
val/labels/
To convert from the original COCO-format RGBT-Tiny dataset:
python Datasets/coco_to_yolo.py
This script reads Datasets/RGBT-Tiny/annotations_coco/ and produces the RGB/IR YOLO
directory structure above.
To create day/night validation splits:
python Datasets/create_day_night_split.py
This classifies video sequences by mean RGB brightness (threshold 90.5) and copies the corresponding val images into:
Datasets/val_day_full/ (6986 day val images, RGB)
Datasets/val_night_full/ (5604 night val images, RGB)
Datasets/val_day_ir_full/ (6986 day val images, IR)
Datasets/val_night_ir_full/ (5604 night val images, IR)
The matching dataset YAML files in Datasets/ reference these directories.
final_project/
README.md
evaluation_report.md Full results, analysis, and recommendations
evaluate_full.py General evaluation utility (size-based metrics)
generate_result_tables.py Render evaluation_report.txt files as PNG tables
baselines/
train_yolov8.py Train RGB or IR single-modality baseline
evaluate_yolov8.py Evaluate a trained baseline model
inference_yolov8.py Single-image inference for baselines
fusion/ Early fusion (4-channel single backbone)
train_fusion_ultralytics.py Training script (Ultralytics pipeline)
evaluate_fusion.py Evaluation script
fusion_dataset.py Custom 4-channel dataloader
fusion_dataset.yaml Dataset config (points to rgb+ir roots)
yolov8n-fusion.yaml Modified YOLOv8n arch (ch=4)
check_ir_channels.py Diagnostic: verify IR image channel structure
__init__.py
mid_fusion/ Mid-level fusion (dual-backbone)
train_mid_fusion_ultralytics.py Training script (Ultralytics pipeline)
eval_mid_fusion.py Evaluation script
mid_fusion_model.py Dual-backbone model architecture
mid_fusion_dataset.py Custom 6-channel dataloader
fusion_modules.py Fusion block implementations (MLP, attention)
inference_mid_fusion.py Single-image inference
mid_fusion_dataset.yaml Dataset config
yolov8_mid_fusion.yaml Architecture reference
Datasets/
RGBT-Tiny_YOLO_rgb_ir/ Main dataset (rgb/ and ir/ subtrees)
RGBT-Tiny_YOLO/ Original mixed YOLO format (not used for training)
RGBT-Tiny/ Original COCO-format annotations
coco_to_yolo.py Dataset conversion script
create_day_night_split.py Day/night split creation
val_day_full_dataset.yaml RGB day split config (used for evaluation)
val_night_full_dataset.yaml RGB night split config
val_day_ir_full_dataset.yaml IR day split config
val_night_ir_full_dataset.yaml IR night split config
val_day_full_fusion_dataset.yaml Fusion day split config
val_night_full_fusion_dataset.yaml Fusion night split config
(+ smaller 50-image subsets for quick testing)
runs/ Training outputs (weights, logs, plots)
yolov8n_rgb_scratch_v2/ RGB baseline weights
yolov8n_ir_scratch/ IR baseline weights
yolov8n_fusion_ultralytics3/ Early fusion (original training)
yolov8n_fusion_ultralytics4/ Early fusion (corrected training)
mid_fusion/
mlp_20260208_172237_phase1/ Mid-fusion phase 1 checkpoint
mlp_20260210_113817_phase2/ Mid-fusion final evaluated model
evaluation_results/ Structured evaluation outputs (metrics.json, plots)
results_tables/ PNG table images generated from evaluation results
test_images/ Sample RGB and IR images for quick testing
archive/ Superseded runs, scripts, and evaluation results
python baselines/train_yolov8.py \
--data Datasets/RGBT-Tiny_YOLO_rgb_ir/rgb_dataset.yaml \
--epochs 250 \
--batch 16 \
--name yolov8n_rgb_scratch_v2
Key arguments:
| Argument | Default | Description |
|---|---|---|
--data |
rgb_dataset.yaml | Dataset YAML path |
--epochs |
250 | Training epochs |
--batch |
16 | Batch size |
--device |
0 | GPU index or cpu |
--name |
yolov8n_rgb | Run name under runs/ |
--pretrained |
(flag) | Start from COCO pretrained YOLOv8n |
Identical to RGB -- just point to the IR dataset:
python baselines/train_yolov8.py \
--data Datasets/RGBT-Tiny_YOLO_rgb_ir/ir_dataset.yaml \
--epochs 250 \
--batch 16 \
--name yolov8n_ir_scratch
The early fusion model concatenates RGB (3-ch) and IR (1-ch grayscale) into a 4-channel input. The first conv layer is initialized with zero weights for the IR channel so the model starts as a near-pure RGB model and learns IR contribution from scratch.
Augmentation note: The ChannelAwareHSV transform applies full HSV jitter only
to the RGB channels (0:2) and brightness-only variation to the IR channel (3+),
avoiding undefined behavior from OpenCV's BGR2HSV on non-3-channel images.
python fusion/train_fusion_ultralytics.py \
--data fusion/fusion_dataset.yaml \
--ir-root Datasets/RGBT-Tiny_YOLO_rgb_ir/ir \
--epochs 250 \
--patience 50 \
--batch 16 \
--name yolov8n_fusion_new
Key arguments:
| Argument | Default | Description |
|---|---|---|
--data |
fusion/fusion_dataset.yaml | Dataset YAML (RGB paths) |
--ir-root |
Datasets/.../ir | Root of IR image tree |
--epochs |
250 | Max training epochs |
--patience |
50 | Early stopping patience |
--batch |
16 | Batch size |
--yaml |
fusion/yolov8n-fusion.yaml | Model architecture YAML |
--lr0 |
0.001 | Initial learning rate |
--name |
yolov8n_fusion_ultralytics | Run name under runs/ |
The mid-fusion model uses two independent YOLOv8n backbones (one for RGB, one for IR), fuses feature maps at P3/P4/P5, then passes through a shared PAFPN neck and detection head. Training uses two phases: Phase 1 freezes the backbones and trains only the fusion head; Phase 2 fine-tunes all layers.
Both backbones are initialized from their respective trained baselines (RGB and IR). These baseline weights must exist before running mid-fusion training.
python mid_fusion/train_mid_fusion_ultralytics.py \
--fusion-type mlp \
--phase1-epochs 30 \
--phase2-epochs 200 \
--patience 50
Key arguments:
| Argument | Default | Description |
|---|---|---|
--fusion-type |
mlp | Fusion module type: mlp, channel_attention, concat |
--rgb-weights |
runs/yolov8n_rgb_scratch_v2/weights/best.pt | RGB backbone init |
--ir-weights |
runs/yolov8n_ir_scratch/weights/best.pt | IR backbone init |
--data |
fusion/fusion_dataset.yaml | Dataset YAML (6-ch RGB+IR) |
--ir-root |
Datasets/.../ir | Root of IR image tree |
--phase1-epochs |
30 | Phase 1 epochs (frozen backbones) |
--phase1-lr |
0.01 | Phase 1 learning rate |
--phase2-epochs |
200 | Phase 2 epochs (full fine-tuning) |
--phase2-lr |
0.001 | Phase 2 learning rate |
--patience |
50 | Early stopping patience (applies to phase 2) |
--batch |
8 | Batch size |
--device |
0 | GPU index or cpu |
--name |
mlp_{timestamp} | Run name under runs/mid_fusion/ |
Phase 1 saves to runs/mid_fusion/{name}_phase1/ and Phase 2 to
runs/mid_fusion/{name}_phase2/. The evaluated model is the Phase 2 best.pt.
To skip Phase 1 and load a previously trained Phase 1 checkpoint:
python mid_fusion/train_mid_fusion_ultralytics.py \
--fusion-type mlp \
--skip-phase1 \
--phase1-weights runs/mid_fusion/mlp_20260208_172237_phase1/weights/best.pt \
--phase2-epochs 200 \
--phase2-lr 0.001 \
--patience 50
All evaluations use conf=0.001 (full PR curve sweep for correct mAP) and iou=0.6.
Each script saves structured results (metrics.json, confusion matrices, PR curves) to
evaluation_results/{model}_{modality}_{timestamp}/.
# Full validation set
python baselines/evaluate_yolov8.py \
--weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
--data Datasets/RGBT-Tiny_YOLO_rgb_ir/rgb_dataset.yaml
# Day split
python baselines/evaluate_yolov8.py \
--weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
--data Datasets/val_day_full_dataset.yaml
# Night split
python baselines/evaluate_yolov8.py \
--weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
--data Datasets/val_night_full_dataset.yaml
# IR baseline (use IR dataset YAMLs)
python baselines/evaluate_yolov8.py \
--weights runs/yolov8n_ir_scratch/weights/best.pt \
--data Datasets/RGBT-Tiny_YOLO_rgb_ir/ir_dataset.yaml
# Full validation set
python fusion/evaluate_fusion.py \
--weights runs/yolov8n_fusion_ultralytics4/weights/best.pt
# Day split
python fusion/evaluate_fusion.py \
--weights runs/yolov8n_fusion_ultralytics4/weights/best.pt \
--data Datasets/val_day_full_fusion_dataset.yaml \
--ir-root Datasets/val_day_ir_full
# Night split
python fusion/evaluate_fusion.py \
--weights runs/yolov8n_fusion_ultralytics4/weights/best.pt \
--data Datasets/val_night_full_fusion_dataset.yaml \
--ir-root Datasets/val_night_ir_full
Key arguments for evaluate_fusion.py:
| Argument | Default | Description |
|---|---|---|
--weights |
(required) | Path to best.pt |
--data |
fusion/fusion_dataset.yaml | Dataset YAML |
--ir-root |
Datasets/.../ir | IR images root (overrides YAML default) |
--conf |
0.001 | Confidence threshold |
--iou |
0.6 | NMS IoU threshold |
--batch |
16 | Batch size |
# Full validation set
python mid_fusion/eval_mid_fusion.py \
--weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
--fusion-type mlp \
--data fusion/fusion_dataset.yaml \
--ir-root Datasets/RGBT-Tiny_YOLO_rgb_ir/ir \
--modality full_val
# Day split
python mid_fusion/eval_mid_fusion.py \
--weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
--fusion-type mlp \
--data Datasets/val_day_full_fusion_dataset.yaml \
--ir-root Datasets/val_day_ir_full \
--modality val_day_full
# Night split
python mid_fusion/eval_mid_fusion.py \
--weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
--fusion-type mlp \
--data Datasets/val_night_full_fusion_dataset.yaml \
--ir-root Datasets/val_night_ir_full \
--modality val_night_full
Key arguments for eval_mid_fusion.py:
| Argument | Default | Description |
|---|---|---|
--weights |
(required) | Path to best.pt |
--fusion-type |
mlp | Must match training fusion type |
--data |
fusion/fusion_dataset.yaml | Dataset YAML |
--ir-root |
Datasets/.../ir | IR images root |
--modality |
full_val | Label added to output directory name |
--conf |
0.001 | Confidence threshold |
--iou |
0.6 | NMS IoU threshold |
--batch |
8 | Batch size |
# Baseline
python baselines/inference_yolov8.py \
--weights runs/yolov8n_rgb_scratch_v2/weights/best.pt \
--source test_images/rgb/sample.jpg
# Mid-fusion
python mid_fusion/inference_mid_fusion.py \
--weights runs/mid_fusion/mlp_20260210_113817_phase2/weights/best.pt \
--fusion-type mlp \
--rgb test_images/rgb/sample.jpg \
--ir test_images/ir/sample.jpg
Full results, per-class breakdowns, day/night analysis, and improvement recommendations
are in evaluation_report.md. Key mAP@0.5 numbers:
| Model | Full Val | Day | Night |
|---|---|---|---|
| RGB Baseline | 0.406 | 0.497 | 0.236 |
| IR Baseline | 0.412 | 0.448 | 0.372 |
| Early Fusion (corrected) | 0.370 | 0.485 | 0.200 |
| Mid Fusion (MLP) | 0.390 | 0.481 | 0.239 |
The IR baseline outperforms all fusion models overall. The primary limitation of both
fusion architectures is fixed-weight combination at inference time, which prevents the
model from suppressing degraded RGB channels at night. See Section 7 and 8 of
evaluation_report.md for architectural analysis and recommendations.