Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

depth — Fine-Tuned Multi-Model Monocular Depth Estimation for ROS 2

A ROS 2 (Humble) package that wraps two monocular depth estimation models, YOLO26 Depth and Depth Anything V2, as interchangeable nodes, using checkpoints fine-tuned in-house rather than the stock base weights. Both share the same RGB input topic and depth output topic, so you can drop either one into a robot's perception stack, or run them side by side to compare.

All fine-tuned checkpoints are hosted on Hugging Face: WasiqSaleem/Fine-Tuned-Depth-Estimation-for-ROS-2.

Contents

Overview

This package has two separate ROS 2 nodes. Each one subscribes to a live RGB image stream and publishes a metric depth prediction.

yolo_depth runs Ultralytics' YOLO26 Depth architecture, loaded with our own fine-tuned checkpoints rather than the stock pretrained weights. depth_anything_v2 runs Depth Anything V2 in metric depth mode, also loaded with our own fine-tuned checkpoints.

Both nodes expose the checkpoint choice and frame skip rate as ROS 2 parameters, so you can swap checkpoints or throttle inference load at launch time without touching the code. frame_skip exists to trade off latency against CPU load — running inference on every incoming frame is unnecessary for most navigation/mapping use cases and keeps the CPU pinned, so skipping frames (e.g. frame_skip:=3 → inference on 1 of every 3 frames) reduces both latency and CPU load at the cost of a lower effective depth update rate.

Models Implemented

Node Model Notes
yolo_depth YOLO26 Depth Fast, single-shot depth head from Ultralytics, fine-tuned checkpoints
depth_anything_v2 Depth Anything V2 (metric depth) ViT-based, strong accuracy, heavier compute, fine-tuned checkpoints

Fine-Tuning

Both models were fine-tuned in two stages rather than used off the shelf:

  • Stage 1 — NYU Depth V2. Both YOLO26 Depth and Depth Anything V2 were fine-tuned on the NYU Depth V2 dataset. DAv2 was trained for 60 epochs (checkpoint: early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth, early-stopped at epoch 56). YOLO26 Depth was trained for 150 epochs.
  • Stage 2 — ScanNet. The stage 1 (NYU Depth V2 fine-tuned) checkpoints were then further fine-tuned on ScanNet, producing the ScanNet-stage checkpoints referenced in the node parameters below. YOLO26 Depth's ScanNet-stage fine-tuning was run for both the Nano and Small model sizes.

For full training configuration, data splits, and fine-tuning methodology, see the separate repo: fine-tuning of depth models.

Package Structure

depth/
├── depth/
│   ├── yolo_depth.py          # YOLO26 Depth ROS 2 node
│   ├── depth_anything_v2.py   # DAv2 ROS 2 node
├── package.xml
├── setup.py
└── setup.cfg

Setup

Clone this package into your ROS 2 workspace src/ folder and build it with colcon build once the model dependencies below are installed.

1. Download the fine-tuned checkpoints

All four checkpoints are hosted on Hugging Face at WasiqSaleem/Fine-Tuned-Depth-Estimation-for-ROS-2. Download them with the Hugging Face CLI:

pip install -U "huggingface_hub[cli]"
huggingface-cli download WasiqSaleem/Fine-Tuned-Depth-Estimation-for-ROS-2 \
    --local-dir ./checkpoints

...or grab individual files directly from the Files and versions tab:

  • early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth — DAv2, NYUv2-stage (297 MB)
  • 20_epochs_scannet_documentation_model.pth — DAv2, ScanNet-stage (297 MB)
  • nano_yolo26-depth-scannet-batch3.pt — YOLO26 Depth Nano, ScanNet-stage (10.6 MB)
  • small_yolo26-depth-scannet-batch3.pt — YOLO26 Depth Small, ScanNet-stage

Then point each node at wherever you saved them — edit the path dictionaries near the top of each node's __init__:

depth/depth_anything_v2.py — self.model_list:

self.model_list = {
    'scannet': "/path/to/your/checkpoints/20_epochs_scannet_documentation_model.pth",
    'nyuv2': "/path/to/your/checkpoints/early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth",
}

depth/yolo_depth.py — self.model_paths:

self.model_paths = {
    'nano': '/path/to/your/checkpoints/nano_yolo26-depth-scannet-batch3.pt',
    'small': '/path/to/your/checkpoints/small_yolo26-depth-scannet-batch3.pt',
}

If you ran the huggingface-cli download command above with --local-dir ./checkpoints, that's simply ./checkpoints/<filename> for all four — no directory guessing needed.

2. Depth Anything V2 (DAv2)

git clone https://github.com/DepthAnything/Depth-Anything-V2
cd Depth-Anything-V2/metric_depth
pip install -r requirements.txt

This node does not use the base DAv2 checkpoints from the official repo — it loads our own fine-tuned checkpoints (see above).

If you're running on CPU, uninstall xformers. It's a GPU-only dependency and will throw an error otherwise. See DepthAnything/Depth-Anything-V2#312 for the exact error.

python3 -m pip uninstall xformers

3. YOLO26 Depth

pip install ultralytics

This node does not use the base YOLO26 Depth weights — it loads our own fine-tuned checkpoints (see above). See the official task docs for background on the base architecture: Ultralytics — Depth Estimation.

Topics

Both nodes share the same topic interface.

Direction Topic Type
Subscribed /camera/color/image_raw sensor_msgs/Image
Published /camera/depth/image_raw_cal sensor_msgs/Image (32FC1)

Running the Nodes

Depth Anything V2

ros2 run depth depth_anything_v2 --ros-args -p model_key:=scannet -p frame_skip:=3
ros2 run depth depth_anything_v2 --ros-args -p model_key:=nyuv2 -p frame_skip:=3

model_key options:

Key Checkpoint Fine-tuning stage
nyuv2 early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth Stage 1: NYU Depth V2 only
scannet 20_epochs_scannet_documentation_model.pth Stage 2: NYU Depth V2 → ScanNet

YOLO26 Depth

ros2 run depth yolo_depth --ros-args -p model_key:=small -p frame_skip:=3
ros2 run depth yolo_depth --ros-args -p model_key:=nano -p frame_skip:=3

model_key options:

Key Checkpoint Fine-tuning stage
small small_yolo26-depth-scannet-batch3.pt Stage 2: NYU Depth V2 → ScanNet (Small)
nano nano_yolo26-depth-scannet-batch3.pt Stage 2: NYU Depth V2 → ScanNet (Nano)

frame_skip works the same way across both nodes and takes any positive integer. It controls how many incoming RGB frames go by per inference run — frame_skip:=3 runs inference on 1 out of every 3 frames, trading a lower effective depth update rate for reduced latency and CPU load.

Model Comparison

Evaluated with the same rosbag/Gazebo ground-truth pipeline described in Testing / Evaluation. Lower is better for AbsRel, RMSE, and LogRMSE. Higher is better for δ1 through δ3.

Depth Anything V2 — base vs. fine-tuned stages

Model AbsRel↓ RMSE↓ LogRMSE↓ δ1↑ δ2↑ δ3↑
DAv2 Base 1.245 2.286 0.779 0.116 0.221 0.407
Early Stop 56/60 (nyuv2) 1.213 1.963 0.796 0.112 0.238 0.413
ScanNet 20 Epochs (scannet) 0.986 1.912 0.699 0.144 0.333 0.508

ScanNet 20-epoch fine-tune vs. DAv2 base:

Metric Change
AbsRel 20.8% reduction
RMSE 16.4% reduction
LogRMSE 10.2% reduction
Delta1 (δ<1.25) +24.0%
Delta2 (δ<1.25²) +51.1%
Delta3 (δ<1.25³) +24.9%

YOLO26 Depth — base vs. fine-tuned stages

Model Frames AbsRel↓ RMSE↓ LogRMSE↓ Delta1↑ Delta2↑ Delta3↑
YOLO Base Model 279 2.4376 4.0132 1.2022 0.0195 0.0562 0.1731
YOLO 150 Epochs (150epochs) 279 1.2619 2.0480 0.8134 0.1107 0.3359 0.5050
ScanNet Batch 1 278 0.5673 2.1515 0.6380 0.2085 0.4435 0.7218
ScanNet Batch 2 (scannet_batch2) 279 0.5047 2.1524 0.6162 0.2239 0.5053 0.7637

ScanNet Batch 2 vs. base model (best result):

Metric Change
AbsRel 79.3% reduction
RMSE 46.4% reduction
LogRMSE 48.7% reduction
Delta1 (δ<1.25) +1049.6% (10.5x)
Delta2 (δ<1.25²) +799.8% (8x)
Delta3 (δ<1.25³) +341.1% (4.4x)

Best YOLO checkpoint vs. best DAv2 checkpoint

Metric Top YOLO Top DAv2 Winner
AbsRel↓ 0.505 0.986 YOLO
RMSE↓ 2.370 1.912 DAv2
LogRMSE↓ 0.641 0.699 YOLO
Delta1↑ 0.242 0.144 YOLO
Delta2↑ 0.506 0.333 YOLO
Delta3↑ 0.730 0.508 YOLO

Testing / Evaluation

The accuracy numbers above come from a custom ROS 2 bag–based evaluation pipeline, distinct from the offline ScanNet/NYUv2 benchmark numbers reported in the fine-tuning repo. RGB and Gazebo ground-truth depth are matched with message_filters.ApproximateTimeSynchronizer, and per-frame AbsRel, RMSE, LogRMSE, and δ1 through δ3 get logged to CSV for offline aggregation. For a general reference on structuring depth model testing in ROS 2, the ROS 2 message_filters package docs cover the synchronization approach this is built on.

License

Apache-2.0

About

Fine-tuned version of Multi-Model-Monocular-Depth-Estimation-for-ROS-2 — ROS 2 (Humble) package for YOLO26 Depth and Depth Anything V2, using checkpoints fine-tuned on NYU Depth V2 and ScanNet (hosted on Hugging Face) instead of base pretrained weights.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages