A ROS 2 (Humble) package that wraps two monocular depth estimation models, YOLO26 Depth and Depth Anything V2, as interchangeable nodes, using checkpoints fine-tuned in-house rather than the stock base weights. Both share the same RGB input topic and depth output topic, so you can drop either one into a robot's perception stack, or run them side by side to compare.
All fine-tuned checkpoints are hosted on Hugging Face: WasiqSaleem/Fine-Tuned-Depth-Estimation-for-ROS-2.
- Overview
- Models Implemented
- Fine-Tuning
- Package Structure
- Setup
- Topics
- Running the Nodes
- Model Comparison
- Testing / Evaluation
- License
This package has two separate ROS 2 nodes. Each one subscribes to a live RGB image stream and publishes a metric depth prediction.
yolo_depth runs Ultralytics' YOLO26 Depth architecture, loaded with our own fine-tuned checkpoints rather than the stock pretrained weights. depth_anything_v2 runs Depth Anything V2 in metric depth mode, also loaded with our own fine-tuned checkpoints.
Both nodes expose the checkpoint choice and frame skip rate as ROS 2 parameters, so you can swap checkpoints or throttle inference load at launch time without touching the code. frame_skip exists to trade off latency against CPU load — running inference on every incoming frame is unnecessary for most navigation/mapping use cases and keeps the CPU pinned, so skipping frames (e.g. frame_skip:=3 → inference on 1 of every 3 frames) reduces both latency and CPU load at the cost of a lower effective depth update rate.
| Node | Model | Notes |
|---|---|---|
yolo_depth |
YOLO26 Depth | Fast, single-shot depth head from Ultralytics, fine-tuned checkpoints |
depth_anything_v2 |
Depth Anything V2 (metric depth) | ViT-based, strong accuracy, heavier compute, fine-tuned checkpoints |
Both models were fine-tuned in two stages rather than used off the shelf:
- Stage 1 — NYU Depth V2. Both YOLO26 Depth and Depth Anything V2 were fine-tuned on the NYU Depth V2 dataset. DAv2 was trained for 60 epochs (checkpoint:
early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth, early-stopped at epoch 56). YOLO26 Depth was trained for 150 epochs. - Stage 2 — ScanNet. The stage 1 (NYU Depth V2 fine-tuned) checkpoints were then further fine-tuned on ScanNet, producing the ScanNet-stage checkpoints referenced in the node parameters below. YOLO26 Depth's ScanNet-stage fine-tuning was run for both the Nano and Small model sizes.
For full training configuration, data splits, and fine-tuning methodology, see the separate repo: fine-tuning of depth models.
depth/
├── depth/
│ ├── yolo_depth.py # YOLO26 Depth ROS 2 node
│ ├── depth_anything_v2.py # DAv2 ROS 2 node
├── package.xml
├── setup.py
└── setup.cfg
Clone this package into your ROS 2 workspace src/ folder and build it with colcon build once the model dependencies below are installed.
All four checkpoints are hosted on Hugging Face at WasiqSaleem/Fine-Tuned-Depth-Estimation-for-ROS-2. Download them with the Hugging Face CLI:
pip install -U "huggingface_hub[cli]"
huggingface-cli download WasiqSaleem/Fine-Tuned-Depth-Estimation-for-ROS-2 \
--local-dir ./checkpoints...or grab individual files directly from the Files and versions tab:
early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth— DAv2, NYUv2-stage (297 MB)20_epochs_scannet_documentation_model.pth— DAv2, ScanNet-stage (297 MB)nano_yolo26-depth-scannet-batch3.pt— YOLO26 Depth Nano, ScanNet-stage (10.6 MB)small_yolo26-depth-scannet-batch3.pt— YOLO26 Depth Small, ScanNet-stage
Then point each node at wherever you saved them — edit the path dictionaries near the top of each node's __init__:
depth/depth_anything_v2.py — self.model_list:
self.model_list = {
'scannet': "/path/to/your/checkpoints/20_epochs_scannet_documentation_model.pth",
'nyuv2': "/path/to/your/checkpoints/early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth",
}depth/yolo_depth.py — self.model_paths:
self.model_paths = {
'nano': '/path/to/your/checkpoints/nano_yolo26-depth-scannet-batch3.pt',
'small': '/path/to/your/checkpoints/small_yolo26-depth-scannet-batch3.pt',
}If you ran the huggingface-cli download command above with --local-dir ./checkpoints, that's simply ./checkpoints/<filename> for all four — no directory guessing needed.
git clone https://github.com/DepthAnything/Depth-Anything-V2
cd Depth-Anything-V2/metric_depth
pip install -r requirements.txtThis node does not use the base DAv2 checkpoints from the official repo — it loads our own fine-tuned checkpoints (see above).
If you're running on CPU, uninstall xformers. It's a GPU-only dependency and will throw an error otherwise. See DepthAnything/Depth-Anything-V2#312 for the exact error.
python3 -m pip uninstall xformerspip install ultralyticsThis node does not use the base YOLO26 Depth weights — it loads our own fine-tuned checkpoints (see above). See the official task docs for background on the base architecture: Ultralytics — Depth Estimation.
Both nodes share the same topic interface.
| Direction | Topic | Type |
|---|---|---|
| Subscribed | /camera/color/image_raw |
sensor_msgs/Image |
| Published | /camera/depth/image_raw_cal |
sensor_msgs/Image (32FC1) |
ros2 run depth depth_anything_v2 --ros-args -p model_key:=scannet -p frame_skip:=3
ros2 run depth depth_anything_v2 --ros-args -p model_key:=nyuv2 -p frame_skip:=3model_key options:
| Key | Checkpoint | Fine-tuning stage |
|---|---|---|
nyuv2 |
early_stopping_56_out_of_60_epochs_multi_loss_documentation_model.pth |
Stage 1: NYU Depth V2 only |
scannet |
20_epochs_scannet_documentation_model.pth |
Stage 2: NYU Depth V2 → ScanNet |
ros2 run depth yolo_depth --ros-args -p model_key:=small -p frame_skip:=3
ros2 run depth yolo_depth --ros-args -p model_key:=nano -p frame_skip:=3model_key options:
| Key | Checkpoint | Fine-tuning stage |
|---|---|---|
small |
small_yolo26-depth-scannet-batch3.pt |
Stage 2: NYU Depth V2 → ScanNet (Small) |
nano |
nano_yolo26-depth-scannet-batch3.pt |
Stage 2: NYU Depth V2 → ScanNet (Nano) |
frame_skip works the same way across both nodes and takes any positive integer. It controls how many incoming RGB frames go by per inference run — frame_skip:=3 runs inference on 1 out of every 3 frames, trading a lower effective depth update rate for reduced latency and CPU load.
Evaluated with the same rosbag/Gazebo ground-truth pipeline described in Testing / Evaluation. Lower is better for AbsRel, RMSE, and LogRMSE. Higher is better for δ1 through δ3.
| Model | AbsRel↓ | RMSE↓ | LogRMSE↓ | δ1↑ | δ2↑ | δ3↑ |
|---|---|---|---|---|---|---|
| DAv2 Base | 1.245 | 2.286 | 0.779 | 0.116 | 0.221 | 0.407 |
| Early Stop 56/60 (nyuv2) | 1.213 | 1.963 | 0.796 | 0.112 | 0.238 | 0.413 |
| ScanNet 20 Epochs (scannet) | 0.986 | 1.912 | 0.699 | 0.144 | 0.333 | 0.508 |
ScanNet 20-epoch fine-tune vs. DAv2 base:
| Metric | Change |
|---|---|
| AbsRel | 20.8% reduction |
| RMSE | 16.4% reduction |
| LogRMSE | 10.2% reduction |
| Delta1 (δ<1.25) | +24.0% |
| Delta2 (δ<1.25²) | +51.1% |
| Delta3 (δ<1.25³) | +24.9% |
| Model | Frames | AbsRel↓ | RMSE↓ | LogRMSE↓ | Delta1↑ | Delta2↑ | Delta3↑ |
|---|---|---|---|---|---|---|---|
| YOLO Base Model | 279 | 2.4376 | 4.0132 | 1.2022 | 0.0195 | 0.0562 | 0.1731 |
| YOLO 150 Epochs (150epochs) | 279 | 1.2619 | 2.0480 | 0.8134 | 0.1107 | 0.3359 | 0.5050 |
| ScanNet Batch 1 | 278 | 0.5673 | 2.1515 | 0.6380 | 0.2085 | 0.4435 | 0.7218 |
| ScanNet Batch 2 (scannet_batch2) | 279 | 0.5047 | 2.1524 | 0.6162 | 0.2239 | 0.5053 | 0.7637 |
ScanNet Batch 2 vs. base model (best result):
| Metric | Change |
|---|---|
| AbsRel | 79.3% reduction |
| RMSE | 46.4% reduction |
| LogRMSE | 48.7% reduction |
| Delta1 (δ<1.25) | +1049.6% (10.5x) |
| Delta2 (δ<1.25²) | +799.8% (8x) |
| Delta3 (δ<1.25³) | +341.1% (4.4x) |
| Metric | Top YOLO | Top DAv2 | Winner |
|---|---|---|---|
| AbsRel↓ | 0.505 | 0.986 | YOLO |
| RMSE↓ | 2.370 | 1.912 | DAv2 |
| LogRMSE↓ | 0.641 | 0.699 | YOLO |
| Delta1↑ | 0.242 | 0.144 | YOLO |
| Delta2↑ | 0.506 | 0.333 | YOLO |
| Delta3↑ | 0.730 | 0.508 | YOLO |
The accuracy numbers above come from a custom ROS 2 bag–based evaluation pipeline, distinct from the offline ScanNet/NYUv2 benchmark numbers reported in the fine-tuning repo. RGB and Gazebo ground-truth depth are matched with message_filters.ApproximateTimeSynchronizer, and per-frame AbsRel, RMSE, LogRMSE, and δ1 through δ3 get logged to CSV for offline aggregation. For a general reference on structuring depth model testing in ROS 2, the ROS 2 message_filters package docs cover the synchronization approach this is built on.
Apache-2.0