Hello, I am very interested in your work and would like to build upon this repository for further research.
During my recent reproduction efforts, I prepared the UMD dataset—compiling 2,004 samples based on the paper's configuration (rather than 1,922). However, even when ensuring consistency in other aspects (such as using GPT-5.4), I observed that the gIoU was 5 points lower and the CIoU was 15 points lower than the metrics reported in the paper.
I am unsure where the reproduction went wrong, especially considering that when using the full-8b-vl model on the RAGNet-3DOI dataset, the performance gap was only 3 to 5 points. I would greatly appreciate any suggestions you might have.
Hello, I am very interested in your work and would like to build upon this repository for further research.
During my recent reproduction efforts, I prepared the UMD dataset—compiling 2,004 samples based on the paper's configuration (rather than 1,922). However, even when ensuring consistency in other aspects (such as using GPT-5.4), I observed that the gIoU was 5 points lower and the CIoU was 15 points lower than the metrics reported in the paper.
I am unsure where the reproduction went wrong, especially considering that when using the full-8b-vl model on the RAGNet-3DOI dataset, the performance gap was only 3 to 5 points. I would greatly appreciate any suggestions you might have.