Hi, thank you for releasing this excellent work!
I am interested in conducting a more detailed analysis of Qwen-RobotNav's performance on the VLN-CE benchmark. Would it be possible to provide per-episode evaluation outputs, in addition to the aggregated metrics?
For example, the following information would be very helpful:
- Episode ID
- Predicted trajectory or visited viewpoints
- Final stopping position
- Success / Oracle Success
- Navigation Error
- SPL
- Model-generated actions or responses at each step
Thank you very much for your time and for sharing this work!
Hi, thank you for releasing this excellent work!
I am interested in conducting a more detailed analysis of Qwen-RobotNav's performance on the VLN-CE benchmark. Would it be possible to provide per-episode evaluation outputs, in addition to the aggregated metrics?
For example, the following information would be very helpful:
Thank you very much for your time and for sharing this work!