arXiv cs.AIPaper
GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting
The approach is clever: translate vision to structured language, then work in language space rather than building a domain-specific 3D encoder. Results on ScanNet++ are competitive but not superior. This is incremental progress on a narrow task. Use it if you're already doing open-vocabulary segmentation without training data, otherwise the practical benefit is limited.