Thanks for maintaining such a great 3D-LLM list and for keeping this repository up to date!
We have one related work that we would love to suggest for inclusion in your awesome list.
Title: CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
arXiv: https://arxiv.org/abs/2601.21798
Project Page: https://cv.jream.top/CG-MLLM-page/
CG-MLLM is a novel Multi-modal Large Language Model (MLLM) that unifies 3D captioning and high-resolution 3D generation in a single framework. Built on a Mixture-of-Transformer architecture, it decouples a Token-level Autoregressive (TokenAR) Transformer for token-level content and a Block-level Autoregressive (BlockAR) Transformer for block-level content, enabling long-context interactions between standard tokens and spatial blocks. It outperforms existing MLLMs in generating high-fidelity 3D objects while showing that learning 3D generation transfers back to improve the model's image-based 3D spatial perception.
We believe it would be a valuable addition to your list, and we'd be grateful if you'd consider including it. Thank you for your time and for your great work on this repository!
Thanks for maintaining such a great 3D-LLM list and for keeping this repository up to date!
We have one related work that we would love to suggest for inclusion in your awesome list.
Title: CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
arXiv: https://arxiv.org/abs/2601.21798
Project Page: https://cv.jream.top/CG-MLLM-page/
CG-MLLM is a novel Multi-modal Large Language Model (MLLM) that unifies 3D captioning and high-resolution 3D generation in a single framework. Built on a Mixture-of-Transformer architecture, it decouples a Token-level Autoregressive (TokenAR) Transformer for token-level content and a Block-level Autoregressive (BlockAR) Transformer for block-level content, enabling long-context interactions between standard tokens and spatial blocks. It outperforms existing MLLMs in generating high-fidelity 3D objects while showing that learning 3D generation transfers back to improve the model's image-based 3D spatial perception.
We believe it would be a valuable addition to your list, and we'd be grateful if you'd consider including it. Thank you for your time and for your great work on this repository!