Skip to content

[Paper Suggestion] CG-MLLM: Captioning and Generating 3D content via Multi-modal LLMs(ICML 2026 Accepted) #78

Description

@dreaming-huang

Thanks for maintaining such a great 3D-LLM list and for keeping this repository up to date!

We have one related work that we would love to suggest for inclusion in your awesome list.

Title: CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models

arXiv: https://arxiv.org/abs/2601.21798

Project Page: https://cv.jream.top/CG-MLLM-page/

CG-MLLM is a novel Multi-modal Large Language Model (MLLM) that unifies 3D captioning and high-resolution 3D generation in a single framework. Built on a Mixture-of-Transformer architecture, it decouples a Token-level Autoregressive (TokenAR) Transformer for token-level content and a Block-level Autoregressive (BlockAR) Transformer for block-level content, enabling long-context interactions between standard tokens and spatial blocks. It outperforms existing MLLMs in generating high-fidelity 3D objects while showing that learning 3D generation transfers back to improve the model's image-based 3D spatial perception.

We believe it would be a valuable addition to your list, and we'd be grateful if you'd consider including it. Thank you for your time and for your great work on this repository!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions