Escaping the Forgetting-Synergy Dilemma in Native Multimodal Pretraining
Xiangyue Liu1, Zijian Zhang2, Miles Yang2, Zhao Zhong2, Liefeng Bo2, Ping Tan1*
1HKUST 2Tencent Hunyuan
Figure 1. (Left) Performance on MMLU (language ability) across composable pretraining stages (LM → +MMU → +T2I). Standard MoE and structurally-isolated MoT suffer catastrophic routing collapse upon integrating continuous generative objectives. Rosetta maintains a stable semantic anchor throughout all stages. (Right) Qualitative image generation results from the Rosetta model.
Figure 2. Rosetta FFN. Three mechanisms enable non-destructive modality expansion: (1) Unified Attention — globally shared QKV projections preserve dense cross-modal interactions. (2) Composable FFN — modality-specific plug-and-play experts (Text / ViT / VAE) are bridged by a single Global Shared Expert that anchors foundational knowledge. (3) Conflict-Free Optimization (MAOP) — surgically neutralizes destructive gradients with zero memory overhead.
Requirements: Python 3.12+, CUDA 12.8+
git clone https://github.com/Lxiangyue/Rosetta.git
cd Rosetta
conda create -n rosetta python=3.12 -y
conda activate rosetta
# 1. Install PyTorch using your CUDA version
pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu128
# 2. Install dependencies
pip install -r requirements.txt
# 3. Install Flash Attention (required; takes ~20 mins to compile)
FLASH_ATTENTION_FORCE_BUILD=TRUE MAX_JOBS=8 pip install flash-attn==2.8.1 --no-build-isolation --no-binary flash-attn --no-cache-dir
Before running training or evaluation, please set up the shared assets, checkpoints, and datasets.
# Required for both training and evaluation (includes VAE, ViT, tokenizer, and evaluation datasets; total 18G)
hf download tencent/Rosetta-inference public_assets.zip --local-dir . && unzip -o public_assets.zip && rm public_assets.zip# For Evaluation (Rosetta Main Model; total 17G)
hf download tencent/Rosetta-inference --include "checkpoints/Rosetta-3.8B-A1B/**" --local-dir .
# For Training (Stage 2/3 Init Checkpoints; total 17+17+20G)
hf download tencent/Rosetta-inference --include "checkpoints/Rosetta-3.8B-A1B-stage2-lm-mmu/**" --local-dir .
hf download tencent/Rosetta-inference --include "checkpoints/MoE-3.8B-A1B-stage2-lm-mmu/**" --local-dir .
hf download tencent/Rosetta-inference --include "checkpoints/MoT-4.5B-A1B-stage3-init/**" --local-dir .# Prepares open-source training data (LM data from LLaVA, MMU+T2I data from BAGEL; total 379M)
bash scripts/download_example_data.shExpected directory layout
Rosetta/
├── checkpoints/
│ ├── Rosetta-3.8B-A1B/ ← For evaluation
│ │ └── hf_weights/ ← safetensors model weights
│ ├── Rosetta-3.8B-A1B-stage2-lm-mmu/ ← For training
│ ├── MoE-3.8B-A1B-stage2-lm-mmu/ ← For training
│ └── MoT-4.5B-A1B-stage3-init/ ← For training
└── example_data/
│ ├── lm/
│ │ └── conversation_58k.json
│ ├── mmu/
│ │ └── images/
│ │ └── llava_ov_si.jsonl
│ ├── t2i/
│ │ └── *.parquet
└── public_assets/
├── image_encoder/ ← VAE (FLUX.2-VAE)
├── vision_encoder/ ← ViT (Qwen3-VL-30B-A3B-Instruct)
├── pretrained_llm/ ← tokenizer + Qwen3-0.6B-Base
├── generation_configs/ ← generation settings for evaluation
└── evaluation/ ← benchmark datasets
├── ARC_C/
├── MMLU/
├── BBH/
├── MBPP/
├── MMMU/
├── MMBench/
├── POPE/
├── AI2D/
├── RealWorldQA/
└── T2I-CompBench/
└── dataset/
└── COCO/
⏳ Time to Reproduce: ~90 mins on 8x H20 GPUs (VRAM > 40GB).
We provide a minimal one-command demo to reproduce the catastrophic forgetting phenomenon and validate Rosetta's structural immunity. By injecting the Text-to-Image (T2I) generation task for just 100 steps of training, you can directly compare Rosetta against MoE and MoT:
bash scripts/run/run_example_train.shThis script automatically evaluates the ARC-Challenge score (Language Ability) at step 0 and 100, generating the following comparison in outputs/example_train/arc_step0_step100.png:
Key insight: While MoE and MoT show a significant performance drop in language ability when learning visual generation, Rosetta robustly averts this forgetting.
🚀 Multinode Training
⏳ Time to Reproduce: ~30 mins on 64x H20
- Create
hosts.txtand replace entries with your own node hostnames/IPs. The first listed node will coordinate the distributed launch:
node-0-ip
node-1-ip
node-2-ip
node-3-ip
node-4-ip
node-5-ip
node-6-ip
node-7-ip
- Run multinode training:
hostfile=hosts.txt HOST_NUM=8 HOST_GPU_NUM=8 \
bash launch/run_multinode.sh scripts/run/run_example_train.sh --gradient-accumulation-steps 1Note: We set
--gradient-accumulation-steps 1to keep the global batch size (gbs) consistent with the single-node 8 GPUs setting.
💻 Single-GPU Training
Note: We provide two options for single-GPU users: the Equivalent Training (acc=64) to match the 8-GPU global batch size for exact score comparison (significantly slower), and the Fast Training (acc=8) for rapid experimentation (scores not directly comparable).
- Equivalent Training (~9 hours on 1x H20)
HOST_GPU_NUM=1 bash scripts/run/run_example_train.sh --gradient-accumulation-steps 64- Fast Training (~90 mins on 1x H20)
HOST_GPU_NUM=1 bash scripts/run/run_example_train.sh --gradient-accumulation-steps 8For complete training recipes, hyperparameters, data mix methods, please refer to TRAIN.md.
Run a quick evaluation (e.g., ARC-Challenge on Rosetta-3.8B-A1B; ~2 mins using 8 GPUS):
bash scripts/eval/eval_arc_c.shWe provide scripts for evaluating 11 benchmarks covering LM, VLM, and T2I tasks, please see EVAL.md.
@misc{liu2026rosettacomposablenativemultimodal,
title={Rosetta: Composable Native Multimodal Pretraining},
author={Xiangyue Liu and Zijian Zhang and Miles Yang and Zhao Zhong and Liefeng Bo and Ping Tan},
year={2026},
eprint={2607.00293},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.00293},
}


