Congratulation on the CVPR acceptance! Your paper notes that you use MM-DiT, but for conditionals you "use standard cross-attention conditioning on the flattened patchified sequence each layer of the network".
Do you only have the noise input as a stream to MMDiT and have the conditionals cross attend to them after every self-attention block? If so, why bother with MMDiT (how is joint attention useful here?) and just use DiT / JiT instead? Or is it that you use the expanded / funneled fields as independent streams?
Happy to wait for the code drop, but could also do with a short response!
Thanks!
Congratulation on the CVPR acceptance! Your paper notes that you use MM-DiT, but for conditionals you "use standard cross-attention conditioning on the flattened patchified sequence each layer of the network".
Do you only have the noise input as a stream to MMDiT and have the conditionals cross attend to them after every self-attention block? If so, why bother with MMDiT (how is joint attention useful here?) and just use DiT / JiT instead? Or is it that you use the expanded / funneled fields as independent streams?
Happy to wait for the code drop, but could also do with a short response!
Thanks!