Intern Lumina U2

A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding and Image Generation

Intern Lumina U2 unifies language, images, video, and 3D for text question answering, image understanding and editing, image generation, video understanding, and 3D understanding.

Intern Lumina U2 capabilities and benchmark overview
Figure 1. Capabilities across image generation and editing, image understanding, video understanding, and 3D understanding.

Experimental Results

Preliminary, partial results. Full comparison tables will appear in the upcoming technical report.

Task familyBenchmarkInternVL-ULLaDA2.0-UniLLaDA-oIntern Lumina U2
Visual reasoning & perception
Document & chartChartQA ↑76.6080.1087.9086.52
Document & chartCharXiv-DQ ↑53.3068.4069.8083.65
PerceptionHallusionBench ↑44.8050.2047.4062.15
KnowledgeMMMU-Pro ↑20.8034.0028.3036.13
Math reasoningMathVision ↑22.1026.7015.7033.22
Math reasoningDynaMath ↑46.9746.8940.5856.37
Temporal understanding
VideoVideo-MME ↑53.0043.7644.9051.26
VideoMVBench ↑59.6148.9747.1659.74
VideoMLVU ↑57.2745.8148.6257.03
Generation & editing
Image generationGenEval ↑0.850.890.860.81
Image generationDPG ↑85.1887.7687.0487.10
Image generationTIIF Short / Long ↑74.90 / 73.9084.60 / 85.95
Image editingImgEdit ↑3.673.923.8307

Method.

A unified architecture
for omni-visual tasks.

Intern Lumina U2 unified multi-modal architecture
Figure 2. Architecture overview. The encoder and two task heads share a MoE Diffusion LLM backbone across visual understanding, image generation, and editing.

Discrete VQ levels for visual tokens.

Comparison of universal, separate, and multi-codebook visual representations
Figure 3. Representation design. The proposed design uses discrete VQ levels to represent each visual position.

Parallel positions, sequential codebooks.

Multi-codebook autoregressive head for training and inference
Figure 4. Prediction head. Spatial positions are processed in parallel and codebooks are decoded in sequence.

Resources

Coming soon.