Gestalt: Large Multimodal Interplay Model

  1. 1Gaoling School of Artificial Intelligence, Renmin University of China, Beijing
  2. 2Beijing Key Laboratory of Research on Large Models and Intelligent Governance
  3. 3Beihang University
  4. 4Beijing University of Posts and Telecommunications
  5. 5Shanghai Artificial Intelligence Laboratory
  6. 6AresoX

†Equal contribution,‡Team leader,✉︎Corresponding author

*Work partially done at Shanghai Artificial Intelligence Laboratory, as an internship.

Overview

OVERVIEW

The Multimodal Interplay Pyramid

Multimodal intelligence grows from preserving what each modality knows to discovering what they can reveal together.

The multimodal interplay pyramid: synergy at the top, redundancy alignment in the middle, and modality-specific information modeling at the base.

Multimodal Synergy

Combine complementary cues to derive information that neither modality provides alone.

Cross-Modal Redundancy Alignment

Align shared information across modalities to establish cross-modal correspondence.

Modality-Specific Information Modeling

Preserve information available only in one modality, providing the foundation for multimodal learning.

A dog wearing a blue collar; the original figure marks its collar region.

Multimodal Synergy: A Case Study

“I have two dogs. The larger one wears a red collar, while the smaller one wears a blue collar.”

The image provides the collar color; the text links collar color to size. Together, they identify the smaller dog.

ARCHITECTURE

From Controlled Exchange To Full Interplay

Gestalt brings vision and language into a shared discrete diffusion framework. Learnable interplay tokens connect two zones: one preserves modality-specific information, and the other enables deeper integration.

The complete Gestalt architecture, with Zone I below Zone II and a unified objective.

DATA & TRAINING

Multimodal Training And Interplay Curriculum

  1. 70M

    Multimodal Pretraining

    Joint Masked Prediction

  2. 8M

    Continual Pretraining

    Conditional Masked Prediction

  3. 13.7M

    Supervised Fine-Tuning

    Four Interplay Categories

    Plus ≈2.5M text-to-image examples across all SFT phases.

Supervised Fine-Tuning

Three-Phase Curriculum

Three phases progressively emphasize language and alignment, visual perception, and multimodal synergy while retaining all four data categories.

All four categories remain present; their sampling ratios change across phases.

IMAGE GENERATION

Fine-Grained Text-To-Image Generation

Qualitative comparisons of images generated under complex semantic constraints.

UniGenBench

Gestalt achieves the strongest overall result among the evaluated models, combining fine-grained semantic alignment with compositional generation.

ModelOverallStyleWorldAttr.ActionRel.Comp.GrammarLayoutLogicText
Gen. Only
DALL-E-370.8295.0892.7184.9868.3677.9073.8868.1971.7657.1118.26
SD-3.5-Large64.3588.1288.1578.7859.6367.6262.2165.2371.1944.9017.66
OmniGen271.3994.3584.8383.0366.5773.0670.4976.4080.6356.5527.99
Unified
Emu350.9589.3676.1666.8143.8051.7046.0050.2556.6727.431.36
Show-o270.3393.1188.4486.3569.0277.3776.4570.3080.6359.711.90
Janus-Pro71.1194.0288.1581.8169.1477.9676.5374.6282.1462.624.08
MMaDA40.1075.8352.7549.9032.4239.0638.3750.0043.0219.420.27
BAGEL71.2692.4489.3184.2167.6275.7074.7174.7581.9059.7112.23
Lumina-DiMOO71.8186.8888.5883.7169.6673.3374.9374.4984.8458.0123.64
Gestalt73.30↑1.4995.85↑0.7781.6582.5168.5778.12↑0.1681.20↑4.6782.23↑5.8385.79↑0.9573.28↑10.663.80

TIIF-Bench

Gestalt leads across short and long instructions, with a larger advantage when longer prompts introduce more interdependent requirements.

ModelOverallBasicAdvancedDesigner
ShortLongShortLongShortLongShortLong
Gen. Only
PixArt-Sigma62.0058.1270.6675.2557.6549.5062.1152.41
FLUX.1 Pro67.3269.8979.0878.9161.1065.3771.8068.80
MidJourney V768.7465.6977.4176.0064.6660.5368.8363.61
SD 3.5 Large71.1566.9678.3479.5667.6761.1864.4366.39
Unified
Emu343.3839.4449.8842.0837.0933.5253.7360.45
MMaDA52.3252.9065.7866.7150.3250.7660.4555.60
Show-o267.3868.7681.1684.0669.2572.9975.3775.75
Janus-Pro66.5065.0279.3378.2559.7158.8265.8460.25
BAGEL71.5071.7081.7980.0570.2472.1968.2867.91
Lumina-DiMOO71.2768.5375.5078.2970.4968.3369.7870.90
Gestalt74.57↑3.0778.86↑7.1683.79↑2.0086.06↑2.0072.60↑2.1177.69↑4.7084.70↑9.3388.81↑13.06

MULTIMODAL UNDERSTANDING

Strong Visual Understanding

Multimodal Understanding Benchmarks

The gains on vision-centric tasks reflect the value of preserving modality-specific information within a unified model.

ModelsGeneralVision-Centric
MME-PGQAMMStar-PPOPERWQAMMVPCVB2dCVB3d
AR-Based
BAGEL1687.0†66.470.988.267.669.3†77.784.2
Diffusion-based
MMaDA1410.7†61.3†43.086.1†48.217.355.354.8
Lumina-DiMOO1534.2†43.3-87.4†35.934.054.352.0
LaViDA-o1431.054.155.9--56.647.373.470.8
LLaDA-o1412.0†58.055.687.266.446.778.075.9
Omni-diffusion1216.7†----76.6†--------
Gestalt1600.2↑66.060.260.886.559.148.7↑1.478.9↑0.986.1↑10.2

LANGUAGE

Retains Strong Language Capability

Text-Only Benchmarks

Gestalt leads the evaluated diffusion-based unified models and achieves language performance comparable to autoregressive models such as Janus-Pro.

ModelMMLUTruthfulQAWinoGrandeHellaSwagARC-EARC-C
AR-Based
Show-o271.7046.9474.0376.6784.4358.53
Janus-Pro49.9041.7267.1768.4165.7440.70
BAGEL28.0240.5150.7528.5927.5323.63
Diffusion-based
MMaDA40.1443.8154.8545.8146.7228.67
Lumina-DiMOO29.7543.3451.6239.1144.4426.45
LLaDA-o25.2552.3051.2230.7236.6624.66
Gestalt49.50↑9.3647.5560.62↑5.7753.44↑7.6362.29↑15.5743.69↑15.02

MULTIMODAL INTERPLAY

An Integrated Vision–Language Space

Visual and textual tokens are more interleaved in Gestalt, suggesting a more integrated representation space than the evaluated diffusion-based and autoregressive baselines.

Gestalt forms a more interleaved distribution of visual and textual representations than the comparison models.

Preserving Unique Information, Deriving Synergy

Visual-Specific And Synergistic Capability

Strong visual-specific performance is paired with leading synergy results among the evaluated diffusion-based models.

ModelVisualSynergy
MIB-VCoreCog-SMMM-IMDbSRBench
AR-Based
BAGEL65.9665.0060.6051.89
Diffusion-based
MMaDA44.2343.2030.3936.72
Lumina-DiMOO56.4342.2031.2245.50
LLaDA-o60.3153.3039.6450.89
LaViDa-O58.0750.6051.5440.89
Gestalt59.8955.20↑1.9067.07↑15.5353.17↑2.28

Interplay-Aware Representations

Interplay tokens form distinct visual-unique, text-unique, and synergistic structures. The synergy distribution partially bridges the other two, suggesting that these tokens adapt to different information demands.

Interplay-token representations reflect visual-unique, text-unique, and synergistic information demands.