While modern multimodal LLMs predominantly rely on pretrained visual encoders (CLIP, SigLIP) for visual priors, this decoupled design introduces cross-modal alignment friction and inference pipeline bloat. Researchers from CASIA and collaborating institutions present the first systematic scaling laws for encoder-free MLLMs learning representations directly from raw pixels. The study establishes that the visual prior advantage diminishes with scale, predicting that encoder-free models catch up and surpass encoder-based counterparts at ~10^22 FLOPs. Furthermore, the LLM naturally assumes vision tasks through early-layer bidirectional patch interactions and specialized MoE expert routing, establishing a unified blueprint for next-generation end-to-end multimodal foundations.
- ✓10^22 FLOPs Crossover Frontier: While text loss-compute frontiers virtually overlap, encoder-free multimodal loss catches up and surpasses encoder-based baselines at approximately 10^22 FLOPs—well within practical enterprise pretraining budgets.
- ✓Parameter-Compute Allocation Shift: Removing visual encoders shifts compute-optimal allocation toward larger model parameters for multimodal objectives, while leaving optimal text allocations unaffected.
- ✓Emergent Vision Mechanisms in Native LLMs: Early Transformer layers natively develop bidirectional visual token cross-attention, while sparse MoE routing converges into specialized visual experts, naturally absorbing the role of external vision towers.
🧭Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
🔗
Project Links & Resources
Direct AccessDirect access to official project resources and documentation🔬
In-Depth Technical Analysis
Core Background & Industry Pain Points Nearly all modern Multimodal Large Language Models (MLLMs)—ranging from frontier proprietary models to mainstream open-source architectures—rely on decoupled pipelines pairing a pretrained visual encoder (e.g., CLIP, SigLIP) with a projection layer and a language model backbone. While this paradigm successfully bootstrap visual capabilities by borrowing frozen visual priors, it creates severe architectural friction: 1. Information Bottlenecks and Cross-Modal Misalignment: Fixed-dimension representations from independently pretrained vision towers cause irreversible semantic loss on high-resolution spatial relationships, dense OCR text, and long-horizon video context; 2. Engineering Pipeline & Distributed Training Overhead: Maintaining separate network topologies with distinct tensor and pipeline parallelisms complicates distributed clusters and introduces serving latency bubbles. While an "encoder-free" architecture ingesting raw pixel patches directly into an autoregressive LLM offers simplicity, its scaling behavior has never been systematically charted. ### Architectural Highlights & Underlying Mechanics Researchers from CASIA and collaborating institutions present the first principled, multi-order scaling laws comparing encoder-free and encoder-based MLLMs across identical compute budgets: 1. Unified Autoregressive Pixel Tokenization: Raw pixel patches are mapped into input embeddings via a lightweight linear projection, entering the Transformer directly as standard tokens without external vision towers; 2. Shifted Compute-Optimal Allocation: While the text loss-compute frontier is virtually indistinguishable between the two paradigms, removing the visual encoder systematically shifts optimal compute allocation toward larger parameter counts for multimodal objectives, whereas text objective allocation remains invariant; 3. Emergent Vision Adaptation in Transformer Layers: As compute scales, the language backbone naturally assumes visual abstraction: early Transformer layers develop bidirectional attention across visual tokens, and sparse MoE routers dynamically direct visual patches to specialized early-stage experts. ### Benchmark & Experimental Validation - 10^22 FLOPs Crossover Threshold: At lower compute scales, encoder-based models lead due to inherited CLIP priors. However, the performance curve of encoder-free models scales at a steeper slope, crossing and surpassing encoder-based baselines at approximately 10^22 FLOPs—well within enterprise pretraining thresholds; - Diminishing Returns of Frozen Priors: The quantitative analysis confirms that the marginal value of pretrained visual encoders degrades exponentially with training compute; - Interleaved Image-Text Coherence: On complex multimodal reasoning benchmarks, the encoder-free architecture maintains consistent attention geometry across mixed text and image sequences, avoiding context fragmentation common to external projection bridges. ### Engineering Takeaways & Practical Guide - Paper & Reference: Full formal formulations and Pareto loss curves are documented in arXiv:2609.35457; - Architectural Recommendations: Large-scale labs training foundation models with pretraining compute exceeding 10^22 FLOPs should actively deprecate decoupled vision towers in favor of unified end-to-end architectures to simplify inference pipelines and training parallelism; - Edge Deployment Considerations: For smaller budget regimes (< 10^21 FLOPs), encoder-based designs remain cost-effective priors, or can serve as distillation targets for unified architectures.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.