On Oct 7 Perplexity released pplx-embed-v2-late, multimodal late-interaction (ColBERT) retrievers in 0.6B (340M active) and 9B (7.4B active) sizes, built on Qwen3.5 with bidirectional attention. Each token gets a 128-dim vector scored by MaxSim, covering text, images and rendered PDFs or slides without OCR. Both sizes share one embedding space, so the 0.6B model can query a 9B-built index. Public ViDoRe v3 image nDCG@10 is 62.3% / 65.2%; MIT-licensed weights are on Hugging Face.

Key Takeaways

  • ✓Public ViDoRe v3 nDCG@10: 0.6B 62.3% image / 61.2% markdown; 9B 65.2% image / 64.7% markdown (0.6B model card)
  • ✓Perplexity says the 0.6B model matches models with about 5x the active parameters on ViDoRe V3 (blog)
  • ✓Shared embedding space: index with 9B, query online with 0.6B to cut query-time cost
  • ✓Training: distilled from an internal 18B ColBERT teacher with a token-level LEAF-style objective; 9B fully fine-tunes the last 8 layers and uses LoRA elsewhere, including the vision encoder
  • ✓Setup: sentence-transformers >= 6.0.0 and transformers >= 5.4.0, MIT license; encode text-only and image-only batches separately (9B model card)
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Perplexity released pplx-embed-v2-late, a family of multimodal late-interaction (ColBERT) retrievers in 0.6B (340M active) and 9B (7.4B active) sizes. Built on Qwen3.5 with bidirectional attention, they emit one 128-dim vector per token and score with MaxSim, retrieving text, images and rendered PDFs or slides without OCR. The two sizes share an embedding space, so a corpus indexed with 9B can be queried with 0.6B. Both were distilled from an internal 18B ColBERT teacher. On public ViDoRe v3 (nDCG@10) the 0.6B scores 62.3% image / 61.2% markdown and the 9B 65.2% / 64.7%; Perplexity says the 0.6B matches models with about 5x the active parameters. Weights are MIT-licensed on Hugging Face and load with sentence-transformers >= 6.0.0 via MultiVectorEncoder; text and image batches must be encoded separately. A full technical report is promised later this year.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.