On 2026-10-01 Cloudflare made AI Search generally available: managed index/retrieval (Workers AI + Vectorize + R2 + Browser Run) ships native multimodal embeddings (Qwen3-VL-Embedding), OCR for scanned PDFs, 10 MiB text/OCR PDF limits, hybrid search by default, and billing from 2026-11-01 with a free allotment.

Key Takeaways

  • ✓GA 2026-10-01: changelog + blog.cloudflare.com/ai-search-ga/
  • ✓Multimodal embeddings: @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2; image queries via REST/public endpoint
  • ✓OCR on all accounts; text/code + OCR PDFs up to 10 MiB (non-OCR PDFs stay 4 MiB)
  • ✓Hybrid search default; Workers AI embed/rerank folded into AI Search pricing (not Workers AI bill)
  • ✓Pricing: $0.75/1M ingest tokens, $2/GB-mo storage, $0.75/1k semantic, $0.10/1k full-text; free tier 5M ingest / 10 GB / 1k+1k queries; bill from Nov 1, 2026
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points

Agents searching private docs, product images, or scanned PDFs usually assemble parsers, vector DBs, rerankers, and billing themselves. Cloudflare’s AI Search packages Workers AI + Vectorize + R2 + Browser Run; after preview it went GA on 2026-10-01 (changelog). Complementary to same-week Web Search API (live web): AI Search indexes your data.

Architecture Highlights & Internals

Pipeline: optional rewrite → embed (multimodal models embed query images directly; text-only models caption via ToMarkdown first) → parallel vector + keyword → fuse / optional rerank → chunks or generation. New instances default to hybrid search. Multimodal models: @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2 with MRL. Optional type inference (URL → website, bucket → R2). Workers AI embed/rerank costs fold into AI Search pricing (not Workers AI bill / Gateway logs); generation, rewrite, and external providers still use account/gateway.

Authoritative Benchmarks & Measured Scores

No official nDCG/recall leaderboard—do not invent. Verified pricing from blog + Limits & pricing: $0.75/1M ingest tokens, +$0.50/1M image processing, $2/GB-mo storage, $0.75/1k semantic, $0.10/1k full-text; free monthly 5M ingest / 10 GB / 1k+1k queries; billing from 2026-11-01. OCR on all accounts; text/code + OCR PDFs up to 10 MiB.

Developer Hands-on Guide

Follow https://developers.cloudflare.com/ai-search/; pick a multimodal embedder for image queries; enable OCR for scans; load-test before Nov 1; do not confuse with Web Search API (public web vs your corpus).