Developer 3s Key Decision Metrics
On 2026-10-01 Cloudflare made AI Search generally available: managed index/retrieval (Workers AI + Vectorize + R2 + Browser Run) ships native multimodal embeddings (Qwen3-VL-Embedding), OCR for scanned PDFs, 10 MiB text/OCR PDF limits, hybrid search by default, and billing from 2026-11-01 with a free allotment.
Key Takeaways
- ✓GA 2026-10-01: changelog + blog.cloudflare.com/ai-search-ga/
- ✓Multimodal embeddings: @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2; image queries via REST/public endpoint
- ✓OCR on all accounts; text/code + OCR PDFs up to 10 MiB (non-OCR PDFs stay 4 MiB)
- ✓Hybrid search default; Workers AI embed/rerank folded into AI Search pricing (not Workers AI bill)
- ✓Pricing: $0.75/1M ingest tokens, $2/GB-mo storage, $0.75/1k semantic, $0.10/1k full-text; free tier 5M ingest / 10 GB / 1k+1k queries; bill from Nov 1, 2026
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Core Background & Industry Pain Points
Agents searching private docs, product images, or scanned PDFs usually assemble parsers, vector DBs, rerankers, and billing themselves. Cloudflare’s AI Search packages Workers AI + Vectorize + R2 + Browser Run; after preview it went GA on 2026-10-01 (changelog). Complementary to same-week Web Search API (live web): AI Search indexes your data.
Architecture Highlights & Internals
Pipeline: optional rewrite → embed (multimodal models embed query images directly; text-only models caption via ToMarkdown first) → parallel vector + keyword → fuse / optional rerank → chunks or generation. New instances default to hybrid search. Multimodal models: @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2 with MRL. Optional type inference (URL → website, bucket → R2). Workers AI embed/rerank costs fold into AI Search pricing (not Workers AI bill / Gateway logs); generation, rewrite, and external providers still use account/gateway.
Authoritative Benchmarks & Measured Scores
No official nDCG/recall leaderboard—do not invent. Verified pricing from blog + Limits & pricing: $0.75/1M ingest tokens, +$0.50/1M image processing, $2/GB-mo storage, $0.75/1k semantic, $0.10/1k full-text; free monthly 5M ingest / 10 GB / 1k+1k queries; billing from 2026-11-01. OCR on all accounts; text/code + OCR PDFs up to 10 MiB.
Developer Hands-on Guide
Follow https://developers.cloudflare.com/ai-search/; pick a multimodal embedder for image queries; enable OCR for scans; load-test before Nov 1; do not confuse with Web Search API (public web vs your corpus).

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.