CrewAI 1.15.23 (2026-09-28) adds native Gemini 3.8 Flash support; crewai eval can score the last AMP-traced run (not just print it); task spans capture declared output format/results; LLM retries throttled providers and Bedrock acall falls back to sync. Tightens the observe→evaluate loop for multi-agent crews.

Key Takeaways

  • ✓Release 1.15.23 / PyPI
  • ✓Native Gemini 3.8 Flash via Google Gen AI SDK (LLM docs)
  • ✓crewai eval scores the last AMP-traced run; TUI Evaluate button + graded areas
  • ✓Task spans record declared output format and results
  • ✓Retries on throttled providers; Bedrock acall→sync fallback; SQLite/S3/Selenium cleanup fixes
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Background Crew frameworks lag new Gemini IDs and often stop at logs—no tight last-trace→score loop. Throttling and Bedrock async/sync mismatches break long runs. ### What shipped 1.15.23 adds native Gemini 3.8 Flash, AMP evaluation of the last traced run via crewai eval, richer task spans (output format + results), provider retry, and Bedrock acall→sync fallback, plus resource-cleanup fixes. ### Benchmarks No framework-only public scores in this patch; judge on your task suite and Google’s model card for 3.8 Flash. ### Get started pip install -U crewai==1.15.23, set Gemini via LLM docs, enable tracing, then crewai eval on the last AMP run. Repo: crewAIInc/crewAI.