Meta previewed WildArtifactBench, an internal eval that scores multimodal agents on complex real-world deliverables using win rates and Elo from human and agentic preference judges instead of ground-truth rubrics. It is releasing 10 tasks, with public artifact pages such as Cat Mesh and Plywood Whale. The design targets practical utility across formats—pages, meshes, mixed media—where a single gold answer is the wrong unit of measurement for agent work.

Key Takeaways

  • The bench uses preference win rates and Elo from human and agentic judges, not gold rubrics.
  • Meta is releasing 10 tasks and public artifact pages such as Cat Mesh and Plywood Whale.
  • It is built to score multimodal agents on real deliverables across mixed output formats.
ADSponsored