Meta previewed WildArtifactBench, an internal eval that scores multimodal agents on complex real-world deliverables using win rates and Elo from human and agentic preference judges instead of ground-truth rubrics. It is releasing 10 tasks, with public artifact pages such as Cat Mesh and Plywood Whale. The design targets practical utility across formats—pages, meshes, mixed media—where a single gold answer is the wrong unit of measurement for agent work.
Key Takeaways
- ✓The bench uses preference win rates and Elo from human and agentic judges, not gold rubrics.
- ✓Meta is releasing 10 tasks and public artifact pages such as Cat Mesh and Plywood Whale.
- ✓It is built to score multimodal agents on real deliverables across mixed output formats.