On September 1, xAI said LatchBio's independent BioSecBench-Refusal analysis found Grok 4.6 the only tested system scoring above 50% on both refusing disguised hazardous biology requests and completing routine biological work. Across harnesses, the trial-weighted harmonic mean averaged about 62.1% (roughly 59.2% red-team refusal and 64.8% routine completion). On BioSecBench-Surveillance it averaged about 53.5%, behind Opus 5 and ahead of GPT-5.6 Sol. LatchBio's blog says refusals are largely model-reasoning driven rather than API filters, without clear degradation on other biology benchmarks.

Key Takeaways

  • Only tested model above 50% on both refusal and routine biology tasks
  • ~62.1% harmonic mean; ~53.5% on Surveillance, behind Opus 5
  • LatchBio: refusals mostly model-reasoning driven without clear bio capability loss
ADSponsored