Command Code CEO Ahmad Awais shared Composio benchmark results claiming Command Code solved 18 of 30 tasks with a 375-second average time, leading the comparison on open models. The post references Composio testing GPT-6 Astra across six coding-agent harnesses and highlights substantial token differences when harnesses fail.

Key Takeaways

  • βœ“Command Code solved 18 of 30 tasks, with the post claiming the best accuracy among the tested harnesses.
  • βœ“Its reported average task time was 375 seconds, also described as fastest and ahead of DeepSeek's own harness.
  • βœ“The cited Composio test covered Codex, Claude Code, OpenCode, Hermes Agent, Pi Agent, and Command Code.
ADSponsored