Command Code CEO Ahmad Awais shared Composio benchmark results claiming Command Code solved 18 of 30 tasks with a 375-second average time, leading the comparison on open models. The post references Composio testing GPT-6 Astra across six coding-agent harnesses and highlights substantial token differences when harnesses fail.
Key Takeaways
- βCommand Code solved 18 of 30 tasks, with the post claiming the best accuracy among the tested harnesses.
- βIts reported average task time was 375 seconds, also described as fastest and ahead of DeepSeek's own harness.
- βThe cited Composio test covered Codex, Claude Code, OpenCode, Hermes Agent, Pi Agent, and Command Code.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.