Elon Musk shared that Grok 4.7 is in final polish, explaining that RL over-penalizing token length caused models to surrender prematurely on hard problems. The team is correcting length penalties and rigor before release in days.
Key Takeaways
- ✓Grok 4.7 completed base pretraining, now undergoing RL reward signal fine-tuning
- ✓Removes token length penalties causing premature model surrender on complex reasoning
- ✓Bolsters multi-step verification and reflection ahead of full rollout in days
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.