Elon Musk shared that Grok 4.7 is in final polish, explaining that RL over-penalizing token length caused models to surrender prematurely on hard problems. The team is correcting length penalties and rigor before release in days.

Key Takeaways

  • Grok 4.7 completed base pretraining, now undergoing RL reward signal fine-tuning
  • Removes token length penalties causing premature model surrender on complex reasoning
  • Bolsters multi-step verification and reflection ahead of full rollout in days
ADSponsored