Together AI is cutting Dedicated Inference on H100s from $5.49/hr to $3.99/hr for September, with the lower rate applied automatically to new and existing deployments. Supported options include Gemma 4, Qwen3/3.5, gpt-oss, Llama, Nemotron 3.5 Lightning, or a bring-your-own LoRA. For teams that pin coding agents or fine-tunes to dedicated GPUs, the change is a straight cut in 24/7 serving cost without a migration.
Key Takeaways
- ✓H100 Dedicated Inference drops from $5.49/hr to $3.99/hr in September, auto-applied to all deployments.
- ✓Catalog includes Gemma 4, Qwen3/3.5, gpt-oss, Llama, Nemotron 3.5 Lightning, and BYO LoRA.
- ✓Existing dedicated endpoints get the cut with no migration.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.