Together AI is cutting Dedicated Inference on H100s from $5.49/hr to $3.99/hr for September, with the lower rate applied automatically to new and existing deployments. Supported options include Gemma 4, Qwen3/3.5, gpt-oss, Llama, Nemotron 3.5 Lightning, or a bring-your-own LoRA. For teams that pin coding agents or fine-tunes to dedicated GPUs, the change is a straight cut in 24/7 serving cost without a migration.

Key Takeaways

  • โœ“H100 Dedicated Inference drops from $5.49/hr to $3.99/hr in September, auto-applied to all deployments.
  • โœ“Catalog includes Gemma 4, Qwen3/3.5, gpt-oss, Llama, Nemotron 3.5 Lightning, and BYO LoRA.
  • โœ“Existing dedicated endpoints get the cut with no migration.
ADSponsored