Together AI is cutting Dedicated Inference on H100s from $5.49/hr to $3.99/hr for September, with the lower rate applied automatically to new and existing deployments. Supported options include Gemma 4, Qwen3/3.5, gpt-oss, Llama, Nemotron 3.5 Lightning, or a bring-your-own LoRA. For teams that pin coding agents or fine-tunes to dedicated GPUs, the change is a straight cut in 24/7 serving cost without a migration.
Key Takeaways
- โH100 Dedicated Inference drops from $5.49/hr to $3.99/hr in September, auto-applied to all deployments.
- โCatalog includes Gemma 4, Qwen3/3.5, gpt-oss, Llama, Nemotron 3.5 Lightning, and BYO LoRA.
- โExisting dedicated endpoints get the cut with no migration.