Ollama’s Pro, Max, and Team plans now bill at published per-token rates against a monthly credit pool: $20/month with $60 of usage, $100 with $300, and $500 Team with $1,000 shared across unlimited users. The plans plug into Claude Code, Codex, and a first-party API, with no service fees and no 5-hour or weekly caps—when the pool runs out, traffic continues at the same token price. Inference stays on dedicated US/EU compute (Singapore for a limited Qwen set) with zero retention and no training on prompts. Existing subscribers can keep the old plan; new signups land on the token model because GPU-time billing became hard to forecast as open models scaled to sizes such as Kimi K3.

Key Takeaways

  • Pro/Max/Team are $20/$100/$500 with $60/$300/$1,000 monthly token credits; Team shares one pool with unlimited seats.
  • No service fees or 5-hour/weekly caps; unused work continues at the published token rate and works with Claude Code and Codex.
  • US/EU dedicated inference with zero retention; existing plans stay until you upgrade, new signups use token pricing.
ADSponsored