Together AI
VerifiedCloud platform for training and running open-source AI models.
Together AI Pricing History
Historical price tier adjustments and billing model changes.
Together AI Chronological Timeline
How to train your own Jev for $17
Together AI released the together/Tev1-4B-experimental classifier model based on Qwen3.5 4B. The company provided a technical workflow for users to fine-tune custom versions of this model on their serverless platform for a cost of $17.
Canary rollouts: upgrade models in production without downtime
A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference....
How a global fintech scaled coding agent traffic with Dedicated Model Inference
Together AI has introduced Dedicated Model Inference (DMI) to provide enterprise clients with isolated compute resources for model hosting. This capability enables engineering teams to manage independent scaling, model selection, and testing environments for high-traffic coding agents.
Migrating from closed to open source models, Together
Together AI has introduced a structured five-stage framework to facilitate the migration of enterprise workloads from proprietary closed-source models to open-source alternatives. The playbook defines a standardized methodology covering discovery, evaluation, adaptation, decision-making, and production deployment.
Together AI Introduces GLM-5.3 Flash Model with 17x Cost Reduction
Together AI has launched the GLM-5.3 Flash model, which provides a 17x reduction in cost compared to the standard GLM-5.3. The model maintains high performance with only a 5.6 point drop in pass@1 accuracy.
Together AI expands fine-tuning service with more models, live metrics, and finer controls
Together AI has integrated new open-weight models and introduced granular training features including Expert LoRA, early stopping, and live experiment tracking. The update also implements pre-flight validation and tokenized dataset previews alongside reduced pricing for specific models.
Introducing preemptible compute: the same compute, half the price
Together AI has introduced preemptible compute instances for its GPU clusters. This new offering provides identical GPU capacity at a 50% discount compared to on-demand rates, subject to a five-minute drain window.
To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
Together AI has ported the ThunderKittens library to support the NVIDIA Vera Rubin NVL72 architecture. The implementation includes a rebuilt NVFP4 GEMM kernel that achieves over 22 PFLOPS performance.
The Open Source AI Stack
Together AI has released a comprehensive open-source AI stack designed to streamline the deployment and fine-tuning of large language models. The stack integrates optimized inference engines, data processing pipelines, and model training frameworks to reduce latency and infrastructure overhead.
Together AI Rolls Out Next-Generation Platform Capabilities & API Architecture
Together AI released substantial architectural upgrades, introducing enhanced API integrations and updated workflow tooling for its Artificial Intelligence ecosystem.
GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing
Together AI has introduced GLM-5.3 Flash as a high-efficiency alternative to the standard GLM-5.3 model. The new model achieves a 17x reduction in cost while maintaining performance within 5.6 points of pass@1 and 2.6 points of pass@4 on the DeepSWE benchmark.
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
Together AI evaluated GLM-5.3 and GPT-5.6 Sol on 904 DeepSWE rollouts, establishing performance benchmarks for coding tasks. The analysis confirms GLM-5.3 achieves higher pass@4 efficiency at 50% of the cost compared to GPT-5.6 Sol.
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
Together AI has introduced the GLM-5.3 model, which achieves parity with Claude Fable 5 on pass@1 benchmarks for DeepSWE. The model demonstrates superior pass@4 performance while reducing operational costs to $3.99 per rollout compared to $21.63 for Claude Fable 5.
DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
Together AI evaluated the performance of DeepSeek V4 Pro 0813 and GPT-5.6 Sol using 904 DeepSWE rollouts. The analysis confirms that GPT-5.6 Sol leads in pass@1 accuracy while DeepSeek V4 Pro 0813 excels in pass@4 and cost-efficiency.
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
Together AI released performance benchmarks for DeepSeek V4 Pro 0813 and Claude Fable 5 on the DeepSWE coding benchmark. The data establishes a cost-performance trade-off where Claude Fable 5 leads in pass@1 accuracy while DeepSeek V4 Pro 0813 excels in pass@4 and cascade routing efficiency.
A/B test models in production
Together AI has introduced native A/B testing capabilities directly at the API endpoint level. This feature allows developers to route traffic between different model versions without modifying application-side routing logic.
Complete Together AI Change Log Index
| Date | Change Title | Type | Impact | Details |
|---|---|---|---|---|
| Sep 23, 2026 | How to train your own Jev for $17 | product_launch | 9/10 | View ➔ |
| Sep 22, 2026 | Canary rollouts: upgrade models in production without downtime | feature | 8/10 | View ➔ |
| Sep 18, 2026 | How a global fintech scaled coding agent traffic with Dedicated Model Inference | product_launch | 8/10 | View ➔ |
| Sep 16, 2026 | Migrating from closed to open source models, Together | product_launch | 7/10 | View ➔ |
| Sep 14, 2026 | Together AI Introduces GLM-5.3 Flash Model with 17x Cost Reduction | product_launch | 7/10 | View ➔ |
| Sep 11, 2026 | Together AI expands fine-tuning service with more models, live metrics, and finer controls | feature | 8/10 | View ➔ |
| Sep 10, 2026 | Introducing preemptible compute: the same compute, half the price | pricing | 8/10 | View ➔ |
| Sep 10, 2026 | To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72! | feature | 9/10 | View ➔ |
| Sep 9, 2026 | The Open Source AI Stack | product_launch | 9/10 | View ➔ |
| Sep 5, 2026 | Together AI Rolls Out Next-Generation Platform Capabilities & API Architecture | feature | 8/10 | View ➔ |
| Aug 28, 2026 | GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing | product_launch | 9/10 | View ➔ |
| Aug 21, 2026 | GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing | feature | 8/10 | View ➔ |
| Aug 21, 2026 | GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing | product_launch | 9/10 | View ➔ |
| Aug 18, 2026 | DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing | product_launch | 9/10 | View ➔ |
| Aug 17, 2026 | DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing | product_launch | 9/10 | View ➔ |
| Aug 17, 2026 | A/B test models in production | feature | 8/10 | View ➔ |