product launchSep 23, 2026
How to train your own Jev for $17
Together AI released the together/Tev1-4B-experimental classifier model based on Qwen3.5 4B. The company provided a technical workflow for users to fine-tune custom versions of this model on their serverless platform for a cost of $17.
➔ Availability of the together/Tev1-4B-experimental model and a verified $17 fine-tuning workflow.
featureSep 22, 2026
Canary rollouts: upgrade models in production without downtime
A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference....
➔ Updated platform deployment with Canary rollouts: upgrade models in production without downtime.
product launchSep 18, 2026
How a global fintech scaled coding agent traffic with Dedicated Model Inference
Together AI has introduced Dedicated Model Inference (DMI) to provide enterprise clients with isolated compute resources for model hosting. This capability enables engineering teams to manage independent scaling, model selection, and testing environments for high-traffic coding agents.
➔ Deployment of Dedicated Model Inference (DMI) providing isolated compute resources, direct scaling control, and independent testing environments.
product launchSep 16, 2026
Migrating from closed to open source models, Together
Together AI has introduced a structured five-stage framework to facilitate the migration of enterprise workloads from proprietary closed-source models to open-source alternatives. The playbook defines a standardized methodology covering discovery, evaluation, adaptation, decision-making, and production deployment.
➔ Availability of a structured five-stage migration playbook for transitioning to open-source models via the Together AI platform.
product launchSep 14, 2026
Together AI Introduces GLM-5.3 Flash Model with 17x Cost Reduction
Together AI has launched the GLM-5.3 Flash model, which provides a 17x reduction in cost compared to the standard GLM-5.3. The model maintains high performance with only a 5.6 point drop in pass@1 accuracy.
➔ Availability of GLM-5.3 Flash, offering 17x lower cost at the expense of 5.6 points pass@1 accuracy.
featureSep 11, 2026
Together AI expands fine-tuning service with more models, live metrics, and finer controls
Together AI has integrated new open-weight models and introduced granular training features including Expert LoRA, early stopping, and live experiment tracking. The update also implements pre-flight validation and tokenized dataset previews alongside reduced pricing for specific models.
➔ Service now includes Expert LoRA, live experiment tracking, early stopping, pre-flight validation, tokenized dataset previews, and reduced pricing on selected models.