Back to Live Feed
Together AI logo

Together AI

Verified
Artificial Intelligence • USA
Official Site

Cloud platform for training and running open-source AI models.

Tracked Changes
16
Pricing Shifts
1
Features & Launches
15
Last Verified Event
Sep 23, 2026

Together AI Pricing History

Historical price tier adjustments and billing model changes.

1 Price Changes
Was: Together GPU Clusters were only available at standard on-demand pricing rates.
Now: Together GPU Clusters now offer preemptible compute instances at 50% of the on-demand rate with a five-minute termination notice.

Together AI Chronological Timeline

2026
product launchSep 23, 2026
96% Verified

How to train your own Jev for $17

Together AI released the together/Tev1-4B-experimental classifier model based on Qwen3.5 4B. The company provided a technical workflow for users to fine-tune custom versions of this model on their serverless platform for a cost of $17.

Before: Users lacked a documented, low-cost path for fine-tuning Jev-like classifiers on the Together AI serverless platform.
After: Availability of the together/Tev1-4B-experimental model and a verified $17 fine-tuning workflow.
View full change record & proof ➔
featureSep 22, 2026
96% Verified

Canary rollouts: upgrade models in production without downtime

A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference....

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Canary rollouts: upgrade models in production without downtime.
View full change record & proof ➔
product launchSep 18, 2026
96% Verified

How a global fintech scaled coding agent traffic with Dedicated Model Inference

Together AI has introduced Dedicated Model Inference (DMI) to provide enterprise clients with isolated compute resources for model hosting. This capability enables engineering teams to manage independent scaling, model selection, and testing environments for high-traffic coding agents.

Before: Reliance on shared, multi-tenant API endpoints with limited control over infrastructure scaling and model versioning.
After: Deployment of Dedicated Model Inference (DMI) providing isolated compute resources, direct scaling control, and independent testing environments.
View full change record & proof ➔
product launchSep 16, 2026
96% Verified

Migrating from closed to open source models, Together

Together AI has introduced a structured five-stage framework to facilitate the migration of enterprise workloads from proprietary closed-source models to open-source alternatives. The playbook defines a standardized methodology covering discovery, evaluation, adaptation, decision-making, and production deployment.

Before: Lack of a standardized, documented methodology for migrating enterprise applications from closed-source to open-source model architectures.
After: Availability of a structured five-stage migration playbook for transitioning to open-source models via the Together AI platform.
View full change record & proof ➔
product launchSep 14, 2026
92% Verified

Together AI Introduces GLM-5.3 Flash Model with 17x Cost Reduction

Together AI has launched the GLM-5.3 Flash model, which provides a 17x reduction in cost compared to the standard GLM-5.3. The model maintains high performance with only a 5.6 point drop in pass@1 accuracy.

Before: Access limited to standard GLM-5.3 model architecture.
After: Availability of GLM-5.3 Flash, offering 17x lower cost at the expense of 5.6 points pass@1 accuracy.
View full change record & proof ➔
featureSep 11, 2026
96% Verified

Together AI expands fine-tuning service with more models, live metrics, and finer controls

Together AI has integrated new open-weight models and introduced granular training features including Expert LoRA, early stopping, and live experiment tracking. The update also implements pre-flight validation and tokenized dataset previews alongside reduced pricing for specific models.

Before: Fine-tuning service lacked native live experiment tracking, early stopping, pre-flight validation, and Expert LoRA support.
After: Service now includes Expert LoRA, live experiment tracking, early stopping, pre-flight validation, tokenized dataset previews, and reduced pricing on selected models.
View full change record & proof ➔
pricingSep 10, 2026
96% Verified

Introducing preemptible compute: the same compute, half the price

Together AI has introduced preemptible compute instances for its GPU clusters. This new offering provides identical GPU capacity at a 50% discount compared to on-demand rates, subject to a five-minute drain window.

Before: Together GPU Clusters were only available at standard on-demand pricing rates.
After: Together GPU Clusters now offer preemptible compute instances at 50% of the on-demand rate with a five-minute termination notice.
View full change record & proof ➔
featureSep 10, 2026
96% Verified

To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

Together AI has ported the ThunderKittens library to support the NVIDIA Vera Rubin NVL72 architecture. The implementation includes a rebuilt NVFP4 GEMM kernel that achieves over 22 PFLOPS performance.

Before: ThunderKittens lacked native support and optimized NVFP4 GEMM kernels for the Vera Rubin NVL72 architecture.
After: ThunderKittens supports NVIDIA Vera Rubin NVL72 with a rebuilt NVFP4 GEMM kernel delivering over 22 PFLOPS.
View full change record & proof ➔
product launchSep 9, 2026
96% Verified

The Open Source AI Stack

Together AI has released a comprehensive open-source AI stack designed to streamline the deployment and fine-tuning of large language models. The stack integrates optimized inference engines, data processing pipelines, and model training frameworks to reduce latency and infrastructure overhead.

Before: Users relied on closed-source, proprietary API endpoints for model inference and lacked integrated tools for end-to-end open-source model management.
After: Users have access to a unified, open-source stack for deploying, fine-tuning, and serving open-weights models with optimized performance.
View full change record & proof ➔
featureSep 5, 2026
97% Verified

Together AI Rolls Out Next-Generation Platform Capabilities & API Architecture

Together AI released substantial architectural upgrades, introducing enhanced API integrations and updated workflow tooling for its Artificial Intelligence ecosystem.

Before: Previous Together AI integration tier and workflow limitations.
After: Modernized Together AI platform capabilities with enhanced integration tooling.
View full change record & proof ➔
product launchAug 28, 2026
96% Verified

GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing

Together AI has introduced GLM-5.3 Flash as a high-efficiency alternative to the standard GLM-5.3 model. The new model achieves a 17x reduction in cost while maintaining performance within 5.6 points of pass@1 and 2.6 points of pass@4 on the DeepSWE benchmark.

Before: Users were limited to the standard GLM-5.3 model for DeepSWE coding tasks.
After: Availability of GLM-5.3 Flash, offering a 17x cost reduction compared to GLM-5.3 with a 5.6 point pass@1 performance delta.
View full change record & proof ➔
featureAug 21, 2026
96% Verified

GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Together AI evaluated GLM-5.3 and GPT-5.6 Sol on 904 DeepSWE rollouts, establishing performance benchmarks for coding tasks. The analysis confirms GLM-5.3 achieves higher pass@4 efficiency at 50% of the cost compared to GPT-5.6 Sol.

Before: Lack of comparative performance and cost-efficiency data for GLM-5.3 and GPT-5.6 Sol on the DeepSWE benchmark.
After: Established performance metrics showing GLM-5.3 at 50% cost of GPT-5.6 Sol for pass@4 tasks and an 85.9% success rate for GLM-first routing cascades.
View full change record & proof ➔
product launchAug 21, 2026
96% Verified

GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

Together AI has introduced the GLM-5.3 model, which achieves parity with Claude Fable 5 on pass@1 benchmarks for DeepSWE. The model demonstrates superior pass@4 performance while reducing operational costs to $3.99 per rollout compared to $21.63 for Claude Fable 5.

Before: Claude Fable 5 was the primary benchmark for DeepSWE rollouts at a cost of $21.63 per unit.
After: GLM-5.3 is available for DeepSWE tasks at $3.99 per rollout with improved pass@4 performance.
View full change record & proof ➔
product launchAug 18, 2026
96% Verified

DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Together AI evaluated the performance of DeepSeek V4 Pro 0813 and GPT-5.6 Sol using 904 DeepSWE rollouts. The analysis confirms that GPT-5.6 Sol leads in pass@1 accuracy while DeepSeek V4 Pro 0813 excels in pass@4 and cost-efficiency.

Before: Lack of comparative performance data for DeepSeek V4 Pro 0813 and GPT-5.6 Sol on the DeepSWE benchmark.
After: Availability of comparative performance metrics and a validated 83.0% success rate for a Pro-first routing cascade.
View full change record & proof ➔
product launchAug 17, 2026
96% Verified

DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

Together AI released performance benchmarks for DeepSeek V4 Pro 0813 and Claude Fable 5 on the DeepSWE coding benchmark. The data establishes a cost-performance trade-off where Claude Fable 5 leads in pass@1 accuracy while DeepSeek V4 Pro 0813 excels in pass@4 and cascade routing efficiency.

Before: Lack of comparative performance and cost-efficiency data for DeepSeek V4 Pro 0813 and Claude Fable 5 on the DeepSWE benchmark.
After: Availability of benchmark data confirming Claude Fable 5's pass@1 lead and DeepSeek V4 Pro 0813's pass@4 performance, alongside an 82.7% success rate for Pro-first cascade routing.
View full change record & proof ➔
featureAug 17, 2026
96% Verified

A/B test models in production

Together AI has introduced native A/B testing capabilities directly at the API endpoint level. This feature allows developers to route traffic between different model versions without modifying application-side routing logic.

Before: Developers were required to implement custom routing logic within their own application code to distribute traffic across different model endpoints.
After: Developers can configure traffic splits directly at the Together AI endpoint, enabling native A/B testing of models without application-level code changes.
View full change record & proof ➔

Complete Together AI Change Log Index

DateChange TitleTypeImpactDetails
Sep 23, 2026How to train your own Jev for $17product_launch9/10View ➔
Sep 22, 2026Canary rollouts: upgrade models in production without downtimefeature8/10View ➔
Sep 18, 2026How a global fintech scaled coding agent traffic with Dedicated Model Inferenceproduct_launch8/10View ➔
Sep 16, 2026Migrating from closed to open source models, Togetherproduct_launch7/10View ➔
Sep 14, 2026Together AI Introduces GLM-5.3 Flash Model with 17x Cost Reductionproduct_launch7/10View ➔
Sep 11, 2026Together AI expands fine-tuning service with more models, live metrics, and finer controlsfeature8/10View ➔
Sep 10, 2026Introducing preemptible compute: the same compute, half the pricepricing8/10View ➔
Sep 10, 2026To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!feature9/10View ➔
Sep 9, 2026The Open Source AI Stackproduct_launch9/10View ➔
Sep 5, 2026Together AI Rolls Out Next-Generation Platform Capabilities & API Architecturefeature8/10View ➔
Aug 28, 2026GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routingproduct_launch9/10View ➔
Aug 21, 2026GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routingfeature8/10View ➔
Aug 21, 2026GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routingproduct_launch9/10View ➔
Aug 18, 2026DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routingproduct_launch9/10View ➔
Aug 17, 2026DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routingproduct_launch9/10View ➔
Aug 17, 2026A/B test models in productionfeature8/10View ➔
product_launchSep 23, 2026
How to train your own Jev for $17
Impact: 9/10View Record
featureSep 22, 2026
Canary rollouts: upgrade models in production without downtime
Impact: 8/10View Record
product_launchSep 18, 2026
How a global fintech scaled coding agent traffic with Dedicated Model Inference
Impact: 8/10View Record
product_launchSep 16, 2026
Migrating from closed to open source models, Together
Impact: 7/10View Record
product_launchSep 14, 2026
Together AI Introduces GLM-5.3 Flash Model with 17x Cost Reduction
Impact: 7/10View Record
featureSep 11, 2026
Together AI expands fine-tuning service with more models, live metrics, and finer controls
Impact: 8/10View Record
pricingSep 10, 2026
Introducing preemptible compute: the same compute, half the price
Impact: 8/10View Record
featureSep 10, 2026
To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
Impact: 9/10View Record
product_launchSep 9, 2026
The Open Source AI Stack
Impact: 9/10View Record
featureSep 5, 2026
Together AI Rolls Out Next-Generation Platform Capabilities & API Architecture
Impact: 8/10View Record
product_launchAug 28, 2026
GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing
Impact: 9/10View Record
featureAug 21, 2026
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
Impact: 8/10View Record
product_launchAug 21, 2026
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
Impact: 9/10View Record
product_launchAug 18, 2026
DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
Impact: 9/10View Record
product_launchAug 17, 2026
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
Impact: 9/10View Record
featureAug 17, 2026
A/B test models in production
Impact: 8/10View Record