Back to Live Feed
Hugging Face logo

Hugging Face

Verified
Artificial Intelligence • USA / France
Official Site

The AI community building the future of open models and datasets.

Tracked Changes
25
Pricing Shifts
0
Features & Launches
25
Last Verified Event
Oct 1, 2026

Hugging Face Chronological Timeline

2026
featureOct 1, 2026
96% Verified

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Hugging Face announced Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs.
View full change record & proof ➔
product launchSep 30, 2026
96% Verified

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Hugging Face has launched a centralized leaderboard specifically for benchmarking Text-to-Speech (TTS) and voice cloning models. The platform provides standardized evaluation metrics for multilingual speech synthesis performance.

Before: Lack of a centralized, standardized, and public benchmark for comparing diverse open-source TTS and voice cloning models.
After: Availability of a public, scalable leaderboard for objective, comparative evaluation of multilingual TTS and voice cloning model performance.
View full change record & proof ➔
product launchSep 29, 2026
96% Verified

NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

NVIDIA has released Kumo Tabular, a specialized machine learning framework designed to optimize tabular data prediction tasks. The model architecture achieves state-of-the-art performance by balancing predictive accuracy with computational efficiency.

Before: Tabular prediction tasks relied on traditional gradient-boosted decision trees or standard deep learning architectures that often lacked the efficiency-to-accuracy ratio required for large-scale enterprise deployment.
After: The availability of the Kumo Tabular framework provides a dedicated, high-efficiency model architecture specifically tuned for tabular data prediction.
View full change record & proof ➔
featureSep 29, 2026
96% Verified

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

Hugging Face has introduced a source-aware verification framework designed to integrate with Model Context Protocol (MCP) agents. This system enables agents to cross-reference generated claims against verified source documents to reduce hallucinations.

Before: MCP agents relied on internal model weights or standard RAG without explicit, automated source-verification loops.
After: MCP agents now utilize a source-aware verification framework to validate claims against specific source documents.
View full change record & proof ➔
featureSep 28, 2026
97% Verified

Hugging Face Releases Inference Endpoints with 1-Click Serverless GPU Auto-Scaling

Hugging Face updated its Inference Endpoints to support automatic scale-to-zero serverless deployment for any open model on the Hub.

Before: Inference endpoints required paying continuous hourly rates for idle GPUs.
After: Scale-to-zero serverless endpoints billing exclusively for active query execution seconds.
View full change record & proof ➔
product launchSep 28, 2026
96% Verified

Holo4: powering generalist computer-use agents

Hugging Face has released Holo4, a specialized model architecture designed to enable generalist computer-use agents. The model provides native capabilities for interpreting and interacting with graphical user interfaces across diverse operating systems.

Before: Agents were primarily restricted to text-based API interactions or required brittle, heuristic-based screen scraping tools.
After: Holo4 provides a dedicated model architecture for generalist computer-use, enabling agents to interpret and manipulate GUI elements directly.
View full change record & proof ➔
featureSep 24, 2026
96% Verified

Accelerating vision-language models with LFM2.5-VL-DSpark

Hugging Face announced Accelerating vision-language models with LFM2.5-VL-DSpark.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Accelerating vision-language models with LFM2.5-VL-DSpark.
View full change record & proof ➔
featureSep 23, 2026
96% Verified

How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows

Hugging Face has integrated NVIDIA Warp and MjWarp into its robotics ecosystem to enable high-performance GPU-accelerated simulation. This integration allows developers to execute physics kernels directly in Python while maintaining C++ performance levels for robotics learning workflows.

Before: Robotics simulation workflows were primarily CPU-bound or required complex, non-native C++ extensions to achieve high-throughput parallelization.
After: Native support for NVIDIA Warp and MjWarp allows for GPU-accelerated physics simulation kernels directly within Python-based robotics pipelines.
View full change record & proof ➔
featureSep 23, 2026
96% Verified

**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**

Hugging Face announced **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**.
View full change record & proof ➔
featureSep 22, 2026
96% Verified

Transformers now runs llama.cpp quants

The Hugging Face Transformers library now natively supports loading and running GGUF-formatted quantized models via the llama.cpp backend. This integration allows users to execute compressed models directly within the Transformers ecosystem without requiring external conversion tools.

Before: Transformers required full-precision models or specific AutoGPTQ/AutoAWQ integrations, lacking native support for GGUF-formatted quantized files.
After: Transformers natively supports loading GGUF-formatted models, enabling direct inference of llama.cpp quantized weights through the standard library interface.
View full change record & proof ➔
featureSep 22, 2026
96% Verified

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Hugging Face has integrated the UK AISI's Inspect framework with the EvalEval platform to standardize model evaluation protocols. This integration enables automated, reproducible execution of safety and capability benchmarks for large language models.

Before: Benchmark results were often non-reproducible due to fragmented evaluation environments and inconsistent execution scripts across different research teams.
After: Standardized, reproducible benchmark execution via the integration of the Inspect framework into the EvalEval platform.
View full change record & proof ➔
featureSep 22, 2026
96% Verified

Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community

Hugging Face has hired Jun Kim, the creator of the oMLX library, to formalize support for the MLX framework within its ecosystem. This move signals a strategic commitment to optimizing Apple Silicon-based machine learning workflows.

Before: MLX support on Hugging Face was community-driven and lacked dedicated internal engineering resources for framework-specific optimization.
After: Hugging Face now employs the primary maintainer of oMLX to provide dedicated support, integration, and optimization for the MLX framework.
View full change record & proof ➔
featureSep 21, 2026
96% Verified

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Hugging Face introduced a novel model pruning methodology that treats transformer block removal as an Ising optimization problem. This approach utilizes physical system modeling to identify and remove redundant layers while minimizing performance degradation.

Before: LLM pruning relied on heuristic-based layer removal or computationally expensive retraining processes.
After: LLM pruning is now supported by an Ising optimization framework for systematic, physics-based block removal.
View full change record & proof ➔
product launchSep 21, 2026
96% Verified

tokenizers v1: encode, decode and scaling, measured

Hugging Face has released version 1.0 of the 'tokenizers' library, marking a transition to a stable API. This release focuses on performance optimizations for encoding and decoding processes and establishes long-term API stability.

Before: The library was in a pre-v1.0 state, implying potential breaking API changes and ongoing architectural flux.
After: The library is now at v1.0, providing a stable, production-ready API with optimized encoding and decoding performance.
View full change record & proof ➔
featureSep 15, 2026
96% Verified

Your Agent Aced the Task. Will It Do It Again?

Hugging Face announced Your Agent Aced the Task. Will It Do It Again?.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Your Agent Aced the Task. Will It Do It Again?.
View full change record & proof ➔
featureSep 10, 2026
96% Verified

Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

Hugging Face has introduced an asynchronous implementation of Group Relative Policy Optimization (GRPO) that supports LoRA fine-tuning across distributed HF Jobs. The architecture eliminates the requirement for NCCL (NVIDIA Collective Communications Library) by utilizing an S3-compatible bucket and a proxy for state synchronization.

Before: GRPO training required synchronous NCCL-based communication, necessitating high-bandwidth, low-latency interconnects and complex cluster configuration.
After: GRPO training can now be performed asynchronously using LoRA across HF Jobs, utilizing an S3-compatible bucket and proxy for state management without requiring NCCL.
View full change record & proof ➔
featureSep 10, 2026
96% Verified

Rebuilding AUTOMATIC1111 with Gradio Workflow

Hugging Face has integrated the AUTOMATIC1111 Stable Diffusion web UI into a native Gradio-based workflow. This transition replaces legacy interface components with Gradio's modular blocks to improve UI responsiveness and component extensibility.

Before: The AUTOMATIC1111 interface relied on a custom, non-standardized frontend implementation that hindered modular extension and UI responsiveness.
After: The interface is now rebuilt using Gradio blocks, enabling standardized component interaction, improved state management, and easier integration with the broader Hugging Face Gradio ecosystem.
View full change record & proof ➔
product launchSep 9, 2026
96% Verified

IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license

IBM has released the Granite Time Series PatchTST-FM-r2 model, a foundation model specifically optimized for time-series forecasting. The model is distributed under the Apache 2.0 license, enabling unrestricted commercial use.

Before: Time-series forecasting models were often restricted by non-commercial licenses or required extensive training from scratch on proprietary data.
After: Availability of the pre-trained Granite Time Series PatchTST-FM-r2 model under an Apache 2.0 license for commercial deployment.
View full change record & proof ➔
featureSep 8, 2026
96% Verified

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Hugging Face has introduced a refined safety filtering methodology that enables models to distinguish between harmful and benign sub-topics within a broader category. This approach replaces blanket topic refusals with granular classification to improve model utility while maintaining safety guardrails.

Before: Models utilized broad, binary safety filters that triggered total refusals for entire sensitive topics.
After: Models utilize granular safety filters capable of distinguishing between harmful and benign sub-topics, allowing for partial topic responses.
View full change record & proof ➔
featureSep 3, 2026
96% Verified

NeoMME: an efficient Multimodal-native and Multilingual Encoder

Hugging Face announced NeoMME: an efficient Multimodal-native and Multilingual Encoder.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with NeoMME: an efficient Multimodal-native and Multilingual Encoder.
View full change record & proof ➔
featureSep 3, 2026
96% Verified

Give Your Coding Agents a Memory You Own

Hugging Face announced Give Your Coding Agents a Memory You Own.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Give Your Coding Agents a Memory You Own.
View full change record & proof ➔
featureSep 3, 2026
96% Verified

Training a coding model to paint watercolours with TRL and OpenEnv

Hugging Face announced Training a coding model to paint watercolours with TRL and OpenEnv.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Training a coding model to paint watercolours with TRL and OpenEnv.
View full change record & proof ➔
featureSep 3, 2026
96% Verified

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

Hugging Face announced Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps.
View full change record & proof ➔
featureSep 2, 2026
96% Verified

Real-Time Intelligence with IBM Time Series Models on Confluent

Hugging Face announced Real-Time Intelligence with IBM Time Series Models on Confluent.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with Real-Time Intelligence with IBM Time Series Models on Confluent.
View full change record & proof ➔
featureSep 1, 2026
96% Verified

BenchMIRT: What are LLM benchmarks actually measuring?

Hugging Face announced BenchMIRT: What are LLM benchmarks actually measuring?.

Before: Previous platform capabilities and architecture.
After: Updated platform deployment with BenchMIRT: What are LLM benchmarks actually measuring?.
View full change record & proof ➔

Complete Hugging Face Change Log Index

DateChange TitleTypeImpactDetails
Oct 1, 2026Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEsfeature8/10View ➔
Sep 30, 2026Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloningproduct_launch8/10View ➔
Sep 29, 2026NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Predictionproduct_launch8/10View ➔
Sep 29, 2026Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agentsfeature8/10View ➔
Sep 28, 2026Hugging Face Releases Inference Endpoints with 1-Click Serverless GPU Auto-Scalingfeature8/10View ➔
Sep 28, 2026Holo4: powering generalist computer-use agentsproduct_launch9/10View ➔
Sep 24, 2026Accelerating vision-language models with LFM2.5-VL-DSparkfeature8/10View ➔
Sep 23, 2026How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflowsfeature8/10View ➔
Sep 23, 2026**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**feature8/10View ➔
Sep 22, 2026Transformers now runs llama.cpp quantsfeature9/10View ➔
Sep 22, 2026How UK AISI and EvalEval Are Making Benchmark Results Reproduciblefeature8/10View ➔
Sep 22, 2026Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX communityfeature7/10View ➔
Sep 21, 2026Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problemfeature8/10View ➔
Sep 21, 2026tokenizers v1: encode, decode and scaling, measuredproduct_launch9/10View ➔
Sep 15, 2026Your Agent Aced the Task. Will It Do It Again?feature8/10View ➔
Sep 10, 2026Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCLfeature8/10View ➔
Sep 10, 2026Rebuilding AUTOMATIC1111 with Gradio Workflowfeature7/10View ➔
Sep 9, 2026IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly licenseproduct_launch9/10View ➔
Sep 8, 2026Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topicfeature8/10View ➔
Sep 3, 2026NeoMME: an efficient Multimodal-native and Multilingual Encoderfeature8/10View ➔
Sep 3, 2026Give Your Coding Agents a Memory You Ownfeature8/10View ➔
Sep 3, 2026Training a coding model to paint watercolours with TRL and OpenEnvfeature8/10View ➔
Sep 3, 2026Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Stepsfeature8/10View ➔
Sep 2, 2026Real-Time Intelligence with IBM Time Series Models on Confluentfeature8/10View ➔
Sep 1, 2026BenchMIRT: What are LLM benchmarks actually measuring?feature8/10View ➔
featureOct 1, 2026
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Impact: 8/10View Record
product_launchSep 30, 2026
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Impact: 8/10View Record
product_launchSep 29, 2026
NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
Impact: 8/10View Record
featureSep 29, 2026
Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Impact: 8/10View Record
featureSep 28, 2026
Hugging Face Releases Inference Endpoints with 1-Click Serverless GPU Auto-Scaling
Impact: 8/10View Record
product_launchSep 28, 2026
Holo4: powering generalist computer-use agents
Impact: 9/10View Record
featureSep 24, 2026
Accelerating vision-language models with LFM2.5-VL-DSpark
Impact: 8/10View Record
featureSep 23, 2026
How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows
Impact: 8/10View Record
featureSep 23, 2026
**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**
Impact: 8/10View Record
featureSep 22, 2026
Transformers now runs llama.cpp quants
Impact: 9/10View Record
featureSep 22, 2026
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Impact: 8/10View Record
featureSep 22, 2026
Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community
Impact: 7/10View Record
featureSep 21, 2026
Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
Impact: 8/10View Record
product_launchSep 21, 2026
tokenizers v1: encode, decode and scaling, measured
Impact: 9/10View Record
featureSep 15, 2026
Your Agent Aced the Task. Will It Do It Again?
Impact: 8/10View Record
featureSep 10, 2026
Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
Impact: 8/10View Record
featureSep 10, 2026
Rebuilding AUTOMATIC1111 with Gradio Workflow
Impact: 7/10View Record
product_launchSep 9, 2026
IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license
Impact: 9/10View Record
featureSep 8, 2026
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
Impact: 8/10View Record
featureSep 3, 2026
NeoMME: an efficient Multimodal-native and Multilingual Encoder
Impact: 8/10View Record
featureSep 3, 2026
Give Your Coding Agents a Memory You Own
Impact: 8/10View Record
featureSep 3, 2026
Training a coding model to paint watercolours with TRL and OpenEnv
Impact: 8/10View Record
featureSep 3, 2026
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Impact: 8/10View Record
featureSep 2, 2026
Real-Time Intelligence with IBM Time Series Models on Confluent
Impact: 8/10View Record
featureSep 1, 2026
BenchMIRT: What are LLM benchmarks actually measuring?
Impact: 8/10View Record