Hugging Face
VerifiedThe AI community building the future of open models and datasets.
Hugging Face Chronological Timeline
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Hugging Face announced Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs.
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Hugging Face has launched a centralized leaderboard specifically for benchmarking Text-to-Speech (TTS) and voice cloning models. The platform provides standardized evaluation metrics for multilingual speech synthesis performance.
NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
NVIDIA has released Kumo Tabular, a specialized machine learning framework designed to optimize tabular data prediction tasks. The model architecture achieves state-of-the-art performance by balancing predictive accuracy with computational efficiency.
Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Hugging Face has introduced a source-aware verification framework designed to integrate with Model Context Protocol (MCP) agents. This system enables agents to cross-reference generated claims against verified source documents to reduce hallucinations.
Hugging Face Releases Inference Endpoints with 1-Click Serverless GPU Auto-Scaling
Hugging Face updated its Inference Endpoints to support automatic scale-to-zero serverless deployment for any open model on the Hub.
Holo4: powering generalist computer-use agents
Hugging Face has released Holo4, a specialized model architecture designed to enable generalist computer-use agents. The model provides native capabilities for interpreting and interacting with graphical user interfaces across diverse operating systems.
Accelerating vision-language models with LFM2.5-VL-DSpark
Hugging Face announced Accelerating vision-language models with LFM2.5-VL-DSpark.
How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows
Hugging Face has integrated NVIDIA Warp and MjWarp into its robotics ecosystem to enable high-performance GPU-accelerated simulation. This integration allows developers to execute physics kernels directly in Python while maintaining C++ performance levels for robotics learning workflows.
**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**
Hugging Face announced **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**.
Transformers now runs llama.cpp quants
The Hugging Face Transformers library now natively supports loading and running GGUF-formatted quantized models via the llama.cpp backend. This integration allows users to execute compressed models directly within the Transformers ecosystem without requiring external conversion tools.
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Hugging Face has integrated the UK AISI's Inspect framework with the EvalEval platform to standardize model evaluation protocols. This integration enables automated, reproducible execution of safety and capability benchmarks for large language models.
Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community
Hugging Face has hired Jun Kim, the creator of the oMLX library, to formalize support for the MLX framework within its ecosystem. This move signals a strategic commitment to optimizing Apple Silicon-based machine learning workflows.
Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
Hugging Face introduced a novel model pruning methodology that treats transformer block removal as an Ising optimization problem. This approach utilizes physical system modeling to identify and remove redundant layers while minimizing performance degradation.
tokenizers v1: encode, decode and scaling, measured
Hugging Face has released version 1.0 of the 'tokenizers' library, marking a transition to a stable API. This release focuses on performance optimizations for encoding and decoding processes and establishes long-term API stability.
Your Agent Aced the Task. Will It Do It Again?
Hugging Face announced Your Agent Aced the Task. Will It Do It Again?.
Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
Hugging Face has introduced an asynchronous implementation of Group Relative Policy Optimization (GRPO) that supports LoRA fine-tuning across distributed HF Jobs. The architecture eliminates the requirement for NCCL (NVIDIA Collective Communications Library) by utilizing an S3-compatible bucket and a proxy for state synchronization.
Rebuilding AUTOMATIC1111 with Gradio Workflow
Hugging Face has integrated the AUTOMATIC1111 Stable Diffusion web UI into a native Gradio-based workflow. This transition replaces legacy interface components with Gradio's modular blocks to improve UI responsiveness and component extensibility.
IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license
IBM has released the Granite Time Series PatchTST-FM-r2 model, a foundation model specifically optimized for time-series forecasting. The model is distributed under the Apache 2.0 license, enabling unrestricted commercial use.
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
Hugging Face has introduced a refined safety filtering methodology that enables models to distinguish between harmful and benign sub-topics within a broader category. This approach replaces blanket topic refusals with granular classification to improve model utility while maintaining safety guardrails.
NeoMME: an efficient Multimodal-native and Multilingual Encoder
Hugging Face announced NeoMME: an efficient Multimodal-native and Multilingual Encoder.
Give Your Coding Agents a Memory You Own
Hugging Face announced Give Your Coding Agents a Memory You Own.
Training a coding model to paint watercolours with TRL and OpenEnv
Hugging Face announced Training a coding model to paint watercolours with TRL and OpenEnv.
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Hugging Face announced Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps.
Real-Time Intelligence with IBM Time Series Models on Confluent
Hugging Face announced Real-Time Intelligence with IBM Time Series Models on Confluent.
BenchMIRT: What are LLM benchmarks actually measuring?
Hugging Face announced BenchMIRT: What are LLM benchmarks actually measuring?.
Complete Hugging Face Change Log Index
| Date | Change Title | Type | Impact | Details |
|---|---|---|---|---|
| Oct 1, 2026 | Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | feature | 8/10 | View ➔ |
| Sep 30, 2026 | Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning | product_launch | 8/10 | View ➔ |
| Sep 29, 2026 | NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction | product_launch | 8/10 | View ➔ |
| Sep 29, 2026 | Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents | feature | 8/10 | View ➔ |
| Sep 28, 2026 | Hugging Face Releases Inference Endpoints with 1-Click Serverless GPU Auto-Scaling | feature | 8/10 | View ➔ |
| Sep 28, 2026 | Holo4: powering generalist computer-use agents | product_launch | 9/10 | View ➔ |
| Sep 24, 2026 | Accelerating vision-language models with LFM2.5-VL-DSpark | feature | 8/10 | View ➔ |
| Sep 23, 2026 | How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows | feature | 8/10 | View ➔ |
| Sep 23, 2026 | **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** | feature | 8/10 | View ➔ |
| Sep 22, 2026 | Transformers now runs llama.cpp quants | feature | 9/10 | View ➔ |
| Sep 22, 2026 | How UK AISI and EvalEval Are Making Benchmark Results Reproducible | feature | 8/10 | View ➔ |
| Sep 22, 2026 | Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community | feature | 7/10 | View ➔ |
| Sep 21, 2026 | Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem | feature | 8/10 | View ➔ |
| Sep 21, 2026 | tokenizers v1: encode, decode and scaling, measured | product_launch | 9/10 | View ➔ |
| Sep 15, 2026 | Your Agent Aced the Task. Will It Do It Again? | feature | 8/10 | View ➔ |
| Sep 10, 2026 | Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL | feature | 8/10 | View ➔ |
| Sep 10, 2026 | Rebuilding AUTOMATIC1111 with Gradio Workflow | feature | 7/10 | View ➔ |
| Sep 9, 2026 | IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license | product_launch | 9/10 | View ➔ |
| Sep 8, 2026 | Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic | feature | 8/10 | View ➔ |
| Sep 3, 2026 | NeoMME: an efficient Multimodal-native and Multilingual Encoder | feature | 8/10 | View ➔ |
| Sep 3, 2026 | Give Your Coding Agents a Memory You Own | feature | 8/10 | View ➔ |
| Sep 3, 2026 | Training a coding model to paint watercolours with TRL and OpenEnv | feature | 8/10 | View ➔ |
| Sep 3, 2026 | Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps | feature | 8/10 | View ➔ |
| Sep 2, 2026 | Real-Time Intelligence with IBM Time Series Models on Confluent | feature | 8/10 | View ➔ |
| Sep 1, 2026 | BenchMIRT: What are LLM benchmarks actually measuring? | feature | 8/10 | View ➔ |