
Work Here?
Hugging Face provides tools and platforms for building and sharing machine learning applications. Its core offering is the Hugging Face Hub, where developers and researchers share, discover, and collaborate on models, datasets, and applications; users access pre-trained models via the Transformers library and deploy them with services like Inference Endpoints or Private Hub. The company stands out through its large open-source community, vast collections of models and datasets, and tight integrations with cloud providers. Its goal is to democratize machine learning by making advanced AI accessible to individuals and organizations alike.
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
1,001-5,000
Company Stage
Acquired
Total Funding
$395.7M
Headquarters
New York City, New York
Founded
2016
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$395.7M
Above
Industry Average
Funded Over
8 Rounds
Flexible Work Environment
Health Insurance
Unlimited PTO
Equity
Growth, Training, & Conferences
Generous Parental Leave
Hugging Face and ServiceNow introduce AutoSynthData for enterprise AI training. Hugging Face and ServiceNow have partnered to launch AutoSynthData, a new framework designed to automate the creation of high-quality synthetic training data for enterprise AI agents, addressing a critical bottleneck in AI development. Published October 1, 2026 A significant challenge in deploying enterprise-grade AI agents lies in acquiring sufficient volumes of high-quality, task-specific training data. Traditional methods of data collection and annotation are often time-consuming, expensive, and raise privacy concerns, especially with sensitive corporate information. To address this persistent hurdle, Hugging Face and ServiceNow have collaborated to introduce AutoSynthData, a novel framework aimed at automating the generation of synthetic datasets tailored for enterprise applications. This initiative promises to accelerate AI development cycles and enhance the efficacy of AI solutions across various business functions. Bridging the data gap with synthetic generation. AutoSynthData is engineered to alleviate the dependency on vast quantities of proprietary, real-world data by generating synthetic alternatives. The core concept involves leveraging a small initial set of real data or expert-defined rules to guide a Language Model (LLM) in producing diverse and realistic synthetic examples. This approach is particularly valuable for enterprise environments where data privacy regulations and the scarcity of labeled datasets can severely impede AI project progress. By providing a scalable and privacy-preserving method for data acquisition, AutoSynthData enables companies to train more robust and specialized AI agents without compromising sensitive information. How AutoSynthData operates. The framework operates through a multi-stage process. Initially, it requires either a small seed dataset of real examples or a set of well-defined rules that describe the characteristics of the desired data. An LLM, acting as the generative engine, then synthesizes new data points based on these inputs. A crucial aspect of AutoSynthData is its iterative refinement mechanism, which includes a feedback loop to enhance data quality and relevance. The generated synthetic data undergoes evaluation against predefined metrics or human review, with insights used to fine-tune the generative process, ensuring the output closely mimics real-world data distributions and meets task-specific requirements. This iterative improvement is vital for maintaining high data fidelity, which is essential for effective model training. Key benefits for enterprise AI development. AutoSynthData offers several compelling advantages for enterprises. Firstly, it drastically reduces the time and cost associated with manual data collection and annotation. Businesses can rapidly generate large datasets tailored to their specific use cases, such as customer service chatbots, internal knowledge management systems, or specialized compliance tools. Secondly, the framework inherently addresses data privacy and security concerns by generating data that does not contain actual sensitive information, making it safer for training models in regulated industries. Lastly, it promotes innovation by lowering the barrier to entry for developing specialized AI applications, allowing organizations to experiment with new AI solutions more readily and efficiently. Broader implications for AI ecosystems. The introduction of AutoSynthData marks a significant step forward in the broader AI ecosystem. It underscores a growing trend toward synthetic data generation as a viable and often superior alternative to traditional data acquisition methods, especially in niche or data-scarce domains. For developers, this means faster prototyping and deployment of AI models. For businesses, it translates into quicker time-to-market for AI-powered products and services, alongside a more secure and compliant approach to AI development. The collaboration between Hugging Face, a leader in AI models and tools, and ServiceNow, a prominent enterprise software provider, highlights the increasing importance of accessible and efficient data solutions for real-world AI implementation. Why it matters. AutoSynthData offers a practical solution to the persistent challenge of data scarcity and privacy in enterprise AI development. By enabling the automated generation of high-quality synthetic training data, the framework empowers organizations to build and deploy more effective and specialized AI agents efficiently and securely. This innovation is poised to accelerate the adoption of AI across various industries, making advanced AI capabilities more attainable for businesses of all sizes.
Olmo-core 3 advances open-source Mixture-of-Experts training. AllenNLP and Hugging Face release Olmo-core 3, enhancing open-source infrastructure for training large Mixture-of-Experts (MoE) models with improved scalability and efficiency. Published October 1, 2026 Advancing open-source MoE training infrastructure. The landscape of large language models (LLMs) is continually evolving, with Mixture-of-Experts (MoE) architectures gaining prominence for their ability to achieve high performance with improved efficiency during inference. Recognizing the need for accessible and scalable training solutions within the open-source community, AllenNLP, in collaboration with Hugging Face, has introduced Olmo-core 3. This new iteration significantly advances the infrastructure available for developing and training large MoE models, offering a robust, open platform designed to streamline the complex process of building these advanced AI systems. Olmo-core 3 aims to democratize access to cutting-edge MoE model development by providing a comprehensive, open-source framework. This initiative is critical as proprietary solutions often limit transparency and community-driven innovation. By open-sourcing the core training infrastructure, AllenNLP and Hugging Face are fostering an environment where researchers and developers can more readily experiment with, understand, and optimize MoE architectures, ultimately accelerating progress in AI. Core innovations in Olmo-core 3. The latest release of Olmo-core introduces several key innovations that address common challenges in training large-scale MoE models. A primary focus has been on improving scalability and efficiency, which are crucial for handling the immense computational demands of these models. Olmo-core 3 implements enhanced parallelism strategies, allowing for more effective distribution of workloads across multiple GPUs and nodes. This is particularly beneficial for MoE models, which inherently involve routing different parts of an input through specialized "expert" sub-networks. One significant improvement lies in its refined support for various MoE configurations. The framework now offers more flexible tools for defining and managing expert layers, gate mechanisms, and load balancing strategies. This allows developers to experiment with different architectural choices and hyperparameter settings more easily, leading to better-performing and more stable MoE models. Furthermore, Olmo-core 3 includes optimized data loading and processing pipelines, minimizing bottlenecks that can hinder training speed and resource utilization. Scalability and efficiency enhancements. Training large MoE models is notoriously resource-intensive, requiring significant computational power and careful management of memory and communication overhead. Olmo-core 3 tackles these challenges head-on through its architectural improvements. The framework now features advanced distributed training primitives that are specifically tailored for MoE workloads. These primitives enable efficient data parallelism, model parallelism, and expert parallelism, ensuring that computational resources are utilized optimally across a cluster. For instance, the enhanced routing mechanisms within Olmo-core 3 minimize unnecessary communication between experts and between nodes, reducing overhead and improving overall training throughput. This also includes more sophisticated load balancing algorithms that dynamically adjust expert assignments to prevent individual experts from becoming overloaded, which can otherwise lead to training inefficiencies and suboptimal model performance. The result is a more scalable and cost-effective solution for developing state-of-the-art MoE models, making advanced AI research more accessible to institutions and individuals with varying resource constraints. Impact on developers and businesses. Olmo-core 3's release has significant implications for both individual developers and larger organizations. For developers, it lowers the barrier to entry for experimenting with and building sophisticated MoE models. The open-source nature means full transparency into the training process, allowing for deeper understanding and customization. This empowers researchers to push the boundaries of MoE architectures without being constrained by proprietary black boxes or the need to develop foundational infrastructure from scratch. Businesses stand to benefit from the improved efficiency and scalability offered by Olmo-core 3. Training MoE models can be costly; by providing a more optimized and openly available infrastructure, the framework can help reduce the financial and technical burden associated with developing advanced AI. This could lead to faster iteration cycles, more efficient resource allocation, and ultimately, the deployment of more powerful and specialized AI applications across various industries, from natural language processing to drug discovery. Why it matters. Olmo-core 3 represents a crucial step forward in the democratization of advanced AI development. By offering a robust, open, and scalable infrastructure for training Mixture-of-Experts models, AllenNLP and Hugging Face are empowering a broader community of researchers and developers. This initiative not only accelerates innovation in MoE architectures but also fosters greater transparency and collaboration within the AI field. As AI models become increasingly complex, open-source tools like Olmo-core 3 are vital for ensuring that progress remains accessible, efficient, and driven by collective effort, ultimately benefiting the entire ecosystem of AI research and application.
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs. Published October 1, 2026 Today Hugging Face is releasing Olmo-core 3, a significant upgrade to its framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It's one of the core systems behind the next generation of Olmo, and part of its ongoing commitment to open up the tools and training infrastructure behind each new model. Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach - they can contain many more learned components, or parameters, without requiring every input to use all of them. But the full model still has to be stored across GPU memory and updated during training, and directing inputs to the right experts - the specialized components within an MoE - across a cluster creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input. Olmo-core 3 is built to close that gap. In one benchmark, Hugging Face increased the expert pool from 8 to 128 while still selecting only four experts per token - the small units of text a language model processes - keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%. The same infrastructure has been benchmarked at over one trillion total parameters. Building a training stack around how MoEs actually work. Olmo-core has evolved with each generation of Olmo. Its work on sparse models goes back to OlmoE, which used an MoE architecture with 64 routed experts. Olmo 3, by contrast, used a dense architecture, meaning nearly all of the model was active for every token and its training stack was built around that design. Olmo-core 3 extends the framework with a training system designed for much larger MoE models. Its earlier MoE implementation in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism (DDP). It keeps experts resident on GPUs and routes the relevant data to them, avoiding that repeated weight gathering. NVIDIA's Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over its earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using its earlier implementation - about 2.7x the throughput. Scaling and optimizing MoE training. Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient. Three techniques determine how the model and its training state are split across hardware: * Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool. * Pipeline parallelism splits the model's layers - the successive stages that transform an input - across groups of GPUs, reducing how much of the model each GPU needs to keep in memory. * A distributed optimizer spreads the optimizer state - the additional data used to calculate and apply updates during training - across GPUs instead of storing a full copy on every GPU. Together, these techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory. Olmo-core 3 also reduces the cost of routing data to the right experts and running their computations. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the extra work needed to rearrange it. GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back. And grouped GEMM combines many small expert computations so GPUs can execute them more efficiently. Finally, Olmo-core 3 supports MXFP8, a lower-precision number format that represents some values with fewer bits. This can reduce computation and the amount of data moved between GPUs, as long as those savings outweigh the cost of converting between number formats. Hugging Face measured MXFP8's effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format Hugging Face used as its baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone. These techniques and optimizations have to work together. Speeding up one part of training can create costs elsewhere; faster computation may require more data movement, while moving fewer bits may not help if converting the data takes too long. Olmo-core 3 is built around those trade-offs across the full training process, giving Hugging Face - and researchers using the open stack - control over how the pieces fit together. Explore its interactive walkthrough to see how data, expert, and pipeline parallelism work together to scale MoE training - from a single GPU to many. Scaling into the trillion-parameter range. Hugging Face has benchmarked Olmo-core 3 across a range of configurations on NVIDIA B300 GPUs, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. Its highest observed throughput was 858 TFLOP/s/GPU - a measure of useful model computation per second on each GPU. These tests used random routing to measure system performance, rather than the quality of a trained model. Hugging Face has also experimented with DeepEP v2, an alternative way of handling communication between experts across GPUs, reaching a configuration with 2.38 trillion total parameters. This was a short-capacity test rather than a full training run, so it demonstrates the scale Olmo-core 3 can reach rather than sustained training performance. At these scales, systems performance is only part of the picture. Its technical report also documents experiments that informed how Hugging Face train MoEs and measure their performance. For example: * A score intended to encourage balanced routing could improve even as the actual workload became less balanced. Hugging Face call this failure token gerrymandering. * Lowering experts' learning rates - the size of their training updates - because they process fewer tokens did not improve results in the model family Hugging Face tested. * GPU calculations took different amounts of time when the values being processed changed, even with the same matrix dimensions. Performance comparisons therefore need matching input values as well as matching shapes. * Overlapping communication and computation on separate GPU streams did not always make training faster. In some tests, it slowed end-to-end execution - a reminder that more overlap does not necessarily mean higher throughput. The report explains these findings alongside the approaches Hugging Face tested and chose not to adopt. Built for the next generation of Olmo, open for everyone. Olmo-core 3 is the foundation for what Hugging Face is building next. Its next-generation Olmo will use an MoE architecture, and Hugging Face is aiming for it to be its most capable Olmo yet, trained on its largest dataset and with its longest context window. The new stack lets Hugging Face scale beyond its previous MoE work while giving Hugging Face more flexibility to adapt training as models and hardware evolve. And it's fully open - researchers and developers can use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other parts of the system. That's part of how Hugging Face think about open model development - model weights are more useful when the infrastructure and training decisions behind them are open too. For a deeper look at the systems design, experiments, ablations, and approaches Hugging Face tested along the way, read its technical report and explore Olmo-core 3 on GitHub.
OpenAI demonstrated GPT-6 integrated with Hugging Face's MicroDuck robot at its DevDay event. The two-legged robot features a camera and voice capabilities powered by GPT-6. The demonstration showcased how OpenAI's software can connect with physical devices and robots. MicroDuck used GPT Voice for speech and ran Code X in the background, allowing it to see and interact with its environment. The demo illustrated OpenAI's vision for AI integration beyond chatbots. The technology could potentially connect to various smart home devices, from thermostats to lighting systems, creating improved user interfaces and experiences. OpenAI emphasised its aim to develop software that works with real-world hardware, enabling personalised AI agents that could follow users and provide cohesive experiences across different platforms and devices.
Transformers now runs llama.cpp quants. Hugging Face adds GGUF loading to Transformers on Apple Silicon, with Qwen3.5 support on the main branch and performance tied to compatible kernels. Up next. Published on 30 September 2026 AI This post was created with the assistance of artificial intelligence (AI). Prime Big Deal Days · Oct 6-7 Offer from Amazon Get the latest gadgets delivered free - and shop member deals * Fast, free delivery on millions of items * Access to Prime Big Deal Days deals on October 6-7 * Prime Video, Amazon Music and more included As an affiliate, Barrier Magz earn on qualifying purchases. Hugging Face has added support for loading GGUF quantized checkpoints through Transformers' `from_pretrained` API, initially for Qwen3.5 models on Apple Silicon. The feature is on the library's main branch; wider hardware and architecture support, stable release timing and full benchmark details have not been specified. Hugging Face has added a way to load GGUF quantized checkpoints directly through its Transformers library, as detailed in the original analysis, bringing llama.cpp's model format into the familiar `from_pretrained` API. The initial rollout targets Qwen3.5 on Apple Silicon and is available on the library's main branch ahead of a stable release. Users can select a GGUF checkpoint hosted on the Hugging Face Hub and pass its file through the `gguf_file` argument to `from_pretrained`, then generate text using Transformers. Hugging Face says the integration reuses llama.cpp's ggml kernels; the announcement also compares performance with llama.cpp across three checkpoints: a small dense model, a larger dense model and a mixture-of-experts model. For supported Apple Silicon setups, Transformers can load compatible ggml/Metal layer kernels when weights remain packed on Metal. The feature uses `ggml-org/ggml-attn` for attention when available. If that kernel cannot be fetched, it falls back to standard SDPA attention with a warning; users can also select SDPA explicitly. Without a compatible quantization kernel, the loader dequantizes the model, which uses more memory. The stated requirements include an Apple Silicon Mac, a PyTorch version supported by the published kernel builds - generally one of the two latest releases - and current Transformers plus a compatible kernels library. The same checkpoints can also be served using `transformers serve`, which provides a local OpenAI-compatible API for clients configured to connect to that endpoint. At a glance announcement When: Available on the Transformers main bran... The development Hugging Face added a main-branch feature that lets Transformers load GGUF checkpoints using llama.cpp's ggml kernels on Apple Silicon. At a glance announcement When: announced April 2026; available via tra... The development Hugging Face announced that the transformers library can now run llama.cpp-style GGUF quantized models natively, using ggml kernels for near-llama.cpp performance on Apple Silicon. GGUF joins the Transformers workflow. The change gives developers already using Transformers and PyTorch a direct route to GGUF checkpoints, a format commonly used by local inference tools including Ollama, LM Studio and Jan. Previously, users generally relied on llama.cpp-derived software to run these files. The new path may let developers use a broader range of Hub checkpoints without changing their model-loading workflow, subject to the current platform and architecture limits. Quantization can reduce the memory needed to run a model. Hugging Face lists Unsloth's Qwen3.5-4B at 8.42 GB in BF16 and 2.74 GB in Q4_K_M. That smaller footprint can make local inference feasible on machines with less memory, though reduced precision can affect output quality. The announcement advises users to evaluate quantization on their own models and tasks; it does not establish that one setting will work equally well across workloads. Performance is another part of the rationale: the integration is designed to use ggml kernels rather than simply unpacking weights into a conventional format. Hugging Face says it benchmarked against llama.cpp, but speed depends on the hardware, model and configuration. The announcement's benchmark references alone do not establish a universal performance result. A new route for quantized checkpoints. GGUF, developed for the llama.cpp ecosystem, packages model weights and metadata in a single file. Its quantization variants trade some numerical precision for a smaller memory footprint. A label such as Q4_K_M indicates a mixed-precision format, with most weights stored at four bits and some tensors kept at higher precision. Hugging Face's example sizes for Qwen3.5-4B are 3.53 GB for Q6_K, 3.14 GB for Q5_K_M and 2.74 GB for Q4_K_M, compared with 8.42 GB for BF16. The company suggests starting with Q4_K_M and trying higher-precision variants when more memory is available, while warning that quality effects depend on the model and task. The feature arrives amid growing interest in running models locally. Hugging Face co-founder Julien Chaumond recently described a demonstration of Qwen3.6 27B running in the Pi coding agent through llama.cpp on a MacBook Pro. That was his assessment of a particular demonstration, not a benchmark establishing comparable performance across tasks or systems. ""We're adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop's memory through the familiar transformers APIs."" - Hugging Face announcement Platform and release scope remain open. The announced support is focused on Apple Silicon and Qwen3.5. Hugging Face has not specified a timeline for CUDA, Linux or Windows support, or said which model architectures will be added next. The feature is on the Transformers main branch, and no date has been announced for a stable release. The announcement refers to comparisons with llama.cpp across three checkpoints, but the results depend on the selected models and hardware. The information provided here does not establish a performance advantage across devices. It also remains unclear how quickly additional kernel builds and architecture support will be made available. Hugging Face cautions that quality changes from quantization depend on the model and task. The listed file sizes show memory differences, but they do not by themselves measure output quality or suitability for a particular workload. Stable release and broader support. The next practical milestone is a stable Transformers release that includes GGUF loading; Hugging Face has not announced when that will happen. Until then, users can access the capability from the main branch, provided their hardware, PyTorch version and kernels library meet the stated requirements. Users following the rollout can check Hugging Face's GGUF documentation and kernels library for changes to supported formats and compatibility. The announcement leaves expansion to other hardware backends and model architectures as open areas; it does not provide dates or a confirmed rollout sequence for them. Key questions. How do users load a GGUF checkpoint in Transformers? On a compatible setup, users can select a Hub checkpoint and pass its file with the `gguf_file` argument to `from_pretrained`. The feature is currently on the Transformers main branch. Which hardware and model family are supported initially? The initial rollout targets Apple Silicon Macs and Qwen3.5. Hugging Face has not announced a timeline for CUDA, Linux or Windows support. What happens if the quantization kernel is unavailable? The loader falls back to dequantizing the model, which uses more memory. If the ggml attention kernel cannot be fetched, the system falls back to SDPA attention with a warning; users can select SDPA directly. Is the feature in a stable Transformers release? No. It is available on the main branch. Hugging Face has not announced a stable release date. Does a smaller GGUF file guarantee the same output quality? No. Quantization reduces file size, but Hugging Face says its effect on quality depends on the model and task. Users should evaluate the model on their own workload. Fall picks. As an affiliate, Barrier Magz earn on qualifying purchases.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
1,001-5,000
Company Stage
Acquired
Total Funding
$395.7M
Headquarters
New York City, New York
Founded
2016
Find jobs on Simplify and start your career today