Featherless AI

Featherless AI

Serverless infrastructure for open-source AI inference

Overview

Featherless AI builds neutral, serverless infrastructure for open-source AI. Its platform handles inference, model and workflow optimization so teams can run tens of thousands of open models without being tied to hyperscalers. The product works by providing serverless inference and optimized runtimes for open-source models (including RWKV) through its Recursal lineage, aiming for a scalable, vendor-free AI stack. It differentiates itself by offering a neutral, open framework that minimizes lock-in, supports broad hardware (e.g., AMD ROCm), and fosters an open-source agent/runtime ecosystem, rather than relying on proprietary cloud constraints. The company’s goal is to widen access to open-source AI at scale, reduce inference costs, and expand its global footprint and ecosystem through partnerships and continued development of its open-model infrastructure.

About Featherless AI

Simplify's Rating
Why Featherless AI is rated
B-
Rated B on Competitive Edge
Rated B on Growth Potential
Rated C on Differentiation

Industries

Data & Analytics

Enterprise Software

AI & Machine Learning

Company Size

11-50

Company Stage

Series A

Total Funding

$25M

Headquarters

San Francisco, California

Founded

2023

Get referred to Featherless AI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • September 2026 launch of managed open-source agents expands beyond inference into workflows.
  • BMW i Ventures, Airbus Ventures, and AMD Ventures validate enterprise and hardware demand.
  • GLM 5.2, DeepSeek V3.2, and MiniMax M2.5 widen demand for open-model hosting.

What critics are saying

  • Hugging Face controls Featherless distribution; endpoint-routing changes in 2026 slash acquisition.
  • Flat-rate pricing gets crushed by OpenAI, Together, and cloud GPU price cuts by 2027.
  • If AMD ROCm slips behind NVIDIA CUDA, Featherless loses its core cost advantage.

What makes Featherless AI unique

  • Featherless serves 30,000+ open models through flat-rate serverless inference, not token billing.
  • Its June 2025 Hugging Face integration gives direct distribution inside model pages.
  • April 2026 Series A deepens native AMD ROCm support and hardware independence.

Help us improve and share your feedback! Did you find this helpful?

Funding

Total Funding

$25M

Above

Industry Average

Funded Over

2 Rounds

Series A funding typically happens when a startup has a product and some customers, and now needs funding to scale. This money is usually used to grow the team, expand marketing, and improve the product. Venture capital firms are frequently the main investors here.
Series A Funding Comparison
Above Average

Industry standards

$15M
$8.2M
Discord
$15M
Canva
$20M
Featherless AI
$30M
Kalshi

Benefits

Remote Work Options

Company Equity

Company News

Featherless AI
Sep 8th, 2026
How the Ethereum Foundation dAI team secures open-weight models using Featherless and Averta.

How the Ethereum Foundation dAI team secures open-weight models using Featherless and Averta. Summary. Featherless and Averta work together to make open model inference enterprise-ready. Featherless provides serverless access to a large and expanding library of open-weight models, while Averta adds the runtime security and governance layer needed to protect AI systems in production. Through the live webhook integration, Featherless can apply Averta Guardrails and policy controls across the model access it provides. The Ethereum Foundation dAI team's use of Featherless and Averta together for research purposes provides clear validation of the partnership. The challenge: open models need enterprise-grade security. For teams building AI products, model access has become a practical infrastructure decision. Closed model providers offer simple APIs, but limit teams to a smaller set of models. GPU hosting platforms offer greater flexibility, but often require teams to manage infrastructure and operational costs themselves. Featherless addresses this gap as a serverless provider with specialized model loading and GPU orchestration capabilities that allow it to keep an exceptionally large catalogue of models online. It combines a broad model range and variety with serverless pricing, giving teams low-cost, easy access to a continually expanding library of open-weight models without requiring them to operate their own inference infrastructure. As Featherless lowers the cost and operational burden of accessing open-weight models, the next challenge is securing the AI systems built on top at runtime. Native guardrails and system prompts are insufficient on their own because they were not designed to govern the full AI execution path. In production, AI systems may handle sensitive data, call tools, and trigger high-risk actions. Each interaction needs enforceable policy, monitoring, governance, and auditability. This is the problem Averta solves: protecting AI systems across the execution path, where classification, policy enforcement, access control, and auditability work together as a single security layer. Without that runtime layer, AI applications remain exposed to predictable failure modes: prompt injection, unauthorised tool use, sensitive data leakage, policy bypasses, and other real-world attacks. These risks become more serious when agents are embedded in workflows, connected to tools, or allowed to trigger actions on behalf of users. Governance expectations are also rising, with standards such as ISO/IEC 42001 and regulations such as the EU AI Act reinforcing the need for AI systems to be risk-managed, observable, auditable, and resilient in production. Enterprise readiness requires both sides of the equation: scalable access to models and a security layer that protects the full AI execution path, from input analysis and policy enforcement to runtime security and auditability. The collaboration: Featherless and Averta. The Featherless and Averta integration connects serverless open model inference with Averta's runtime guardrails and policy controls. Through the live webhook integration, Featherless now routes AI activity through Averta's runtime checkpoints, applying policy decisions around prompts, tool calls, tool results, and outputs across the AI execution path. Featherless manages the policies applied through this integration, maintaining a consistent security layer across the model access it provides. This allows teams to use Featherless' open-weight model access while applying Averta Guardrails to the production workflows, agents, and applications connected to those models. The partnership between Averta and Featherless exists to secure the open-weight models Featherless serves. Security is moving up the priority list for organizations integrating AI, and it is easier to build in at the inference layer than to bolt on afterward. Supporting Ethereum Foundation dAI team research. The Ethereum Foundation dAI team is using Featherless and Averta together for research purposes, providing a clear validation point for the partnership. Featherless delivers serverless access to open-weight models, while Averta Guardrails applies runtime security and policy enforcement to the AI workflows using them. For the EF dAI team, this means security decisions can be made across the AI execution path: before requests reach the model, before tools are exposed or called, when tool results return, and before final outputs reach users. Averta Guardrails helps enforce policy, monitor behaviour, and provide auditability at each of these points. The deployment shows how Featherless and Averta work together in a high-trust environment: Featherless provides the model access layer, and Averta provides the security and governance layer needed to evaluate agentic AI workflows safely and reliably. Outcomes and impact. The webhook integration between Featherless and Averta is live, enabling AI activity from Featherless to be routed into Averta Guardrails and policy controls. The Ethereum Foundation dAI team is using Featherless and Averta together for research activities, demonstrating how open model inference can be paired with runtime AI security in a high-trust environment. Averta Guardrails is built for the AI execution path, where classification has to be fast, precise, and consistent. The classification engine reaches 98.8% precision on adversarial and benign traffic, evaluated against held-out attack corpora with less than 40ms p99 latency. It provides 100% action coverage across prompts, tool calls, and outputs, with zero unclassified executions silently allowed. The security impact is visible in attack success rate testing across frontier models. Across the models tested, Averta Guardrails reduced average attack success rates from 73% to 4% for prompt attacks, 58% to 6% for unauthorised actions, 41% to 2% for data exfiltration, and 67% to 5% for policy violations. Averta's broader platform extends this security and governance layer across the AI stack. MCP Gateway helps organisations manage enterprise agents, agent identity and ownership, agent activity, and agent tool access across popular integrations such as Linear, GitHub, and Notion, across the entire workforce. Averta RED adds expert-led and automated AI red teaming for launch reviews, regression checks, and ongoing validation, simulating the prompt injection, tool abuse, and data exfiltration paths agents are likely to face in production. Together, these results show the role Averta Guardrails plays alongside Featherless' serverless inference layer: adding runtime protection, policy enforcement, and auditability to open model inference in production. For organisations with broader agent security needs, Averta also offers MCP Gateway for agent and tool governance and Averta RED for continuous testing and validation.

UK Tech News
Jun 30th, 2026
Featherless.ai partners with Z.ai to provide global access to new open-source AI model, GLM 5.2.

Featherless.ai partners with Z.ai to provide global access to new open-source AI model, GLM 5.2. Featherless.ai has partnered with Z.ai to deliver worldwide access to the recently launched GLM 5.2 AI model. With this collaboration, organisations can seamlessly deploy and run the model through Featherless.ai, without encountering any hindrances related to the infrastructure necessary for hosting large-scale open-source architectures. Featherless.ai is a Day Zero launch partner for GLM-5.2. The model is available through the OpenAI-compatible API, in FP8 with up to a 256K context window on public cloud and up to 1M on private cloud deployments. Featherless.ai hosts the model directly, which means no need for GPU provisioning, with no logs and with hosting in the EU and US. The launch takes place at a time when regulation and competition in AI are fierce. The U.S. government has just imposed export controls on the highest-level frontier model of Anthropic, the Fable 5. As an open-source alternative to closed-source technology, GLM 5.2 reduces the gap in competition. It is projected that GLM 5.2's initial evaluation will place it at a comparable level with that of Claude Opus 4.7-4.8. Eugene Cheah, CEO and co-founder of Featherless.ai, said: "Any attempt to enforce tight export controls or closed architectures as a way to keep the frontier features in check will only result in the development of an alternative. In the situation where an engineering team or company feels that the access might be revoked suddenly, there is an incentive to switch to technologies that they can operate and change entirely independently of anyone else. That's precisely why open source technology evolves so quickly." The structural abilities of GLM 5.2 are significantly improved and include a Mixture-of-Experts design with 744 billion parameters that use 39 billion parameters per token. The GLM-5.2 structure is similar in physical size to its predecessor, GLM-5.1. Performance is improved on a generation-by-generation basis thanks to careful coding agent training and an optimised 1 million token context window. To ensure quality during long coding agent sessions, the architecture includes IndexShare, which allows using a common lightweight indexer in each fourth sparse attention layer. Integration helps save compute costs by 2.9 times at high context lengths. Also, an updated multi-token prediction layer improves speculative decoding acceptance by about 20%, whereas built-in thinking effort control lets choosing between High and Max levels for reasoning vs. real-time trade-off. The model is released under the permissive MIT license and available through an OpenAI-compatible API. The architecture improvements yield tangible performance gains under several rigorous evaluation criteria. In long-term programming tasks such as FrontierSWE, PostTrainBench, and SWE-Marathon, GLM-5.2 becomes the top-ranked open-source model in the industry. It is the industry's first true open-source, drop-in replacement for enterprise software development capable of competing directly with closed systems like Claude Opus 4.8 and GPT-5.5 - something major engineering leaders and tech executives (Vercel, Tesla) have also taken notice of. Featherless.ai is introducing an exclusive private cloud deployment that pairs GLM-5.2 with AMD infrastructure to create a secure, fully compliant and highly cost-efficient software engineering environment. As the only platform to have optimised GLM-5.2 to run natively on AMD hardware, Featherless.ai enables organisations to completely bypass ongoing NVIDIA chip shortages and inflated procurement costs while maximising AMD's substantial compute-per-dollar advantage. As empirical testing reveals, Terminal-Bench 2.1 performance scores went from 63.5 to 81.0, and SWE-bench Pro scores rose from 58.4 to 62.1. The scores in complex development assessments were even more dramatic, with FrontierSWE rising from 30.5 to 74.4 and SWE-Marathon going from 1.0 to 13.0. In pure logical reasoning tasks, the scores went up as well, with AIME 2026 going from 95.3 to 99.2 and GPQA-Diamond scores rising from 86.2 to 91.2.

TipRanks
Jun 13th, 2026
Featherlessai raises $20M from AMD Ventures and Airbus Ventures to host millions of AI models

Featherlessai, an AI infrastructure company specialising in hosting open and fine-tuned language models, has raised $20 million in a funding round co-led by AMD Ventures and Airbus Ventures. The company announced the round at AMD Dev Day in San Francisco. It plans to use the capital to host and run millions of open models and large language models, whilst deploying an ecosystem of AI agents and applications including OpenClaw and Hermes Agent. Management said the funding will accelerate hiring, infrastructure development and product delivery. AMD Ventures' participation suggests potential synergies in high-performance computing and specialised AI hardware, whilst Airbus Ventures may provide access to aerospace and industrial enterprise use cases where AI infrastructure is increasingly deployed.

Byteiota
May 4th, 2026
Featherless.ai raises $20M: open-source AI challenges proprietary giants.

Featherless.ai raises $20M: open-source AI challenges proprietary giants. 14 hours ago 0 Featherless.ai raised $20 million in Series A funding co-led by AMD Ventures and Airbus Ventures on May 1, 2026. The Singapore-founded startup runs a serverless platform supporting 30,000+ open-source AI models - production-grade infrastructure challenging OpenAI and Anthropic's proprietary APIs. The investor roster tells the real story: AMD wants an alternative to NVIDIA's AI dominance, Airbus is betting specialized models beat general-purpose LLMs, and BMW needs cost-effective automotive AI. Corporate titans are backing open-source because they see the writing on the wall. Strategic bets against the closed ecosystem. This isn't venture capital diversification. AMD Ventures, Airbus Ventures, and BMW i Ventures are making strategic counter-moves to OpenAI, Anthropic, and NVIDIA monopolies. AMD's partnership ensures popular open-source models run natively on its ROCm platform - a direct challenge to NVIDIA's CUDA lock-in. Airbus Ventures is betting that "millions of specialized, fine-tuned models" will replace the few general-purpose systems everyone's racing to scale. That's a bet against GPT-5 and Claude Opus 5 before they even launch. Featherless co-founder Eugene Cheah, creator of the RWKV model architecture, framed it plainly: "Open-source is the only real check on that, and it only works if the infrastructure to run it actually exists." The funding builds that infrastructure - global serverless deployment, AMD ROCm optimization, and a marketplace for specialized models launching soon. The economics aren't close - they're overwhelming. The gap between closed and open AI models isn't narrowing. It's inverting. Closed models cost six times more than open alternatives, according to MIT Sloan research. Reallocating demand from closed to open could cut global AI spending by 70% - $25 billion annually. One fintech company cut monthly AI costs from $47,000 to $8,000 (83% reduction) by switching to open models via platforms like Featherless. Performance? Llama 3, Mistral, Qwen, and DeepSeek now match GPT-4 and Claude on most benchmarks. DeepSeek trained its V3 model - competitive with GPT-4o - for $5.5 million. OpenAI reportedly spends hundreds of millions per training run. Open-source models deliver ~90% of closed model performance at release and catch up within months. Context windows hit 128K tokens across major families. Inference tooling matured: Ollama, llama.cpp, and platforms like Featherless run production workloads without DevOps overhead. AMD's cost advantage is real. AMD's ROCm platform delivers 25-40% cost savings versus NVIDIA equivalents, with a performance gap that's shrunk to 10-30% (and closing). AMD's MI355X chips now score within single-digit percentage points of NVIDIA's B200 on inference benchmarks. For memory-bandwidth-heavy workloads - large model prefill, long-context generation - ROCm is genuinely competitive. Featherless's native AMD ROCm support gives developers a credible NVIDIA alternative. If you're running inference at scale, that 25-40% cost difference compounds fast. AMD isn't a charity case anymore. It's a strategic hedge. Specialized models vs. General LLMs: the Airbus bet. Airbus Ventures isn't investing in "better ChatGPT." Its thesis: the next phase of AI adoption won't be GPT-5 getting bigger and more general. It'll be millions of domain-specific models - aerospace diagnostics, fraud detection, autonomous vehicle planning - trained on specialized datasets and fine-tuned for narrow tasks. Why specialized models win: better performance on specific tasks, lower inference costs (smaller models), faster responses, and no data leaves your infrastructure. Featherless is launching a marketplace where developers can discover and deploy these specialized models. The platform already supports 30,000+ models. That number will grow as enterprises realize they don't need GPT-5 for every task. What developers get today. Featherless offers 30,000+ open models (language, vision, audio, multimodal) via a serverless platform with flat-rate pricing. No per-token billing. No surprise invoices. Models load in under five seconds. The API is OpenAI-compatible, so migration from proprietary services takes hours, not weeks. Multi-region infrastructure (EU/US) handles data sovereignty requirements. If you're paying OpenAI or Anthropic per-token today, Featherless is a credible exit strategy: 90% cost savings with minimal quality loss. Developers report "removing the fear of the token meter" and gaining predictable monthly costs. The platform is already Hugging Face's fastest-growing inference provider. What comes next. The $20 million funds four priorities: expanding the model library, shipping an open-source agent runtime, deepening AMD hardware integration, and scaling enterprise deployments with private environments. The vision is "AI independence" - developers and companies controlling their AI stack instead of renting compute from Big Tech. Corporate VCs don't back open-source infrastructure on ideological grounds. They back it because closed ecosystems create strategic vulnerabilities. AMD needs competition to NVIDIA. Airbus needs specialized models for aerospace applications. BMW needs cost-effective, auditable AI for vehicles. Featherless is building the infrastructure that makes open-source AI production-ready at scale. 2026 isn't the year open-source AI "catches up" to proprietary models. It's the year it overtakes them on cost, flexibility, and strategic independence. The economics are overwhelming, the performance gap is closed, and the infrastructure now exists. Developers finally have leverage. I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

Tech in Asia
Apr 30th, 2026
Singapore's Featherless raises $20M series A for flat-rate AI inference platform

Featherless, a Singapore-based AI startup, has raised $20 million in a Series A round. The company plans to use the funds to expand infrastructure, add a marketplace for specialised models and improve support for different chip and computing setups. The startup supports over 30,000 open models across language, vision and audio, with infrastructure in the US and Europe. Featherless offers flat-rate subscriptions with unlimited monthly requests and a fixed cap on concurrent requests, breaking from traditional per-token billing that creates uneven costs. The company integrates with Hugging Face and claims to be the largest large language model inference provider there, hosting more than 6,700 open-weight models. Billing can run through either Featherless or Hugging Face.

Recently Posted Jobs

Sign up to get curated job recommendations

Featherless AI is Hiring for 6 Jobs on Simplify!

Find jobs on Simplify and start your career today

Don't see your dream role? Check out thousands of other roles on Simplify. Browse all jobs →