Fal

Fal

NLP-based sentiment & anomaly detection

Compute Operations

Full-TimeUpdated on 9/24/2026
$160k - $200k/yr

+ Equity

Senior
San Francisco, CA, USA
In Person

Relocation assistance to San Francisco is offered.

H1B Sponsorship Available

About the job

Requirements
  • At least 5 years of experience in vendor management, cloud partnerships, strategic sourcing, or infrastructure procurement within the artificial intelligence, graphics processing unit, cloud, or hyperscale infrastructure industry.
  • Direct experience working with neocloud, hyperscaler, or graphics processing unit infrastructure providers while owning commercial relationships and vendor performance.
  • Ability to manage contract records, reconcile invoices, calculate credits, and maintain accurate vendor documentation.
  • Ability to own vendor escalations, run quarterly business reviews and scorecards, and communicate with executive vendor contacts and internal stakeholders.
  • Experience managing multi-vendor allocation by priority during shortages.
  • Experience establishing a vendor-management function from scratch at a neocloud or hyperscaler.
Responsibilities
  • Own vendor performance across graphics processing unit hyperscalers and neocloud providers by defining and tracking service-level agreements, managing escalations, remedies, and service credits, and holding partners accountable to their commitments.
  • Build strategic relationships with key infrastructure partners through regular business reviews and executive engagement with Tier 1 vendors.
  • Align capacity to demand by partnering with Capacity Planning to secure committed graphics processing unit capacity, forecast growth, and mitigate supply constraints.
  • Own vendor governance, including contract compliance, invoice reconciliation, renewals, and performance reporting.
  • Run regular business reviews with vendor partners to discuss performance, open credit claims, capacity planning, and outstanding issues.

About the company

Fal.ai helps businesses improve data analytics using NLP and ML, focusing on sentiment analysis and anomaly detection within dbt data models. It analyzes text from customer reviews, support tickets, and surveys to label sentiment as positive, negative, or neutral, and flags unusual patterns in data transformations. The platform integrates with existing data infrastructure using dbt models and is offered via tiered subscriptions that include basic sentiment analysis, advanced anomaly detection, and premium support. Its goal is to help data-driven organizations make informed decisions, improve customer satisfaction, and get continuous analytics updates.

Company Size

51-200

Company Stage

Late Stage VC

Total Funding

$943.9M

Headquarters

Seattle, Washington

Founded

2021

Get referred to Fal

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • September 2026 H3 Max and Turbo are priced at $0.02-$0.04 per second through September 30.
  • fal says 2.5 million developers use its infrastructure, expanding distribution for new launches.
  • fal Agent and Sound Effects 1.0 broaden workflows from generation into orchestration and audio.

What critics are saying

  • The Maricopa County case names FAL in January 2026 over explicit AI content.
  • H3 Max depends on MiniMax's open weights, exposing fal to supplier leverage and cloning.
  • Model commoditization turns fal into a low-margin router if customers bypass its stack.

What makes Fal unique

  • fal's September 2026 Lucent acquisition deepens creative-workflow control beyond raw model hosting.
  • fal's inference stack serves 1,300+ production endpoints across 1,000+ models, per September 2026 PR.
  • H3 Max combines post-training and inference co-design, winning September 2026 video benchmarks.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Company Equity

Relocation Assistance

Growth & Insights and Company News

Headcount

6 month growth

↓ -3%

1 year growth

↑ 2%

2 year growth

↑ 14%
PR Newswire
Sep 14th, 2026
fal Acquires Lucent, Welcomes Co-Founders Dimi Panagiotopoulos and Alex Koumpas to the team

/PRNewswire/ -- Dimi Panagiotopoulos Lucent's Co-founder and CTO, is a blend of engineer and creative technologist. By the age of 21 he had built and shipped...

Online Casino Reports
Sep 13th, 2026
AI video Generation pushes us closer to the world of Blade Runner 2049.

AI video Generation pushes Online Casino Reports closer to the world of Blade Runner 2049. Shane. - September 13, 2026 Ready to live in the world of Blade Runner? AI can now create video faster than you can watch it. Online Casino Reports explore what could be when instant AI video supports generative online casinos, interactive entertainment and even virtual social experiences. Earlier this year, Shanghai-based AI company MiniMax built an exciting new AI model. MiniMax H3 was considered a breakthrough in generative AI in the video space and was described as follows: MiniMax H3 is an open-weights, general-purpose multimodal AI model that generates up to 15 seconds of high-definition 2K video with native stereo audio from text, image, video, and audio inputs. Then an American-based competitor, Fal, adapted the model and elevated it to new heights with the release of their own version called MiniMax H3 Max. The Max iteration pushed the capabilities of the original model significantly further, delivering: * Lightning Speed: Generates any 5-second video clip in less than 3 seconds, making it 35 to 50 times faster than the base model. * Native Audio: The model creates synchronised stereo sound and dialogue during the same pass. * Resolution Options: The model supports 480p and 768p outputs at 24 frames per second. * Generation Modes: Create videos using both Text-to-Video and Image-to-Video. This means it can generate new content faster than humans can watch it. In a world where seamless integration and the appearance of live engagement are key, this is a watershed moment. Incredible live stream test. A Fal engineer showcased the potential of MiniMax H3 Max by using it to run a live stream. He revived an old idea, channel flipping, by connecting the AI to different tabs, each running its own ongoing stream of randomised prompts, and then simply flipping between them. Each flip of the "channel" presented him with a completely new scene from a completely new show. This turns Bruce Springsteen's "57 Channels and Nothin' On" from a tongue-in-cheek song lyric into science-fiction reality. Then came the really interesting part: viewers in the chat could type in their own ideas and instructions, and the AI would begin incorporating them into the video. The audience could effectively determine the direction of the content, influencing the scenes, events and even the eventual outcome of something being created in real time. Imagine not liking how Game of Thrones Season ended and simply prompting your own conclusion. Have AI create the action-packed, storyline-wrapping finale you always felt the show deserved while you watch. Better yet, ever complained that there's just nothing worth watching anymore? This technology could eventually allow anyone to create their own channels for private viewing, or even distribute them to a wider audience. Suddenly, everyone could claim the title of arteur. Have Claude write you a story. Feed it into MiniMax H3 Max. While you're at it, have ChatGPT design you a business card that reads: Indie AI Filmmaker & Showrunner And congratulations. You're officially a one-man studio. Future online casino application. While most conversations around this technology are focused on television and movie production, for the online gambling industry it represents something potentially much more interesting: a step toward truly interactive, AI-hosted casinos. Online Casino Reports has already seen impressive work from companies such as RAVATAR in creating virtual hosts with distinct personalities, voices and other characteristics. Once those virtual characters become the blueprint for AI video that can respond to user engagement on the fly, Online Casino Reports could have an AI host that players can actually speak to and joke with, without ever needing to type into a chat window. With near-instant video responses to spoken interaction, the experience moves beyond what Online Casino Reports currently see with chatbots and starts to resemble K (played by Ryan Gosling) interacting with his virtual companion Joi (played by Ana de Armas) in Blade Runner 2049. This virtual "person" would know your preferences. Remembers your previous conversations. Responds instantly. Develops a personality around your interactions. And, over time, builds up a shared history with you. There is also the wishlist Online Casino Reports drew up years ago when Online Casino Reports discussed the future of AI and casino gaming. Give an AI access to the mathematical models powering different game types, and it could potentially begin creating custom gambling experiences on demand. Feel like fishing that day? Ask your AI host to create a casino game where you fish the Great Lakes, pulling different species from the water for different prizes. Want something that feels more like a video slot? Use a slot-game engine. Prefer fixed odds? Use a fixed-odds engine. This evolves an online casino from a collection of premade games of chance into something that can generate an entirely new experience every time you sign into your player account. The future is coming, but what a price tag. As exciting as all of this is, there is one major hurdle the technology has to clear before it can become an everyday reality: the exorbitant cost. At the moment, running the model can cost more than $1,000 per hour, meaning it simply isn't ready to be deployed as a mainstream commercial product. But, as Online Casino Reports has seen with everything from "AI can't do hands" to "AI people could never fool us," new technology can move remarkably quickly from an impressive idea to an accessible, everyday tool. While the future of fully interactive AI-powered gaming may not be here quite yet, there is already plenty for you to explore today. Here are some of its top choices and where you can experience them for yourself: * Play iDealer Boulevard Blackjack by ICONIC21 at Cooked Casino to enjoy a cutting-edge AI-hosted gaming experience which showcases RAVATAR's market-leading AI-powered dealers. * Jump onto Stake Casino today and check out their SlotGPT feature. Create your own online slot game instantly by entering a prompt - which can be as simple or complex as you like. * Classic Baccarat is being revamped by leading providers Microgaming and Pragmatic Play, with the introduction of new social features, winning hands and speedier game rounds. Open an account at 1xBet Casino to play All-In Speed Baccarat and Big Small Baccarat. * Chicken games are one of the most popular crash game spinoffs anywhere online. ElCasino took things to the next level by creating a live casino version of the popular tap game genre called Chicken Live (free demo). If it's your first time joining one of these top online casinos or crypto gambling sites, be sure to check out the available welcome bonuses and promotional offers. As a new customer, you could be eligible for several generous offers. Uncover the premier online casinos to enjoy the games highlighted in this article Mentioned in this article. 1xBet Casino Payout Speed 2 - 3 Days Payment Methods 18+ Only, GambleAware.org, T&C Apply.

HyperAI
Sep 5th, 2026
MIT phd Li Muyang launches Nunchux AI to accelerate multimodal inference.

MIT phd Li Muyang launches Nunchux AI to accelerate multimodal inference. Nunchux AI has officially launched as a new player in the generative AI infrastructure space, founded by MIT doctoral graduate Li Muyang, Carnegie Mellon University associate professor Zhu Junyan, and researchers Yujun Lin and Zhekai Zhang. The company targets the growing industry demand for efficient, low-cost, and reliable multimodal inference as deployment scales rapidly. Backed by early-stage investments from Emergence Capital and E14 Fund, MIT Media Lab, Nunchux AI builds upon a decade of research into neural network compression and system optimization. Li's academic work, particularly the SVDQuant framework for 4-bit weight and activation quantization, forms the technical foundation of the venture. This research materialized in Nunchaku, an open-source inference engine that drastically reduces memory footprint and latency, enabling high-resolution image generation on consumer hardware. The engine has secured significant developer adoption, recording nearly 4,000 GitHub stars and native compatibility with major open-source models and development frameworks. Commercially, the startup will debut Modelverse, a unified API platform integrating over thirty image, video, and virtual avatar models. The service offers two operational tiers: Radical Speed for ultra-low latency and Radical Value for cost optimization. Unlike existing inference aggregators that primarily manage GPU orchestration, Nunchux AI pursues vertical integration by combining custom quantization algorithms with optimized inference kernels. This approach delivers superior performance and lower overhead on identical hardware. The Nunchaku core will remain open-source, while commercial revenue will be driven by premium API access, enterprise-grade customization, and dedicated developer support. The venture extends the entrepreneurial legacy of MIT professor Han Song's HAN Lab, which previously spawned companies such as DeePhi Tech, OmniML, Eigen AI, and Inco AI. Academic production continues in parallel, with recent publications including the FourTune paper explicitly listing the company as an affiliation. In a crowded inference market featuring competitors like fal.ai and Replicate, Nunchux AI differentiates itself through algorithmic efficiency rather than mere aggregation. The founding team maintains that reducing latency and deployment costs is critical for preserving creative velocity and democratizing access to advanced generative models, while strictly preserving output fidelity. This news is intelligently aggregated by AI to deliver industry updates efficiently. It does not constitute opinions or advice. MIT Technology Review

PR Newswire
Sep 1st, 2026
fal launches H3 Max video model with faster-than-real-time generation, ranks #1 on Design Arena and Artificial Analysis benchmarks

Fal, a generative media platform, has launched H3 Max, a new video generation model that ranks first on independent benchmarks from Artificial Analysis and Design Arena. The model generates five-second videos in approximately three seconds, achieving faster-than-real-time generation. Built on the open-weights MiniMax H3 model, H3 Max was developed through post-training for improved prompt adherence and visual quality. Fal reports the system delivers roughly 35 times the throughput of the official MiniMax H3 endpoint and averages 15 times faster than comparable quality models. H3 Max is available through the fal API and Playground. During a promotional launch period ending 7 September, pricing starts at $0.04 per second at 768p resolution, rising to $0.08 per second thereafter. Fal developed the model by co-designing training and inference optimisation, treating them as a single problem rather than separate layers.

Lore
Aug 28th, 2026
Lore issue #200: OpenAI's first AI chip beats Nvidia's GB300 on efficiency.

Lore issue #200: OpenAI's first AI chip beats Nvidia's GB300 on efficiency. PLUS: Nvidia nears a $13B Hugging Face deal, Sam Altman says AGI could arrive this year, and Hugging Face unveils a $399 open-source robot Aug 28, 2026 Good morning, welcome to this week's Lore Brief, your 3-minute brief of the most important moves in AI and tech. This issue is brought to you by Factory, the fastest way to ship software with autonomous engineering agents. * OpenAI's Jalapeño chip beats Nvidia GB300 on efficiency | First published results show the custom inference chip delivering 1.5 to 1.9 times more work per watt than Nvidia's GB300 while cutting end-to-end latency by as much as 3.6 times. The 700-watt part goes up against a 1,400-watt flagship and still returns answers faster on open models like DeepSeek R1 and Kimi K2.5. Deployment inside OpenAI's own infrastructure is planned by year-end. Read more here | * fal's H3 Max claims the top spot in video generation | The post-trained model ranked first for quality prompt understanding and aesthetics against a dozen leading systems. It can produce a 5-second 720p clip in under three seconds and is 50 percent off this week. fal built it on MiniMax's open H3 base then tuned it for speed without giving up looks. Read more here | * Nvidia closes in on a Hugging Face acquisition | Talks point to a deal around $13 billion that would give the chipmaker a major foothold in open-source AI. Hugging Face last raised at $4.5 billion in 2023 and now does roughly $150 million in annual revenue. The move would help Nvidia stay central as closed labs build their own chips. Read more here Lip Sync tutorial for AI videos: The same style ad like the one of the new Mac mini but made with AI: * OpenAI leaders say AGI is getting close | Sam Altman told TIME the company could have an internal system that qualifies as AGI by the end of 2026. Mark Chen put the lab at about 80 percent of the way there. Astra already acts as an automated research intern that can run week-long experiments inside OpenAI's own codebase. Read more here | * A practical field guide to living with Grok Bot | Two weeks after launch Matt Van Horn published every hack he has found from giving a bot its own inbox to letting it place phone calls in Portuguese. The write-up covers planning layers named roles cookie sync and the habit of drafting before anything gets sent. The honest caveat is that unsupervised crews can multiply their own errors fast. Read more here | * TIME drops its 2026 list of AI's most influential people | The annual TIME100 AI ranking is out again with the usual mix of lab chiefs and public figures. Online reaction quickly turned to the oddities including Paris Hilton while Jensen Huang is missing from the list. Readers treated that combination as the punchline. Read more here | * Claude memory now follows you from chat into Cowork | One shared memory means a task in Cowork can start from what you already discussed in chat including projects preferences and past clients. You can read edit or delete every saved topic in Settings. Sensitive subjects stay out unless you turn them on. Read more here | * Hugging Face unveils a $399 open-source robot | Microduck can walk pick things up get back up after falling and even roller-skate. Users teach it new tricks with reinforcement learning instead of waiting for a closed lab to ship a firmware update. Clem Delangue frames it as affordable hardware for physical AI and world models. Read more here | * Skild AI's S1 learns 10-minute robot tasks from one video | Show the foundation model a single demonstration and it can complete jobs it never saw in training from making coffee to flipping pancakes. No fine-tuning is required and it can recover from mistakes the human in the video never made. The company says matching that accuracy with older VLAs would take 50 to 100 hours of extra data. Read more here | * Google ships Gemini 3.5 Transcribe for cleaner live speech | The new model turns messy audio into formatted text while cutting filler words and handling self-corrections like "Tuesday no Wednesday." It recognizes more than 85 languages and drops word error rates well below earlier Chirp systems. Developers can use it now through the Gemini API and it is already landing in the macOS Gemini app. Read more here That's it for this week's Lore Brief. See you next week!