Full-Time

Member of Technical Staff

Performance Optimization

Fireworks AI

Fireworks AI

201-500 employees

AI inference platform for model deployment

Compensation Overview

$175k - $220k/yr

+ Equity

Company Does Not Provide H1B Sponsorship

San Mateo, CA, USA

In Person

Category
Software Engineering (1)
Required Skills
LLM
CUDA
Pytorch

Get referred to Fireworks AI

Find people who can refer or advise you

Requirements
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience
  • 5+ years of experience working on performance optimization or high-performance computing systems
  • Proficiency in CUDA or ROCm and experience with GPU profiling tools (e.g., Nsight, nvprof, CUPTI)
  • Familiarity with PyTorch and performance-critical model execution
  • Experience with distributed system debugging and optimization in multi-GPU environments
  • Deep understanding of GPU architecture, parallel programming models, and compute kernels
Responsibilities
  • Optimize system and GPU performance for high-throughput AI workloads across training and inference
  • Analyze and improve latency, throughput, memory usage, and compute efficiency
  • Profile system performance to detect and resolve GPU- and kernel-level bottlenecks
  • Implement low-level optimizations using CUDA, Triton, and other performance tooling
  • Drive improvements in execution speed and resource utilization for large-scale model workloads (LLMs, VLMs, and video models)
  • Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency
  • Improve support for mixed precision, quantization, and model graph optimization
  • Build and maintain performance benchmarking and monitoring infrastructure
  • Scale inference and training systems across multi-GPU, multi-node environments
  • Evaluate and integrate optimizations for emerging hardware accelerators and specialized runtimes
Desired Qualifications
  • Master’s or PhD in Computer Science, Electrical Engineering, or a related field
  • Experience optimizing large models for training and inference (LLMs, VLMs, or video models)
  • Knowledge of compiler stacks or ML compilers (e.g., torch.compile, Triton, XLA)
  • Contributions to open-source ML or HPC infrastructure
  • Familiarity with cloud-scale AI infrastructure and orchestration tools (e.g., Kubernetes)
  • Background in ML systems engineering or hardware-aware model design

Fireworks AI provides an AI inference platform that helps organizations run, customize, and deploy machine learning models. It supports deployment, fine-tuning, and inference workflows, with access to open-source models and on-demand deployment options through a subscription model. The platform enables users to configure and optimize models for production environments, manage variants, and scale AI workloads while controlling costs. Unlike some competitors, the emphasis is on end-to-end production-grade inference and fine-tuning across a range of clients, from tech firms to research institutions and enterprises, with tools tailored for rapid deployment and cost efficiency. The long-term goal is to expand the use of AI in production by building compound AI systems, growing the team, and delivering broader platform capabilities that accelerate AI adoption in real-world settings.

Company Size

201-500

Company Stage

Series D

Total Funding

$1.8B

Headquarters

Redwood City, California

Founded

2022

Get referred to Fireworks AI

Find people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Fireworks crossed $1B ARR in July 2026, growing 5x year-over-year while processing 40 trillion tokens daily.
  • The March 2026 Hathora acquisition enables sub-second global latency and smarter request routing for AI inference.
  • Cost advantage of 5–10x cheaper open-model inference drives adoption among firms like Coinbase.

What critics are saying

  • Cursor's $60B SpaceX acquisition eliminates Fireworks' largest revenue source representing 50% of FY2025 revenue.
  • Amazon and Google erode growth by expanding open-model hosting with lower total cost of ownership.
  • CoreWeave and Lambda force margin cuts through cheaper GPU rental pricing in high-volume token segments.

What makes Fireworks AI unique

  • Fireworks AI delivers 12x faster inference than vLLM using its proprietary FireAttention stack.
  • The platform supports four parallelization techniques optimized for different model types simultaneously.
  • It offers serverless inference and dedicated Deployments with autoscaling and quantization compression.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Professional Development Budget

Growth & Insights and Company News

Headcount

6 month growth

-5%

1 year growth

-3%

2 year growth

1%
Netzender
Jul 16th, 2026
Nvidia-backed Fireworks hits $17.5 billion valuation as companies pursue cheaper AI models.

Nvidia-backed Fireworks hits $17.5 billion valuation as companies pursue cheaper AI models. Jul 16, 2026 - 19:17 Fireworks founders pose for a photo at the startup's headquarters in San Mateo, California. From left in front row: Chenyu Zhao, CEO Lin Qiao and Benny Chen. Back row: James Reed, Pawel Garbacki, Dmytro Ivchenko and Dmytro (Dima) Dzhulgakov. The cost of the latest artificial intelligence models is increasingly breeding anxiety among finance executives, who have started directing employees to consider open-source alternatives. That's boosting cloud startup Fireworks, which competes with Amazon and Google to host models that developers can weave into applications. The Nvidia-backed company said Thursday that it has exceeded $1 billion in annualized revenue, five times what it had last year, and it has now raised a $1.5 billion round at a $17.5 billion valuation. "We're seeing super-linear demand," Lin Qiao, Fireworks' co-founder and CEO, told CNBC in an interview at the company's headquarters in San Mateo, California. "This is a once-in-a-lifetime opportunity to have this kind of market." Fireworks is much smaller than Anthropic and OpenAI, which investors have valued above $800 billion each this year, nor is it close to the top names in technology, whose market capitalizations are counted in the trillions. But the startup's revenue milestone suggests that companies aren't completely satisfied with the models coming out of the top labs. The achievement also presents new evidence that Amazon, Microsoft and Google are not totally dominating in cloud computing. Shares of easy-to-use cloud infrastructure vendor DigitalOcean are up 149% so far this year as growth has accelerated. CoreWeave, which rents out Nvidia graphics processing units, or GPUs, raised $1.5 billion in an initial public offering last year and is now worth $42 billion. By managing computing infrastructure for models, Fireworks does business in the inference cloud market, alongside startups such as Baseten and Together AI. It's also started providing GPUs for training AI models, like neoclouds CoreWeave, Lambda and Nebius. Rather than go it alone, Fireworks has started forming alliances. In March, it announced a partnership with Microsoft, which fields its own Foundry service for running open models. The arrangement allows customers of the Windows and Office company to draw on models through Fireworks, which relies on computing power from more than 20 suppliers, including Microsoft. "Through Microsoft we can get much bigger reach," Qiao said. Fireworks gives developers an easy way to adopt models from Chinese companies such as DeepSeek, MiniMax and Z.ai. Open-weight models OpenAI released last year are also available. The idea is for clients to bring their own data that frontier labs don't have and refine models until they deliver state-of-the-art performance for specific tasks, Qiao said. While Anthropic and OpenAI serve up "generalized intelligence," Fireworks can unlock "specialized intelligence," she said. The argument might sound familiar to those following the discourse of technology figureheads. Microsoft CEO Satya Nadella wrote in a Sunday blog post that "a company should be able to use a model without giving up the knowledge that makes it unique." Nadella was referring to Palantir CEO Alex Karp's remarks on CNBC earlier this month. Technical customers "want to know they own the means of production," Karp said. "It's not being transferred to someone else." Dollars and cents are a factor, too. Cryptocurrency exchange operator Coinbase has been adopting cheaper models where it makes sense, CEO Brian Armstrong wrote in a June X post. "Our cost compared with the equivalent-quality closed model is five to 10 times cheaper," Qiao said. A former Meta director, Qiao and six of her co-founders started Fireworks in 2022. The company employs around 200 people. Qiao expects the head count to reach 600 by the end of 2026. "This is the year when we'll really hit the gas," Qiao said. Fireworks hired former Salesforce executive George Hu as its president in April. The startup plans to assemble a formidable sales team after years of having customers sign themselves up. The new money will also help Fireworks obtain more GPUs and hire more technologists. Developers are increasingly counting on Fireworks to handle requests. Fireworks now handles 40 trillion AI tokens per day, Qiao said. Google disclosed in May that its AI models were processing about 19 billion tokens per minute for developers, implying more than 27 trillion per day. OpenAI announced in March that its developer tools were working through 15 billion tokens per minute, which would suggest about 22 trillion per day. Each token equates to about three-quarters of a single word. As of last year, about half of Fireworks' revenue came from AI coding startup Cursor, which has become less dependent on OpenAI and Anthropic and built a custom model named Composer. "We are much more diversified right now," Qiao said. In June, Elon Musk's SpaceX agreed to acquire Cursor in a $60 billion stock deal, with the transaction set to close this quarter. Other Fireworks clients include Elastic, GitLab and MongoDB. Atreides Management, Index Ventures and TCV led Fireworks' new round. Nvidia also participated, as did Evantic and Lightspeed Venture Partners.

Fireworks AI
Mar 10th, 2026
Fireworks AI

Fireworks Acquires Hathora to Accelerate Global Computer Orchestration

SiliconANGLE Media
Mar 9th, 2026
Fireworks AI acquires Hathora to build global real-time compute infrastructure for agentic AI

Fireworks AI has acquired Hathora, a real-time compute and server orchestration platform, to strengthen its global compute infrastructure for AI inference and training. Chief executive Lin Qiao described the deal as a talent-and-infrastructure acquisition rather than a customer acquisition. Hathora, launched in 2023, built a container orchestration platform across 14 regions serving multiplayer games and real-time AI workloads. Qiao drew parallels between gaming infrastructure's latency demands and AI inference requirements, noting gamers tolerate reduced graphics but not lag. The acquisition supports Fireworks' vision of "millions of models" continuously customised for specific use cases, rather than relying on single generalised models. Qiao emphasised that Fireworks focuses on automated customisation beyond just inference, positioning the company to handle the increasing velocity of agentic AI interactions.

Hathora
Mar 4th, 2026
Hathora Is Joining Fireworks AI

Hathora is joining Fireworks AI. 04 Mar 2026 Today, Hathora Inc.'d like to share that Hathora has been acquired by Fireworks AI. Its team will be joining Fireworks to work on compute orchestration for AI inference at scale. Over the past four years, Hathora Inc. built a global container orchestration platform spanning 14 regions, two bare metal providers and four clouds. Hathora Inc. powered server infrastructure for live titles like Splitgate 2, Stormgate, and Predecessor, and more recently expanded into real-time AI workloads with its voice model marketplace. The throughline was always the same: low-latency compute orchestration across heterogeneous infrastructure, without compromising on performance. Fireworks AI is where that work can have the most impact. The team Hathora Inc. built at Hathora is obsessed with infrastructure, and at Fireworks, they can continue to do what they do best. Founded by the team behind PyTorch at Meta, Fireworks processes more than 10 trillion tokens a day for over 10,000 customers and has built one of the fastest-growing AI inference platforms in the world. The challenge of orchestrating GPU compute across providers at the latency, reliability, and performance their customers demand is exactly the problem Hathora Inc. has spent four years solving. For its gaming customers, support will continue through May 5, 2026, and Hathora Inc. has partnered with Nitrado's GameFabric to provide a clear migration path and hands-on support through the transition. Hathora Inc. has already been in direct contact with its active customers. Details on timing and migration support have been shared directly. Thank you to its team, who took a bet on two first-time founders. To Upfront Ventures, Founders Fund and Lunar Ventures for backing Hathora Inc. early. And to its customers, from the game studios who shaped its platform to the AI teams who pushed Hathora Inc. forward. Onwards, Harsh & Sid

The Wall Street Journal
Oct 28th, 2025
Fireworks AI Raises $254M, Valued $4B

Fireworks AI, a startup focused on providing developers with access to advanced AI chips and models, announced it has raised $254 million in a recent funding round. This investment values the company at $4 billion, according to the Wall Street Journal.