B

Baseten

Deploys and serves scalable AI models

Capacity Manager - Programs

Full-TimeUpdated on 10/5/2026
$265k - $285k/yr+ Equity
Senior, Expert
San Francisco, CA, USA
Hybrid

About the job

Requirements
  • At least 8 years in technical program management, delivery management, or infrastructure program leadership, including at least 3 years managing large-scale physical infrastructure or data center delivery programs.
  • Demonstrated experience delivering on-premises data center builds, including power, cooling, networking, and rack-and-stack, from planning through go-live.
  • Direct experience with GPU cloud or neo cloud capacity delivery, including understanding GPU cluster architectures such as InfiniBand/RoCE fabrics, rack power density, liquid-cooling considerations, and the commercial and operational mechanics of capacity contracts with neo cloud or colocation providers.
  • A track record managing complex, multi-vendor, multi-workstream programs with hard external deadlines and significant financial or business exposure.
  • Ability to communicate complex technical and operational risks clearly in decision-ready updates for senior leadership.
  • Proven vendor and partner management experience, including negotiating delivery commitments and holding partners accountable.
  • Ability to operate with ambiguity in a fast-moving, high-growth environment and apply structured problem-solving.
Responsibilities
  • Own delivery of on-premises infrastructure builds, including colocation expansions, power and cooling readiness, rack-and-stack, network fabric bring-up, and hardware acceptance testing, coordinating colo providers and partners, network engineering, hardware operations, and vendor teams.
  • Drive neo cloud delivery programs by managing capacity delivery from GPU cloud and neo cloud partners, including contract milestones, capacity ramps, service-level agreements, and go-live readiness.
  • Build and maintain master delivery schedules across concurrent, multi-site, multi-vendor programs, integrating power and shell timelines, hardware lead times, logistics, and software and platform readiness into a single critical path.
  • Manage cross-functional stakeholders across engineering, data center operations, procurement, finance, legal, and executive leadership, translating technical dependencies into clear program status, risks, and decisions.
  • Own infrastructure-delivery vendor and partner relationships, including neo cloud providers, colocation partners, hardware original equipment manufacturers, and logistics and installation contractors; hold partners accountable to contractual delivery commitments.
  • Identify critical-path risks, including power availability, hardware lead times, permitting, and supply-chain issues; drive mitigation and escalate proactively.
  • Establish delivery governance through standardized program frameworks, milestone definitions, risk, action, issue, and dependency logs, capacity tracking, and reporting cadences, including weekly operations reviews and executive or board-level updates.
  • Align capacity planning with delivery timelines and business commitments, including customer contracts and internal growth targets.
  • Lead post-mortems and continuous-improvement retrospectives after major builds or delivery milestones to improve the cost, speed, and predictability of future builds.
Desired Qualifications
  • Experience at a hyperscaler, neo cloud provider, colocation company, or AI infrastructure company.
  • Familiarity with GPU hardware lifecycles, including NVIDIA H100, H200, and GB200 class systems, power and thermal constraints, and compute supply-chain dynamics.
  • PMP, PgMP, or equivalent program-management certification.
  • Experience building delivery or program-management functions from the ground up.

About the company

Baseten provides a machine learning infrastructure platform for deploying, serving, and scaling AI models in production. Its Inference Stack lets teams deploy custom or open-source models as scalable APIs with features like automatic scaling, resource and version management, and observability. It supports multi-cloud deployments (AWS, GCP, or a client’s own cloud) and includes Baseten Training for containerized training jobs and Baseten Model APIs for quick prototyping. Baseten aims to simplify the end-to-end lifecycle of AI models in production by handling deployment, scaling, and management, with tiered plans to fit different customer needs.

Company Size

201-500

Company Stage

Series F

Total Funding

$2.1B

Headquarters

San Francisco, California

Founded

2019

Get referred to Baseten

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • June 2026 Series F raised $1.5 billion at a $13 billion valuation.
  • Bloomberg reported September 2026 talks for a $26 billion round, signaling intense demand.
  • Company said June 2026 revenue grew 20x, while inference volume grew 40x.

What critics are saying

  • Baseten remains Nvidia-dependent; GPU shortages and pricing squeeze margins during 2026 capacity expansion.
  • AWS, Google, OpenAI, and Hugging Face can bundle competing inference stacks faster.
  • Open-weight models commoditize inference pricing, turning Baseten into a low-margin compute reseller by 2027.

What makes Baseten unique

  • September 2026 Blaxel acquisition unifies inference, training, and agent sandboxes in one platform.
  • August 2026 Hugging Face added Baseten as a serverless inference provider.
  • Base Labs launched September 2026 with Hugging Face and Goodfire on open-weight safety.

Help us improve and share your feedback! Did you find this helpful?

Benefits

💰 Competitive compensation: We aim to provide 90th percentile (or better) salaries and equity grants for every team member commensurate with their experience.

🌎 Remote-first work environment: The Baseten team is welcome to work from wherever they want; fully remote, in our San Francisco office, or a mix of both. We provide a $1,000 stipend for you to make your home office comfortable and productive.

🏓 Regular in-person team summits: We get together as a team three times a year to plan, workshop, and most importantly, get to know each other better.

🌴 Unlimited PTO: We ask that everyone take at least 4 weeks of vacation. And we have a company-wide break between Christmas and New Year's Day.

🏥 Full healthcare coverage: Medical, dental and vision insurance for you and your family.

🍼 Paid parental leave: 16-weeks fully paid parental leave (adoptive and non-birth parents included) and flexibility with schedules while returning to work.

📈 401(k): Company-sponsored 401(k) for you to contribute to.

🧠: Learning and development budget: We encourage you to take classes, attend conferences, and invest in your craft and we’ll cover expenses to make it happen.

Growth & Insights and Company News

Headcount

6 month growth

↓ -3%

1 year growth

↓ -2%

2 year growth

↓ -4%
The Information
Sep 25th, 2026
Figma CMO Sheila Vashee joins AI inference startup Baseten

Sheila Vashee, former chief marketing officer at Figma, has joined AI inference startup Baseten. Vashee spent three years at the design software firm before departing this summer during an executive shakeup. Her exit in August contributed to a decline in Figma's stock price. Baseten is described as a rapidly growing upstart in the AI infrastructure space, though specific details about the company's recent growth or funding were not provided in the source material.

Bloomberg
Sep 23rd, 2026
Startups Modal, Baseten in funding talks to help businesses run AI.

Startups Modal, Baseten in funding talks to help businesses run AI. September 23, 2026 at 3:28 PM PDT Takeaways by Bloomberg aisubscribe. Two startups that provide platforms to help businesses run artificial intelligence models are in talks to raise capital that could at least double their valuations, a sign of surging investor demand for infrastructure services. Modal Labs is in talks to raise new financing at a roughly $15 billion valuation, according to two people familiar with the matter, roughly triple what it was worth in a round four months ago. Baseten is in talks for a funding round that could value it at $26 billion, people said, up from $13 billion in June. The firms are each tapping into the market for inference services, or the process of running AI systems after they've been trained. By some estimates, spending on chips and computing for inference is expected to eclipse that of training AI services, reflecting stronger adoption of the technology and the shift to more computationally intensive tools like agents. At the same time, the growing number of startups providing inference platforms require more capital in order to secure the computing power and chips needed to offer their services to customers. Baseten and Modal declined to comment. The people described the discussions on condition of anonymity as the information is not public. Axios of the Baseten round. Baseten, founded in 2019, delivers software and computing capacity to other startups. For example, Baseten offers its customers "the ability to run these models when they don't necessarily have access to the infrastructure teams of the frontier labs doing it for them," said Tuhin Srivastava, the company's chief executive officer, in a previous interview with Bloomberg. Modal, which was founded in 2021, lets customers run inference tasks using both graphics chips like those from Nvidia Corp. and regular computer chips. At least some of the demand for AI inference startups may be due to the rise of cheap, effective open-source models from China and other markets, according to Dr. Lan Xuezhao, founder and managing partner of venture firm Basis Set in a previous interview. More companies are turning to these startups to help power their own efforts to build, refine and run internal models on open-source technology, while seeking the infrastructure to support it, Xuezhao said.

AI Cyber Australia
Sep 17th, 2026
Baseten unveils safety framework for open-weight AI models in collaboration with Hugging Face and Goodfire AI.

Baseten unveils safety framework for open-weight AI models in collaboration with Hugging Face and Goodfire AI. On Wednesday, Baseten, in partnership with Hugging Face and Goodfire AI, launched a new safety infrastructure standard via its research arm, Base Labs. This initiative addresses safety concerns tied to open-weight models, which are models that can become hazardous when safeguards are removed through a process called abliteration. Hugging Face currently hosts over 6,000 such models. Key Points: * Base Labs aims to establish transparent standards for training and deploying open models, integrating safety directly rather than adding it later. * Baseten emphasizes the advantages of openness for AI safety, promoting the visibility and transparency of model behaviors. * The partnership with Goodfire aims to embed safety within open models; Goodfire focuses on making AI processes more interpretable. * Baseten and Goodfire are both significantly funded; Baseten recently raised $1.5 billion, while Goodfire secured $150 million. * Baseten is inviting contributions from the developer community to bolster this framework and ensure open models are safely accessible. The announcement is part of a broader effort to create safer, more transparent AI environments as concerns about the potential misuse of AI continue to grow.

Intelpro
Sep 17th, 2026
Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire.

Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire. Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models. Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models. The announcement lands amid debate for the safety of open-weight models - which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models. Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models. The company is framing their future work as a "standard" for open models that is transparent and built into how models are trained and deployed, rather than bolted on afterward. "We believe openness to be an advantage for AI safety," the company said on X. "Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source." The companies haven't disclosed how the partnership will work technically, though Goodfire framed the goal in a reply to Baseten's post: "Safety must be built into open models and provided by those who serve them." Goodfire, which specializes in opening AI's "black box" to explain how models make decisions, is the likeliest candidate for the "built into" part. Baseten, an AI inference provider, raised a $1.5 billion Series F in June, vaulting its valuation to $13 billion. Partner Goodfire AI is similarly well-capitalized, having raised a $150 million Series B led by B Capital earlier this year to advance its model interpretability platform. Looking ahead, Baseten is putting out an open call to the broader developer ecosystem to contribute to the framework. "Together, we are building an ecosystem of open models that are safe and accessible to all," the company noted.

CR Bookings
Sep 17th, 2026
Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire.

Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire. admin September 17, 2026 Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for... Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models.