Full-Time

Account Executive

FriendliAI

FriendliAI

51-200 employees

Serverless generative AI inference platform

No salary listed

San Francisco, CA, USA

Hybrid

Hybrid work arrangement in San Francisco.

Bachelor's

Category
Sales & Account Management (1)
Required Skills
LLM
MLOps
Sales
Machine Learning
AWS

Get referred to FriendliAI

See people who can refer or advise you

Requirements
  • At least 5 years of business-to-business sales experience.
  • A proven track record of closing deals with annual recurring revenue of at least $300,000 and consistently exceeding revenue targets.
  • A strong technical background in artificial intelligence and machine learning infrastructure, cloud platforms, or developer tools.
  • Deep understanding of large language model deployment, inference serving, and artificial intelligence and machine learning workflows.
  • Active participation in artificial intelligence and machine learning technology communities with demonstrated networking abilities.
  • Experience managing technical proofs of concept and multi-stakeholder enterprise sales cycles.
  • Excellent communication skills for engaging both C-level executives and technical teams.
Responsibilities
  • Own the end-to-end enterprise sales process, from generating pipeline through closing high-value strategic deals focused on artificial intelligence inference and serving solutions.
  • Identify and capitalize on growth opportunities within existing accounts and serve as a trusted advisor for artificial intelligence infrastructure optimization.
  • Manage technical proofs of concept demonstrating performance and cost advantages, collaborating with internal engineering teams and enterprise clients.
  • Generate and manage a robust pipeline by identifying enterprises with significant inference workloads and strategically positioning the company's solutions.
  • Lead technology community engagements and represent the company at artificial intelligence and machine learning conferences, meetups, and industry events.
  • Cultivate strategic relationships with artificial intelligence engineers, machine learning platform leaders, and technical decision-makers at target enterprises through community involvement.
  • Organize technical events, workshops, technical demonstrations, and meetups showcasing the company.
  • Collaborate on technical content, case studies, and speaking opportunities that establish the company's market leadership.
  • Bring the voice of enterprise customers into product discussions so the product roadmap aligns with market needs and competitive positioning.
  • Work closely with engineering teams to understand and articulate complex technical differentiators related to inference optimization.
  • Analyze the competitive market landscape and provide insights on enterprise artificial intelligence infrastructure trends and requirements.
Desired Qualifications
  • Experience at an artificial intelligence infrastructure, MLOps, or developer tools company.
  • An existing network within enterprise artificial intelligence and machine learning teams.
  • A technical degree or equivalent practical experience in the artificial intelligence and machine learning space.
  • A track record of organizing or speaking at technology meetups and industry conferences.

Preparing concise company summary.

Company Size

51-200

Company Stage

Early VC

Total Funding

$26.7M

Headquarters

Redwood City, California

Founded

2021

Get referred to FriendliAI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Aolani signed September 10, 2026, adding GPU capacity for Asia demand.
  • May 2026 San Francisco expansion targets faster sales to Bay Area AI customers.
  • April 2026 Brian Yoo and Samsung B300 access strengthen go-to-market and supply.

What critics are saying

  • FriendliAI raised only $20 million seed funding in August 2025, limiting war-chest depth.
  • Samsung SDS and Aolani partnerships increase dependence on third-party GPU supply in 2026.
  • vLLM, Fireworks, and Together commoditize inference pricing, compressing margins before 2027.

What makes FriendliAI unique

  • Byung-Gon Chun pioneered continuous batching, now standard across AI inference serving.
  • FriendliAI controls kernels, scheduling, caching, and distribution, not just model hosting.
  • OpenRouter and Artificial Analysis rank FriendliAI among the fastest providers in 2026.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Health Insurance

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

-1%

1 year growth

0%

2 year growth

-4%
Headline
Sep 10th, 2026
Aolani and FriendliAI partner to advance AI inference at scale.

Aolani and FriendliAI partner to advance AI inference at scale. Sep 10, 2026 SINGAPORE - Media OutReach Newswire - 10 September 2026 - Aolani, a Singapore-founded neocloud powering AI growth, today announced a partnership to supply GPU cloud infrastructure to FriendliAI, the San Francisco-headquartered inference cloud for frontier AI to support the rapidly growing demand for inference services. The global market for AI inferencing is expanding quickly as AI applications become part of everyday business workflows and organisations move from experimentation to deployment at scale. At the forefront of production-scale AI, FriendliAI serves this exact demand to help developers and enterprises deploy open-weight and custom AI models. FriendliAI was founded by researchers who invented continuous batching, which is now a standard across AI inference serving. The company has built its inference stack end to end, from optimised GPU kernels to global distribution, so that production AI workloads run fast and reliably at scale. FriendliAI consistently ranks as one of the fastest inference providers on OpenRouter, with enterprise clients including LG, Kilo Code, and Liner running their production inference on the platform. Efficient time-to-value and dependable compute are increasingly important to keep services responsive as usage grows. As access to reliable compute infrastructure becomes a strategic differentiator for companies scaling production workloads, more AI natives are turning to Asia for high-performance compute capacity, attracted by the region's expanding digital infrastructure, strategic connectivity, and growing AI ecosystem. As one of the leading neoclouds offering purpose-built next-generation AI infrastructure, Aolani helps AI natives scale more efficiently. Aolani's infrastructure capabilities across orchestration, automation and lifecycle management actively supports FriendliAI's services. This partnership equips FriendliAI with the compute to serve the rapid customer demand, both across the globe and increasingly in Asia. Nicholas Chia, Chief Executive Officer at Aolani said: "We're seeing inference needs grow faster than companies can find compute to support and service their customers. To narrow the supply and demand gap, we actively partner with companies like FriendliAI to deliver compute capacity on time, at scale, and to rigorous standards. We look forward to partnering with the FriendliAI team to grow its services to bring fast and reliable inference to developers worldwide." Byung-Gon Chun, Founder and CEO of FriendliAI said: "We are seeing exponential growth in demand for our frontier AI inference services. Businesses need the freedom to choose the AI models that best suit their applications and the ability to run them efficiently in production. Our job is to deliver high-performance, reliable inference so developers can focus on building their AI applications. Aolani stood out as a trusted infrastructure partner that can help us scale at the pace our customers need. We look forward to working with Aolani to support our mission." Media Contact H/Advisors on behalf of Aolani Hashtag: #Aolani #Neocloud #FriendliAI #AIinference #Inferencecloud The issuer is solely responsible for the content of this announcement. About Aolani. Where AI gets built in Asia. Founded in Singapore, Aolani is backed by compliant and purpose-built infrastructure to deliver the performance capabilities for next-generation AI. Aolani's AI factories enable organisations to build with confidence, scale ambitiously, and move at hyper-speed in the world's fastest-growing AI market. For more information, visit www.aolanicloud.com and follow @AolaniCloud on LinkedIn. About FriendliAI. FriendliAI is the inference cloud for frontier AI. Headquartered in San Francisco with a team in Seoul, FriendliAI runs open models in production at scale for AI-native startups and enterprises through its Model APIs, Dedicated Endpoints, and Bring Your Own GPU (BYOG) offering. The team built the full inference stack end to end, delivering the speed, reliability and efficiency that agentic AI workloads demand. For more information, visit friendli.ai. Looking for Local Media Coverage in the United States of America? We have a place for all 50 States at the State News Network

FriendliAI
Jul 1st, 2026
How Kilo Code and FriendliAI bring open source AI coding agents to production with NVIDIA Nemotron.

How Kilo Code and FriendliAI bring open source AI coding agents to production with NVIDIA Nemotron. * Kilo and FriendliAI are partnering to make production-grade AI coding agents faster, more accurate, and significantly more cost-efficient with NVIDIA Nemotron open models. * Configure Nemotron 3 Ultra in Kilo Code using FriendliAI as your inference provider, directly inside your IDE. * Using FriendliAI, Kilo achieved up to 7x faster inference compared to several other providers and cut costs by up to 72% on complex agent workloads by routing to NVIDIA Nemotron 3 Ultra. Introduction. Kilo and FriendliAI are partnering to help engineers build AI agents at scale with NVIDIA Nemotron open models. Production agents thrive when built using specialized models, both closed and open, and its priority is to make developing high-performance agents accessible, performant, and cost-effective. NVIDIA is a leader in open models that enable developers, enterprises and nations to build AI applications and agents that they can trust, control, and customize. NVIDIA open Nemotron family combines strong reasoning performance with efficient deployment, providing open weights, training data, and recipes for building specialized AI agents. The Nemotron family offers reasoning models in 3 sizes: Nano, Super, and Ultra for different deployment needs. Nano provides cost efficiency with high accuracy specialized sub-agents, Super delivers highest efficiency with leading accuracy for reasoning and tool calling for multi-agent applications, and Ultra is designed for applications demanding the highest reasoning accuracy for complex agentic tasks. Nemotron 3 Ultra, newest member of the Nemotron family, is a 550B-parameter Mixture-of-Experts model with 55B active parameters, built for frontier reasoning and orchestration in agentic systems. Kilo is an all-in-one agentic engineering platform for software developers. Their open-source coding harness, Kilo Code, is a flexible AI coding assistant centered on model freedom - giving developers the ability to plug in their existing model providers alongside open-weight models. You can experience top open models on Kilo Code, like Nemotron 3 Ultra, directly in your terminal or VS Code. FriendliAI sits directly within Kilo's orchestrator and maximizes efficiency. Its inference platform is optimized for real-time coding workloads, giving Kilo Code the speed and efficiency it needs to scale. Through continuous batching, also known as iteration batching, and memory optimization, FriendliAI enables Kilo Code to feed massive multi-file codebases into NVIDIA Nemotron models without degrading performance, running out of memory, or spiking costs. Together, Kilo, FriendliAI, and NVIDIA provide an open stack for building production AI coding agents - from intelligent orchestration and optimized inference to frontier reasoning. In this blog, FriendliAI Inc. will walk through configuring Kilo Code with FriendliAI and Nemotron 3 Ultra in your existing workflow. Configure NVIDIA Nemotron 3 Ultra in Kilo Code with FriendliAI. Kilo Code is built around a simple idea: developers should be free to choose the models and infrastructure that best fit their workflows. Rather than locking teams into a single AI provider, Kilo Code enables developers to connect their preferred inference platforms and open models directly into their coding environment. By combining Kilo Code, NVIDIA Nemotron 3 Ultra, and FriendliAI, developers gain access to a frontier reasoning model optimized for complex software engineering tasks while maintaining the performance, scalability, and cost efficiency required for production workloads. Connecting FriendliAI to Kilo Code. Getting started requires only a few configuration steps: * Deploy Nemotron 3 Ultra on FriendliAI Begin by deploying Nemotron 3 Ultra through the Friendli Suite. Once deployed, you'll receive an Endpoint ID and API credentials that can be used by external applications. 2. Configure FriendliAI as a Custom Provider in Kilo Code Within Kilo Code, navigate to the Provider settings and select a Custom Provider configuration. Enter your FriendliAI endpoint URL, Endpoint ID, and API key to connect your deployment. 3. Start Building with Nemotron 3 Ultra After configuration is complete, Nemotron 3 Ultra becomes available directly within Kilo Code's model selector. Developers can immediately begin using the model for code generation, repository analysis, debugging, tool use, long-context reasoning, and multi-step agentic workflows - all without leaving their IDE. With the integration complete, Kilo Code can route requests directly to your FriendliAI deployment, providing access to Nemotron 3 Ultra inside your existing development workflow. Why FriendliAI for Nemotron 3. Large reasoning models are most valuable when they can process substantial context, reason across multiple files, and respond quickly enough to keep developers in flow. Running these workloads efficiently requires more than simply hosting a model. Production inference depends on optimized scheduling, batching, memory management, and efficient GPU utilization. FriendliAI's inference platform is optimized for high-throughput, low-latency agentic workloads. This enables coding agents powered by NVIDIA Nemotron 3 Ultra to work with large codebases and long-context prompts while maintaining responsiveness and controlling infrastructure costs. How Kilo uses FriendliAI to optimize cost and performance. Over the past year, Kilo Code has tested several different inference providers hosting a range of both open and proprietary models. According to Kilo's internal evaluations using GLM-5 usage, they were especially impressed by FriendliAI, which consistently delivered up to 7x faster inference than several other providers while significantly reducing error rates. FriendliAI is now a core component of the Kilo stack, enabling high-performance access to the latest open models. To further reduce friction, Kilo developed an auto-routing feature that automatically selects the optimal model for specific tasks like planning, coding, or data analysis. This routing is organized into four distinct modes: * Auto: Frontier, which offers maximum capability with the best available models when cost is not an issue. * Auto: Balanced, which offers strong performance at a lower cost. * Auto: Efficient, which offers the lowest cost per task, with capability matched to difficulty. * Auto: Free, which features the best available free models. These modes leverage models from labs like OpenAI and Anthropic alongside open models from MiniMax, Z.ai, Alibaba Qwen, and NVIDIA. The centerpiece of this collaboration is NVIDIA Nemotron 3 Ultra, an open model built for long-running agents. Nemotron 3 Ultra, recently led the open-weight category of PinchBench leaderboard, a benchmark that evaluates models' performance on real-world agentic tasks in OpenClaw. By intelligently routing complex reasoning tasks to Nemotron 3 Ultra, Kilo reports cost reductions of up to 72% on complex agentic coding tasks while maintaining frontier-level performance. As agentic engineering continues to evolve to support an even wider range of tasks, Kilo is excited to continue working with both FriendliAI and NVIDIA to optimize cost and performance hand-in-hand. This strategic partnership integrates Kilo Code's sophisticated agent orchestration layer with FriendliAI's optimized inference platform and NVIDIA's remarkably flexible Nemotron open models. There's no longer a reason to overpay for dependable AI. Build the stack that works for you. Production AI coding agents work best when each layer of the stack handles what it's built for. Kilo Code orchestrates across models intelligently, FriendliAI delivers production-grade inference performance that keeps developers productive. NVIDIA Nemotron 3 Ultra open model provides advanced reasoning that you can customize, fine-tune, and deploy for specialized AI agents. Together, they're already powering real results - Kilo users are running 7x faster on FriendliAI and cutting costs by up to 72% on complex agentic tasks by routing to Nemotron 3 Ultra. Open, optimized AI development is no longer a trade-off between performance and price. Together, the three layers give developers everything needed to ship production-grade AI coding agents without compromising on performance or cost, while leaving more room for model freedom. Explore NVIDIA Nemotron 3 Ultra, bring frontier open models directly into your IDE with Kilo Code, and get started with FriendliAI to deploy and serve your favorite models with optimized inference.

Business Wire
May 11th, 2026
FriendliAI opens San Francisco office to scale frontier AI inference for open-weight models

FriendliAI, a frontier AI inference cloud provider, has opened a 7,000-square-foot San Francisco office at 20 Hawthorne Street, placing the company closer to Bay Area AI customers and developers. The expansion comes as AI agents require five to 30 times more tokens per task than chatbots, whilst open-weight models now match closed models at lower costs. Founded by Seoul National University Professor Byung-Gon Chun, FriendliAI pioneered continuous batching, now an industry standard for inference optimisation. Independent benchmarks from Artificial Analysis and OpenRouter rank FriendliAI as the top inference provider for models including GLM-5.1 and Gemma 4. The company is on track to grow revenue tenfold this year, targeting another tenfold increase next year. FriendliAI plans significant US team expansion across go-to-market, partnerships and engineering functions.

Business Wire
Apr 15th, 2026
FriendliAI partners with Samsung Cloud Platform to deliver frontier AI inference on NVIDIA B300 GPUs

FriendliAI is partnering with Samsung SDS to deliver frontier model AI inference services using Samsung Cloud Platform's NVIDIA B300 GPU infrastructure. The collaboration combines FriendliAI's optimised inference stack with Samsung SCP's scalable GPU resources to serve global startups and enterprises. The partnership will enable customers to run frontier open-weight models, including GLM MiniMax M2.5, NVIDIA Nemotron 3 Super and DeepSeek v3.2, with production-grade reliability. FriendliAI's platform offers speeds up to three times faster than vLLM and 50% to 90% cost savings compared to closed model APIs. Customers will access the service through FriendliAI's Serverless Endpoints with token-based pricing, whilst benefiting from Samsung SCP's high-availability GPU infrastructure. The collaboration aims to maximise GPU utilisation whilst delivering cost-effective scalability for AI workloads.

FriendliAI
Apr 14th, 2026
FriendliAI and Samsung Cloud Platform forge strategic alliance to power frontier Model AI Inference on NVIDIA B300 GPUs.

FriendliAI and Samsung Cloud Platform forge strategic alliance to power frontier Model AI Inference on NVIDIA B300 GPUs. * FriendliAI is collaborating with Samsung SDS, a leading GPU infrastructure-as-a-service (IaaS) provider in South Korea, to deliver frontier model AI inference services to global startups and enterprises. * The initiative brings together FriendliAI's high-performance inference stack and Samsung Cloud Platform (SCP)'s scalable NVIDIA B300 GPU infrastructure. FriendliAI is collaborating with Samsung SDS, a leading GPU infrastructure-as-a-service (IaaS) provider in South Korea, to deliver frontier model AI inference services to global startups and enterprises. The initiative brings together FriendliAI's high-performance inference stack and Samsung Cloud Platform (SCP)'s scalable NVIDIA B300 GPU infrastructure. FriendliAI is known for building an AI inference platform that delivers unmatched speed, cost efficiency, and reliability. By integrating FriendliAI's platform with Samsung SCP's B300 GPU infrastructure, the two companies will enable customers to run the latest frontier open-weight models - including GLM-5.1, MiniMax M2.5, NVIDIA Nemotron 3 Super, and DeepSeek v3.2 at maximum performance with production-grade reliability and competitive token pricing. Bringing Frontier Open-Weight Model AI Inference to Global Companies The partnership is designed to provide organizations worldwide with a seamless and scalable experience when deploying frontier AI models. Customers will be able to access FriendliAI's inference platform powered by Samsung SCP's NVIDIA B300 GPU IaaS, combining state-of-the-art inference optimization with high-performance GPU infrastructure. Key benefits include: Day-0 Support for Frontier Open-Weight Models: Immediate support for models such as GLM-5.1, MiniMax M2.5, NVIDIA Nemotron 3 Super, and DeepSeek v3.2 enables companies to stay at the forefront of AI innovation without building and maintaining custom inference stacks. High-Performance Inference: FriendliAI's optimized platform is built on its in-house inference engine featuring custom kernels, advanced quantization techniques, intelligent caching, and adaptive speculative decoding. Cost-Effective Scalability: The partnership enables high-speed inference with low operational cost and maximum scalability. Friendli Serverless Endpoints provide a flexible, token-based pricing model and pay only for the tokens they process. Global Reach and Reliability: FriendliAI's inference infrastructure, integrated with Samsung SCP's B300 IaaS platform, provides a global footprint with high-availability guarantees, ensuring low-latency and stable service delivery for customers worldwide. "FriendliAI is thrilled to partner with Samsung SCP to empower AI inference for companies around the world," said Byung-Gon Chun, Founder and CEO of FriendliAI. "Our inference optimization technology, built for unmatched speed and cost efficiency, perfectly complements Samsung SCP's cutting-edge NVIDIA B300 GPU cloud infrastructure. This collaboration allows us to deliver the latest frontier open-weight models at full performance, helping customers unlock new business opportunities with agentic AI." EUNYOUNG KIM, Executive Vice President at Samsung SDS, added, "Our mission at Samsung SCP is to provide flexible and high-performance GPU infrastructure for the AI era. Partnering with FriendliAI enables us to bring state-of-the-art AI inference capabilities to global customers. FriendliAI's platform ensures that our NVIDIA B300 GPUs are fully utilized, offering clients the optimal combination of performance, cost efficiency, and reliability for frontier AI workloads." About FriendliAI FriendliAI is The Frontier AI Inference Cloud. Built by the researchers who invented continuous batching, FriendliAI's highly optimized inference engine efficiently runs state-of-the-art open-weight and custom models at production scale with 99.99% reliability. By maximizing GPU utilization, FriendliAI delivers speeds up to 3x faster than vLLM and 50% to 90% cost savings relative to closed model APIs, empowering engineers to deploy frontier AI with uncompromising speed and model ownership. Learn more at https://friendli.ai About Samsung Cloud Platform (SCP) Samsung Cloud Platform is a leading GPU infrastructure-as-a-service (IaaS) operating clusters based on NVIDIA B300 GPUs. Samsung SCP offers scalable, high-availability GPU resources designed to accelerate compute-intensive workloads such as AI inference and training. Its mission is to empower organizations with flexible, cost-effective infrastructure that keeps pace with the rapid evolution of artificial intelligence. Learn more at https://cloud.samsungsds.com/