Full-Time

Software Engineer

Infrastructure

Thunder Compute

Thunder Compute

11-50 employees

Cloud-based GPU provider for developers

No salary listed

H1B Sponsorship Available

San Francisco, CA, USA

In Person

In-person at downtown San Francisco office; relocation support and visa sponsorship available.

Category
Software Engineering (1)
Required Skills
Kubernetes
Python
Computer Networking
Docker
TypeScript
Go
Next.js
Observability
Linux/Unix

Get referred to Thunder Compute

Find people who can refer or advise you

Requirements
  • Exceptional Go ability, including concurrency, distributed systems design, API design, and production service development
  • Deep understanding of Kubernetes, containers, Linux, networking, storage, or cloud infrastructure
  • Experience building and operating critical production systems
  • Strong systems debugging and operational ability
  • Ability to reason through unfamiliar systems across multiple layers of the stack
  • Working knowledge of Python; familiarity with TypeScript or Next.js is helpful
  • Strong work ethic and the ability to independently push a project from an experimental prototype through 100% completion under tight deadlines
  • Attention to detail and the ability to deliver production-ready, thoroughly tested code without significant oversight
  • Strong ownership over correctness, reliability, performance, and operational outcomes
  • Ability to debug ambiguous problems without a clear reproduction, existing playbook, or obvious owner
  • Willingness to work directly with customers and investigate difficult production failures
  • Strong communication skills and the ability to coordinate across engineering, customers, and external infrastructure providers
Responsibilities
  • Building control-plane services for provisioning and managing virtual GPU instances
  • Designing reliable systems for GPU allocation, scheduling, and lifecycle management
  • Improving our unconventional Kubernetes deployment, which acts as a form of hypervisor for customer workloads
  • Building infrastructure for networking, storage, authentication, billing, and usage metering
  • Automating the deployment and operation of GPU hosts across cloud providers and customer data centers
  • Debugging failures across customer workloads, Kubernetes, our control plane, the network, and physical GPU infrastructure
  • Designing systems for failure recovery, capacity management, observability, and incident response
  • Improving the security, reliability, and operational simplicity of the platform as it scales
  • Working directly with customers to diagnose problems and deploy Thunder Compute in new environments
  • You will spend your days bouncing between the weeds of complicated production infrastructure that is live and used by customers
  • One week, you may be debugging a networking failure across a Kubernetes cluster; the next, you may be redesigning the provisioning system to make deployments faster and more reliable
Desired Qualifications
  • Experience with Kubernetes internals, container runtimes, cloud networking, distributed storage, infrastructure security, or large-scale control planes
  • Experience building high-stakes production infrastructure at a trading firm such as Citadel Securities or Jane Street; a cloud provider such as AWS, CoreWeave, or Lambda; an AI infrastructure company; or a similarly demanding engineering environment
  • Strong computer science fundamentals demonstrated through academic work, distributed systems research, open-source contributions, or exceptional professional experience
  • Experience designing and operating infrastructure across multiple cloud providers or on-premise environments
  • Experience taking a new infrastructure system from an early prototype into a reliable production platform

Thunder Compute provides cloud-based GPU resources that developers and businesses can access through a pay-as-you-go model. Users run code on high-end GPUs in the cloud, with seamless integration to local development workflows so no special tools are required. It differentiates itself by leveraging a global network of consumer GPUs optimized for performance and by offering easy integration to existing tools to address GPU shortages. The goal is to make powerful GPU resources affordable and accessible, letting users scale compute up or down as needed.

Company Size

11-50

Company Stage

Seed

Total Funding

$5.1M

Headquarters

Atlanta, Georgia

Founded

2024

Get referred to Thunder Compute

Find people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Monthly growth over 100% shows strong developer adoption for cost-effective scalable GPU access.
  • Pay-as-you-go model with transparent billing provides instant scalability for startups and researchers.
  • Cheapest on-demand A100 and H100 GPUs undercut AWS by 68%, driving broad customer expansion.

What critics are saying

  • RunPod will undercut H100 pricing by 15-20% via per-second billing and RTX 4090 clusters.
  • Vast.ai spot model will push H100 prices below $1.80/hr by Q3 2026, hurting adoption.
  • NVIDIA export restrictions on H100/A100 by late 2026 could cause immediate inventory collapse.

What makes Thunder Compute unique

  • Proprietary GPU virtualization reduces 60-90% idle time, delivering up to 5x efficiency gain.
  • Native integration with VS Code, Cursor, PyCharm plus CLI tools enables seamless workflow automation.
  • Offers A100-80GB at $1.09/hr and H100 at $2.19/hr with transparent, stable pricing.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

Growth & Insights

Headcount

6 month growth

0%

1 year growth

-7%

2 year growth

30%