Full-Time

Model Serving Engineer

Updated on 8/1/2026

Fundamental

Fundamental

51-200 employees

Predictive analytics platform for structured data

No salary listed

Remote in USA + 1 more

More locations: Europe

Remote

Relocation support is available for employees moving to an office location.

Category
Software Engineering (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
Python
Distributed Systems
Neural Networks
Computer Networking
Docker
Observability
DevOps
Helm

Get referred to Fundamental

See people who can refer or advise you

Requirements
  • A Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • At least 5 years of experience in model serving, machine learning infrastructure, or a closely related backend engineering role.
  • Deep expertise in Python concurrency, including global interpreter lock behavior, multithreading, thread safety, and multiprocessing.
  • Experience building asynchronous and message-driven systems.
  • Experience with high-performance, large-scale distributed systems.
  • Ability to read and reason about machine learning model implementations at a computational level, including compute behavior, batching, memory usage, and inference characteristics.
  • Experience profiling and optimizing performance across CPU, memory, input/output, and ideally GPU workloads, and translating findings into architectural improvements.
Responsibilities
  • Optimize Python inference code for performance under real concurrency constraints, including global interpreter lock contention, multithreading, multiprocessing, asynchronous execution, and long-running production workloads.
  • Work closely with research to understand model internals and support the continuous evolution of the architecture, especially around complex computational behavior under production load.
  • Collaborate with research and infrastructure teams to reason about hardware utilization and serving tradeoffs across GPU, CPU, memory, networking, batching, and concurrency.
  • Define and evolve the architecture behind the distributed inference and asynchronous execution stack, including orchestration, worker coordination, and end-to-end concurrency patterns.
  • Own the Triton serving layer for NEXUS, including how models are packaged, configured, and executed as part of the production inference pipeline.
  • Build observability and performance tooling across the serving stack, and use production metrics to drive tuning decisions around latency, throughput, and resource efficiency.
  • Solve cross-cutting serving challenges that emerge from deploying the same model across environments with different scale, isolation, and reliability constraints.
  • Evaluate and integrate new inference runtimes, serving strategies, and infrastructure approaches as the model ecosystem evolves.
Desired Qualifications
  • Understanding of GPU architecture, performance characteristics, and resource utilization in high-performance compute workloads.
  • Experience working with tabular and structured-data machine learning systems.
  • Understanding of neural networks and modern deep learning architectures.
  • Experience with Kubernetes and cloud infrastructure.
  • Familiarity with DevOps and production infrastructure tooling, including containers, Helm, observability, and continuous integration and continuous delivery systems.

Fundamental builds NEXUS, a Large Tabular Model designed to analyze structured data from spreadsheets, databases, and CRM systems. It ingests raw tabular data and automatically learns patterns, enabling forecasting, classification, anomaly detection, and scenario simulation in one foundation model. The product is deployed through the AWS dashboard via a strategic partnership and trained on Amazon SageMaker HyperPod to handle datasets with billions of rows, with customers including Fortune 100 firms across financial services, healthcare, retail, and energy. The goal is to help enterprises turn raw data into actionable insights to improve forecasting accuracy, reduce risk, and optimize business outcomes, by focusing on structured data rather than unstructured text or images.

Company Size

51-200

Company Stage

N/A

Total Funding

$255M

Headquarters

San Francisco, California

Founded

2024

Get referred to Fundamental

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • NEXUS reduces data preparation time from months to days via autonomous cleaning.
  • Fundamental secured seven-figure contracts with Fortune 100 companies before public launch.
  • Single-line code integration on Amazon SageMaker enables fast deployment into existing AWS stacks.

What critics are saying

  • AWS will bundle a native SageMaker Tabular Foundation model in Q4 2026 at zero inference cost.
  • Google Cloud AutoML Tables 2.0 in August 2026 will undercut pricing for Fortune 100 contracts.
  • Palantir's Aporia module launching Q3 2026 will directly encroach on tabular forecasting use cases.

What makes Fundamental unique

  • NEXUS is a Large Tabular Model trained on billions of tabular datasets, not text.
  • It delivers deterministic, reproducible predictions unlike stochastic LLMs for compliance-heavy industries.
  • NEXUS processes tables with billions of rows natively without chunking or context loss.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Parental Leave

Relocation Assistance

Company Equity

Growth & Insights

Headcount

6 month growth

-11%

1 year growth

-11%

2 year growth

-11%