Full-Time

Senior Software Engineer

Network Observability

Clockwork Systems

Clockwork Systems

1-10 employees

Advanced clock synchronization for distributed systems

Compensation Overview

$140k - $210k/yr

+ Stock Options

Palo Alto, CA, USA

In Person

Bachelor's, Master's

Category
Software Engineering (1)
Required Skills
TCP/IP
Datadog
Rust
Python
Grafana
Distributed Systems
Computer Networking
OpenTelemetry
Go
Prometheus
C/C++
Splunk
Linux/Unix

Get referred to Clockwork Systems

See people who can refer or advise you

Requirements
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • Strong hands-on programming experience in C++, Go, Python, Rust, or similar systems programming languages.
  • Proven experience delivering complex infrastructure projects and contributing to the design and implementation of distributed systems.
  • Experience building distributed systems, backend services, telemetry pipelines, or observability platforms.
  • Hands-on experience with RDMA, RoCE, InfiniBand, or other high-performance network fabrics.
  • Familiarity with libibverbs, RDMA verbs, RDMA CM, queue pairs, completion queues, memory registration, and related RDMA concepts.
  • Strong knowledge of Linux networking, TCP/IP, DNS, HTTP, routing, MTU, congestion control, packet loss, latency, and performance tuning.
  • Experience with traceroute-style diagnostics, path discovery, network reachability checks, synthetic probes, or active network measurements.
  • Experience with monitoring and visualization platforms such as Prometheus, Grafana, Datadog, Splunk, OpenTelemetry, or similar tools.
  • Strong debugging skills across software, operating system, server, and network layers.
  • Experience operating production systems in Linux-based environments.
  • Strong technical judgment and ability to build reliable, scalable, and maintainable systems.
Responsibilities
  • Design, develop, and scale high-performance network monitoring platforms for RDMA, RoCE, InfiniBand, and TCP/IP infrastructure.
  • Build backend telemetry services, observability dashboards, alerts, diagnostics, anomaly detection, SLA monitoring, and traffic analysis workflows.
  • Troubleshoot complex production issues across application, OS, server, RDMA, and network layers while optimizing low-latency collection, aggregation, and alerting.
  • Collaborate with cross-functional teams to deliver scalable solutions, improve engineering practices, automate operational workflows, and contribute to the technical direction of the platform.
Desired Qualifications
  • Experience supporting AI/ML, HPC, storage, or GPU cluster infrastructure workloads.
  • Experience with large-scale RoCE or InfiniBand deployments.
  • Experience with NCCL, distributed training infrastructure, or AI cluster diagnostics.
  • Experience with eBPF, XDP, DPDK, perf, tcpdump, Wireshark, ethtool, iproute2, rdma-core, or Linux kernel networking tools.
  • Experience with cloud infrastructure on AWS, GCP, or Azure.
  • Experience with Kubernetes, service discovery, configuration management, and infrastructure automation.
  • Knowledge of security, compliance, and infrastructure best practices.
  • Experience designing time-series data systems, alerting pipelines, or high-cardinality telemetry platforms.

Clockwork Systems provides clock synchronization technology for mission-critical distributed systems, ensuring precise timing across operations. Its solutions run in both cloud and on-premises environments and are delivered through software licensing, subscriptions, and professional services. The company differentiates itself with deep timing expertise and end-to-end timing across networks, framed by tools like Latency Sensei for cloud latency monitoring. Its goal is to help customers achieve reliable, accurate synchronization to boost performance and reduce timing-related issues in time-sensitive applications.

Company Size

1-10

Company Stage

Early VC

Total Funding

$41.6M

Headquarters

Palo Alto, California

Founded

2018

Get referred to Clockwork Systems

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Clockwork raised $20.6M in September 2025, extending runway for product expansion.
  • Oracle Cloud marketplace availability and AWS private beta broaden distribution beyond early adopters.
  • Customer proof from Uber, eBay, Wells Fargo, and RBC supports enterprise credibility.

What critics are saying

  • NVIDIA, hyperscalers, and networking vendors can bundle adjacent observability and resilience features by 2027.
  • Enterprise buyers will scrutinize YOCO credits after any migration shortfall, pressuring renewals.
  • If FleetIQ and TorchPass fail to create durable demand, Clockwork remains a niche timing vendor.

What makes Clockwork Systems unique

  • Clockwork sells software-only timing and fault-tolerance across cloud, on-prem, and hybrid clusters.
  • Its 2025-2026 stack combines Latency Sensei, Cloud Deluxe, FleetIQ, and TorchPass.
  • YOCO contracts promise 90% failure recovery with no lost training progress or recompute.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Competitive Salary

Company News

Associated Press
Mar 11th, 2026
Clockwork.io launches TorchPass to eliminate GPU failure waste in AI training, saving $6M per 2,048-GPU cluster

Clockwork.io has launched TorchPass Workload Fault Tolerance, a software solution that eliminates costly GPU training failures through Live GPU Migration technology. The system allows AI training workloads to continue running through hardware failures, network disruptions and node crashes without requiring checkpoint restarts. The company claims TorchPass can save over $6 million annually in a typical 2,048-GPU deployment by reducing wasted training progress by 95%. In large clusters, it cuts lost time from approximately three hours per day to under ten minutes. Independent testing by SemiAnalysis found TorchPass delivered faster fault-tolerant performance than standard checkpoint-restart approaches and higher Model FLOPs Utilisation than leading open-source alternatives. The solution typically completes recovery in approximately three minutes whilst training continues uninterrupted. TorchPass is now available as part of Clockwork.io's FleetIQ platform.

The SaaS News
Sep 11th, 2025
Clockwork Raises $20.57 Million in Funding | The SaaS News

Clockwork Raises $20.57 Million in Funding

TechStartups
Sep 10th, 2025
Clockwork Systems Raises $20.6M for FleetIQ

Stanford spinout Clockwork has raised $20.6 million, led by NEA with participation from notable investors, to address AI's GPU inefficiency. The funding coincides with the launch of FleetIQ, a software solution aimed at enhancing GPU performance by improving communication between GPUs, clusters, and clouds. This innovation seeks to reduce crashes, shorten restarts, and increase utilization rates, making AI infrastructure more efficient and sustainable.