Full-Time

Principal Engineer

AI Systems & Platform Internals

Accellor

Accellor

201-500 employees

Guides transformational growth via tech solutions

No salary listed

San Francisco, CA, USA

Hybrid

Hybrid work arrangement in San Francisco.

Category
Software Engineering (1)
Required Skills
Kubernetes
Rust
Python
Grafana
Distributed Systems
Software Testing
Data Visualization
TensorFlow
CUDA
PyTorch
Machine Learning
Computer Networking
Java
OpenTelemetry
Docker
RAG
TypeScript
Microservices
Go
Prometheus
Terraform
Observability
REST APIs
C/C++
DevOps
Linux/Unix
Reinforcement Learning

Get referred to Accellor

See people who can refer or advise you

Requirements
  • 10–12 years of experience in software engineering, systems architecture, machine-learning infrastructure, distributed systems, platform engineering, inference systems, cloud infrastructure, or large-scale backend engineering.
  • Strong hands-on engineering experience with Python and at least one systems or backend language such as C++, Go, Rust, Java, or TypeScript.
  • Deep understanding of distributed systems, production infrastructure, reliability engineering, scalability, observability, and fault-tolerant architecture.
  • Experience designing or operating large-scale systems involving application programming interfaces, microservices, distributed compute, orchestration, job scheduling, caching, high-availability infrastructure, and production monitoring.
  • Strong understanding of artificial intelligence and machine-learning systems, especially model serving, inference workflows, context engineering, retrieval systems, evaluation pipelines, and production model deployment.
  • Practical understanding of GPU systems, accelerator-based workloads, CUDA/Triton-style programming, distributed inference, GPU profiling, memory optimization, and communication libraries such as NCCL or RCCL.
  • Experience with machine-learning frameworks and serving stacks such as PyTorch, JAX, TensorFlow, Triton, vLLM-style serving, Apache Ray, Kubernetes-based serving, or internal model-serving systems.
  • Ability to debug complex problems across model behavior, runtime systems, distributed infrastructure, networking, GPU execution, context quality, retrieval quality, evaluation harnesses, and production services.
  • Ability to write clear architecture documents, evaluate trade-offs, review implementation quality, and align teams around technically sound decisions.
Responsibilities
  • Design and evolve large-scale AI systems supporting ChatGPT, OpenAI API, Codex, agentic workflows, multimodal models, and research workloads.
  • Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.
  • Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.
  • Architect high-throughput, low-latency inference systems across large-scale GPU clusters.
  • Improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.
  • Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving.
  • Analyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.
  • Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement.
  • Design and guide context engineering frameworks covering prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability.
  • Design and build cost optimization frameworks for large-scale large-language-model and generative-AI workloads.
  • Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, asynchronous execution, fallback strategies, and cost telemetry across AI workflows.
  • Collaborate with research and training infrastructure teams on distributed training, checkpointing, orchestration, fault tolerance, observability, data movement, evaluation infrastructure, and experiment execution.
  • Support workflows across pre-training, post-training, reinforcement learning, agent training, evaluation harnesses, and large-scale experiments.
  • Architect validation and release systems for model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases.
  • Define release gates covering correctness, numerical stability, latency, throughput, token usage, cost regression, context quality, retrieval quality, safety behavior, reliability, and model output quality.
  • Define telemetry, tracing, dashboards, alerts, logs, profiling views, runbooks, service-level objectives, and post-incident learning loops.
  • Provide visibility into prompts, context payloads, retrieved sources, token consumption, model selection, cache behavior, inference latency, GPU utilization, evaluation scores, safety events, cost, and failures.
  • Support architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and agent deployment.
  • Work with Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams.
  • Act as a senior technical authority who resolves ambiguity, identifies systemic risks, and drives architecture decisions.
  • Mentor engineers and technical leads on distributed systems, performance engineering, context engineering, cost optimization, production readiness, AI platform design, and architecture trade-offs.
  • Represent architecture decisions through design documents, RFCs, diagrams, technical reviews, operational plans, and leadership-level summaries.
Desired Qualifications
  • Experience working on large-language-model inference, multimodal inference, agent infrastructure, AI assistants, coding agents, or frontier-model serving platforms.
  • Experience with tensor parallelism, pipeline parallelism, model sharding, KV-cache optimization, batching, speculative decoding, streaming inference, and long-context serving.
  • Experience designing context engineering platforms, prompt/version management systems, model-routing frameworks, semantic caching layers, token-budgeting systems, or large-language-model cost dashboards.
  • Experience profiling GPU workloads using Nsight Systems, Nsight Compute, rocprof, perf, Prometheus, Grafana, OpenTelemetry, or custom profiling systems.
  • Experience with large-scale distributed training, reinforcement-learning infrastructure, checkpointing, machine-learning compiler optimizations, model graph transformations, or training runtime systems.
  • Experience designing release gates, regression detection systems, canary systems, continuous-integration/continuous-delivery validation frameworks, and production safety controls for performance-sensitive infrastructure.
  • Experience with evaluations, model quality measurement, hallucination detection, grounding evaluation, safety testing, and model behavior monitoring.

Accellor guides innovative companies through enterprise-wide transformation to stay relevant and grow quickly. It offers tailored consulting and hands-on services that improve how a business works for customers, employees, and partners, from operations to strategy. It begins by listening to questions, evaluating the current environment, and then selecting a right mix of solutions rather than a cookie-cutter approach. By combining mobile, machine learning, and cloud platforms, Accellor delivers faster results across the whole organization, with about 200 staff across five countries.

Company Size

201-500

Company Stage

N/A

Total Funding

N/A

Headquarters

Fremont, California

Founded

2009

Get referred to Accellor

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • accellormedical.com was active through May 2026, signaling ongoing commercial operations.
  • OMS said Accellor’s customer base supports expansion beyond Ontario into Canada.
  • AliMed products and 70,000 SKUs deepen cross-sell opportunities with existing healthcare accounts.

What critics are saying

  • OMS announced Accellor’s acquisition in December 2023; independence already ended.
  • Once OMS integrates inventory into Ottawa, Accellor’s brand value disappears by 2026.
  • Large distributors like Cardinal Health and McKesson crush margins and lock hospital contracts.

What makes Accellor unique

  • Accellor Medical markets 70,000 products, a broad catalog for Ontario acute-care buyers.
  • Its 2024 AliMed partnership widened access to bracing, alarms, and patient-handling products.
  • The company built reputation around sourcing and customer service, not proprietary devices.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Parental Leave

Paid Vacation

Paid Holidays

Flexible Work Hours

Remote Work Options

Performance Bonus

Employee Referral Bonus

Professional Development Budget

Training Programs