Full-Time

Staff Python / PyTorch Developer

Frontend Inference Compiler

Cerebras

Cerebras

501-1,000 employees

AI accelerator hardware replacing GPUs

No salary listed

Dubai - United Arab Emirates

In Person

Bachelor's, Master's, PhD

Category
Software Engineering (1)
Required Skills
Python
PyTorch
C/C++

Get referred to Cerebras

See people who can refer or advise you

Requirements
  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Mathematics, or a related field.
  • 8+ years of experience in large-scale software engineering, with a focus on deep learning or related domains.
  • Proficiency in Python for building and maintaining scalable systems.
  • Advanced proficiency in C++, with an emphasis on multi-threaded programming, performance optimization, and system-level development.
  • Demonstrated experience driving cross-functional projects.
  • Experience building and scaling large-scale inference systems for LLMs or multimodal models.
  • Familiarity with LLM serving frameworks, such as vLLM, SGLang, and TensorRT-LLM.
  • Solid understanding of software architectural patterns for large-scale, high-performance applications.
  • Hands-on experience with ML frameworks, such as PyTorch, and a strong understanding of their underlying architectures.
  • Strong problem-solving skills, with the ability to balance technical depth with practical implementation constraints.
  • Exceptional communication and presentation skills, with the ability to work both independently and collaboratively across multidisciplinary teams.
Responsibilities
  • Drive and provide technical guidance to a team of software engineers working on complex machine learning integration projects.
  • Design and implement ML features (e.g., structured outputs, biased sampling, predicted outputs) that improve performance of generative AI models at inference time.
  • Design and implement high-throughput, low-latency multimodal inference models that support delivery of image, audio, and video inputs and outputs.
  • Maintain our scalable serving backend for handling many concurrent requests per minute.
  • Scale our inference service by implementing detailed observability throughout the entire stack.
  • Analyze and improve latency, throughput, memory usage, and compute efficiency on the service and the implementation of various features.
  • Optimize software to accelerate generative LLM inference by achieving high throughput and low latency.
  • Stay up-to-date with advancements in machine learning and deep learning, and apply state-of-the-art techniques to enhance our solutions.
  • Evaluate trade-offs between different approaches, clearly articulate design choices, and develop detailed proposals for implementing new features.
  • Uncover, scope, and prioritize significant areas of technical debt across the software stack to ensure continued high quality of the inference service.
  • Build and maintain robust automated test suites to ensure software quality, performance, and reliability.
  • Contribute to an agile team environment by delivering high-quality software and adhering to agile development practices.
  • Lead cross-functional initiatives across the company to deliver high-quality inference solutions.

Cerebras Systems creates AI acceleration hardware and software. Its CS-2 system is designed to replace traditional GPU clusters for AI workloads, speeding up training and inference while simplifying the setup by eliminating the need for parallel programming, distributed training, and cluster management. The product works as a single, large processor-based accelerator with accompanying software and cloud services to run AI models efficiently, reducing latency and time to results. Compared with competitors, Cerebras differentiates itself with the largest processor in the industry and an integrated hardware-software stack that aims to streamline AI workflows rather than relying on multi-GPU clusters. The company’s goal is to help research labs, healthcare, finance, and other industries achieve faster, more cost-effective AI development and deployment by offering a turnkey high-performance AI compute solution.

Company Size

501-1,000

Company Stage

IPO

Headquarters

Sunnyvale, California

Founded

2016

Get referred to Cerebras

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Q2 2026 core revenue rose 103% to $209.9 million, driven by cloud expansion.
  • Core cloud and services revenue jumped 287% to $127.7 million, signaling real usage.
  • Cerebras secured over 600MW capacity through 2027 and raised 2026 core revenue guidance to $890 million.

What critics are saying

  • OpenAI controls Cerebras demand; the 750MW commitment creates brutal concentration through 2028.
  • Kaplan Fox and Schall opened 2026 securities investigations, inviting costly distraction and disclosure risk.
  • Nvidia Blackwell Ultra and AMD Helios attack Cerebras’ inference thesis before CS-4 scales.

What makes Cerebras unique

  • Cerebras’ wafer-scale WSE-3 Turbo packs memory bandwidth and latency unmatched by GPU clusters.
  • CS-4 launched August 18, 2026 with 750 PFLOPs and 4,400 tokens-per-second inference.
  • OpenAI’s GPT-5.6 Sol Ultrafast runs on Cerebras, proving frontier-model adoption in production.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Professional Development Budget

Flexible Work Hours

Remote Work Options

401(k) Company Match

401(k) Retirement Plan

Mental Health Support

Wellness Program

Paid Sick Leave

Paid Holidays

Paid Vacation

Parental Leave

Family Planning Benefits

Fertility Treatment Support

Adoption Assistance

Childcare Support

Elder Care Support

Pet Insurance

Bereavement Leave

Employee Discounts

Company Social Events

Growth & Insights and Company News

Headcount

6 month growth

-2%

1 year growth

-4%

2 year growth

0%
Yahoo Finance
Aug 22nd, 2026
Cerebras core revenue up 103% to $210M as cloud quadruples but GAAP hardware sales drop 23%

Cerebras Systems reported second-quarter core revenue of $209.9 million, up 103% year-on-year and exceeding guidance of approximately $194 million. However, GAAP revenue of $180.1 million fell short of the $194.23 million consensus. The quarter revealed a sharp revenue mix shift. Cloud and other services revenue roughly quadrupled to $126 million, whilst GAAP hardware revenue declined 23% to $54.1 million. Core hardware revenue rose 17% year-on-year to $82.1 million but dropped 26% sequentially from $111.6 million. Shares fell 11.9% to $231.01 following the report. The company raised its 2026 core revenue forecast to $880–890 million from $855–865 million and reported $25.4 billion in remaining performance obligations, supported by a multiyear OpenAI agreement.

Yahoo Finance
Aug 21st, 2026
Cerebras unveils CS-4 AI system with 750 PFLOPS, 30x faster inference than GPUs

Cerebras unveiled its fourth-generation CS-4 AI system, claiming inference speeds exceeding 4,400 tokens per second per user on GPT-OSS-120B — up to 30 times faster than GPU-based solutions. The CS-4 delivers 750 PFLOPS of AI compute and 7.2 Tbps of I/O, with 10 times more throughput per watt than the previous CS-3. The company reported second-quarter 2026 core cloud and services revenues of $127.7 million, up 287% year over year. Total core revenues rose 103% to $209.9 million. Cerebras disclosed $25.4 billion in remaining performance obligations and over 600 megawatts of data-centre capacity under contract for delivery by end of 2027. The company faces competition from NVIDIA, which reported $75.2 billion in first-quarter fiscal 2027 data-centre revenues, and AMD's Helios platform.

Tech in Asia
Aug 19th, 2026
Cerebras unveils CS-4 rack with three WSE-3 Turbo chips for AI inference

Cerebras Systems unveiled its CS-4 server rack for AI inference on 18 August in San Francisco. The Sunnyvale, California-based chipmaker said the system uses three WSE-3 Turbo chips and new networking components. CS-4 is the first product based on Cerebras' Nexus architecture, a modular design for compute, power, and input/output. The system features programmable input/output and direct wafer links to connect wafers within and across racks with lower latency. The chips are manufactured using TSMC's 5-nanometre process. The launch follows Cerebras reporting an adjusted loss of $6.9 million on sales of $180.1 million last week. The system will be available in the third quarter. Cerebras competes with Nvidia in inference hardware for workloads such as chatbot response generation.

Yahoo Finance
Aug 19th, 2026
Cerebras unveils CS-4 AI system with OpenAI and AMD partnerships for faster inference

Cerebras Systems has launched its CS-4 AI system, promising up to twice the token-generation speed of its predecessor, six times higher system-level performance, and up to 10 times more tokens per watt in certain applications. The company announced partnerships with OpenAI and AMD to enhance inference capabilities. The AMD collaboration features a split architecture where GPUs handle model prefill whilst Cerebras manages token decoding. Cerebras is expanding its data-centre capacity, with 600 megawatts of power expected online or under contract by the end of next year. The company is targeting AI-agent, design, coding, and cybersecurity applications that benefit from lower latency. CEO Andrew Feldman emphasised that inference speed has become a product-level consideration. Cerebras hardware powers OpenAI's GPT-5.6 Sol Ultrafast mode, which runs models at up to 14 times standard speed.

Associated Press
Aug 19th, 2026
Cerebras launches CS-4 AI accelerator, 30x faster than GPUs with 750 PFLOPs compute

Cerebras Systems has launched the CS-4, its latest AI accelerator, delivering up to 30 times faster performance than GPU-based solutions. The rack-scale system incorporates three Wafer Scale Engine 3 Turbo processors and offers 750 petaflops of AI compute. The CS-4 provides up to twice the speed of its predecessor whilst delivering 10 times more throughput per watt. On GPT-OSS-120B, the system achieved over 4,400 tokens per second per user in testing. Built on the new Cerebras Nexus Platform Architecture, the CS-4 features a modular "backpack" design that reduces deployment time from days to hours. The system supports models exceeding 50 trillion parameters with wafer-to-wafer latency as low as two microseconds. First shipments begin this quarter. Cerebras Systems trades on NASDAQ under the ticker CBRS.