C

Cerebras

AI accelerator hardware replacing GPUs

Senior Research Engineer - Inference ML

Full-TimePosted on 10/3/2026
No salary listed
Mid, Senior
Bachelor's, Master's, PhD
Toronto, ON, Canada+1 moreMore locations: Sunnyvale, CA, USA
HybridHybrid role.

About the job

Requirements
  • One of the following education and experience combinations is required: a bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Electrical Engineering, or a related technical field and 7 or more years of machine-learning software development experience; a master's degree in Computer Science or a related technical field and 4 or more years of software development experience; a PhD in Computer Science or a related technical field and 2 or more years of relevant research or industry experience; or equivalent practical experience.
  • At least 4 years of experience testing, maintaining, or launching software products, including at least 2 years of experience with software design and architecture.
  • At least 3 years of software development experience focused on machine learning, such as deep learning, large language models, or computer vision.
  • Strong programming skills in Python and/or C++.
  • Experience with generative AI and machine-learning systems.
  • Evidence of machine-learning research impact, such as publications at NeurIPS, ICLR, ICML, ACL, EMNLP, or MLSys, or comparable contributions to widely used open-source projects or high-quality preprints.
  • Proficiency with at least one major machine-learning framework: PyTorch, Transformers, vLLM, or SGLang.
  • Deep understanding of transformer-based models in language and/or vision, with demonstrated experience implementing and optimizing them.
  • Ability to translate research ideas into robust code, including implementing new model variants, training strategies, and end-to-end evaluation workflows.
  • Strong foundation in performance optimization on specialized hardware such as GPUs, TPUs, or high-performance computing interconnects.
  • Deep understanding of modern machine-learning architectures and intuition for optimizing their performance, particularly for inference workloads using sparse attention, pruning and compression, and speculative decoding.
  • Track record of owning problems end-to-end and acquiring knowledge needed to deliver results.
  • Self-directed approach with the ability to identify and tackle impactful problems.
Responsibilities
  • Design, implement, and optimize state-of-the-art transformer architectures for natural language processing and computer vision on Cerebras hardware.
  • Research and prototype inference algorithms and model architectures that use Cerebras hardware capabilities, emphasizing speculative decoding, pruning and compression, sparse attention, and sparsity.
  • Train models to convergence, perform hyperparameter sweeps, and analyze results to inform next steps.
  • Bring up new models on the Cerebras system, validate functional correctness, and troubleshoot integration issues.
  • Profile and optimize model code using Cerebras tools to maximize throughput and minimize latency.
  • Develop diagnostic tools or scripts to surface performance bottlenecks and guide optimization strategies for inference workloads.
  • Collaborate across software, hardware, and product teams to drive projects from inception through delivery.
Desired Qualifications
  • Master's degree or PhD in Computer Science, Computer Engineering, or a related technical field.
  • Experience independently driving complex machine-learning or inference projects from prototype to production-quality implementations.
  • Hands-on experience with machine-learning frameworks such as PyTorch, Transformers, vLLM, or SGLang.
  • Experience with large language models, mixture-of-experts models, multimodal learning, or AI agents.
  • Experience with speculative decoding, neural network pruning and compression, sparse attention, quantization, sparsity, post-training techniques, and inference-focused evaluations.
  • Familiarity with large-scale model training and deployment, including performance and cost trade-offs in production systems.
  • Triton or CUDA experience.

About the company

Cerebras Systems creates AI acceleration hardware and software. Its CS-2 system is designed to replace traditional GPU clusters for AI workloads, speeding up training and inference while simplifying the setup by eliminating the need for parallel programming, distributed training, and cluster management. The product works as a single, large processor-based accelerator with accompanying software and cloud services to run AI models efficiently, reducing latency and time to results. Compared with competitors, Cerebras differentiates itself with the largest processor in the industry and an integrated hardware-software stack that aims to streamline AI workflows rather than relying on multi-GPU clusters. The company’s goal is to help research labs, healthcare, finance, and other industries achieve faster, more cost-effective AI development and deployment by offering a turnkey high-performance AI compute solution.

Company Size

1,001-5,000

Company Stage

IPO

Headquarters

Sunnyvale, California

Founded

2016

Get referred to Cerebras

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Q2 2026 core revenue hit $209.9 million, up 103% year over year.
  • Cerebras booked $25.4 billion in RPO on June 30, 2026.
  • Finland's September 1, 2026 Mikkeli campus adds 165MW contracted capacity.

What critics are saying

  • Data-center shortages force Cerebras to rent back customer systems, crushing near-term margins.
  • The August 2026 copyright class action in N.D. California targets model-training practices.
  • If power builds slip, Cerebras misses 2027 delivery and the backlog becomes stranded paper.

What makes Cerebras unique

  • Cerebras' wafer-scale chips still beat cluster-based AI systems for low-latency inference and training.
  • OpenAI signed a 750MW deal in December 2025, validating Cerebras' architecture.
  • AWS Bedrock integration and AMD disaggregation extend Cerebras beyond one-off hardware sales.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Professional Development Budget

Flexible Work Hours

Remote Work Options

401(k) Company Match

401(k) Retirement Plan

Mental Health Support

Wellness Program

Paid Sick Leave

Paid Holidays

Paid Vacation

Parental Leave

Family Planning Benefits

Fertility Treatment Support

Adoption Assistance

Childcare Support

Elder Care Support

Pet Insurance

Bereavement Leave

Employee Discounts

Company Social Events

Growth & Insights and Company News

Headcount

6 month growth

↓ -1%

1 year growth

↓ -2%

2 year growth

↑ 0%
Yahoo Finance
Sep 30th, 2026
Cerebras secures $25.4B in signed work as cloud revenue jumps 103% despite data center bottleneck

Cerebras Systems holds $25.4 billion in signed contracts yet to be delivered, according to its June 30, 2026 figures. The AI hardware and cloud computing company projects core revenue of $880 million to $890 million for 2026, with expectations to more than triple in 2027. Second-quarter 2026 core revenue reached $209.9 million, up 103% year-over-year, driven primarily by cloud services including OpenAI deployment. The stock trades at 51.1 times sales, compared to 3.1 times for the S&P 500. Data centre capacity remains the key bottleneck. Cerebras has secured over 600 megawatts of capacity through end-2027. The company temporarily rents back some customer-operated systems to meet demand, pressuring gross margins short-term. Management expects AWS integration on Bedrock platform in first quarter 2027, with OpenAI remaining a significant but declining revenue portion.

TechCrunch
Sep 30th, 2026
Cerebras CEO Andrew Feldman to discuss AI scaling limits at TechCrunch Disrupt 2026

Cerebras Systems CEO Andrew Feldman will speak at TechCrunch Disrupt 2026 about AI scaling challenges. The company, which raised $5.5 billion in its May IPO, takes a different approach to AI computing through wafer-scale processors rather than conventional chip architectures. Cerebras recently signed a multiyear agreement with OpenAI to deploy 750 megawatts of systems from 2026 through 2028. The company has more than 600 megawatts of data centre capacity live or under contract for delivery by end of 2027 and is increasing manufacturing capacity more than tenfold during 2026. Feldman will discuss compute demand, energy and infrastructure constraints, and potential limits of current AI hardware. The event takes place 13-15 October at Moscone West in San Francisco.

Yahoo Finance
Sep 29th, 2026
BigBear.ai vs Cerebras Systems: Which AI stock is the better 2026 buy?

BigBear.ai and Cerebras Systems represent two distinct approaches to the AI market. BigBear.ai provides decision intelligence solutions primarily for US federal and defence agencies, using AI to help navigate supply chains, manage autonomous systems, and enhance cybersecurity. In FY 2025, BigBear.ai generated roughly $127.7 million in revenue, down approximately 19.3% year-over-year. The company reported a net loss of about $293.9 million, yielding a negative 230.2% net margin. Free cash flow was negative $46.3 million. Around 51% of total sales came from customers contributing over 10% of revenue individually. Cerebras Systems builds wafer-scale chips designed for demanding AI workloads. The company serves sectors including medical research, energy, and agentic AI through hardware sales and cloud services. BigBear.ai has 579 employees, many holding high-level security clearances, and maintains a debt-to-equity ratio of 0.0x with a current ratio of approximately 1.8x.

Yahoo Finance
Sep 28th, 2026
Gimlet Labs partners with Cerebras to deliver AI inference speeds up to 3,000 tokens per second

Gimlet Labs and Cerebras Systems announced a collaboration to deliver ultrafast AI inference through Gimlet Cloud. The partnership combines Cerebras' wafer-scale compute with Gimlet's inference technology to achieve speeds of up to 3,000 tokens per second for real-time AI applications. Gimlet Cloud integrates the Cerebras Wafer Scale Engine with GPUs, using advanced disaggregation technology to optimise each phase of inference. The first Cerebras-powered Gimlet Cloud datacentre is expected to launch later this year. The collaboration aims to improve user experience in real-time AI applications such as voice assistants and AI agents, where reduced latency creates more responsive interactions. Gimlet Labs will serve as a launch partner for Cerebras' CS-4 technology, providing developers with production-scale access to the combined infrastructure.

Yahoo Finance
Sep 15th, 2026
Cerebras Systems turns profitable with 46.6% margin while CoreWeave posts $1.2B loss despite $5.1B revenue

Cerebras Systems and CoreWeave represent distinct approaches to AI infrastructure, offering investors different risk-reward profiles for 2026. Cerebras builds wafer-scale processors for intensive computing workloads, whilst CoreWeave provides specialised cloud services for generative AI. Cerebras reported revenue of $510 million in fiscal 2025, up 75.7% year-on-year, and achieved profitability with net income of $237.8 million. However, free cash flow was negative at $392.8 million. CoreWeave showed more explosive growth, with revenue reaching $5.1 billion, up 167.9%. The company remains unprofitable, posting a net loss of $1.2 billion. CoreWeave faces concentration risk, with Microsoft accounting for roughly 67% of revenue, though agreements with Meta Platforms suggest diversification efforts. Both companies serve the rapidly expanding AI computing market but differ significantly in scale, profitability, and business model maturity.