Full-Time
AI accelerator hardware replacing GPUs
No salary listed
Toronto, ON, Canada + 1 more
More locations: Sunnyvale, CA, USA
Hybrid
Three days on-site per week required.
See people who can refer or advise you
Cerebras Systems creates AI acceleration hardware and software. Its CS-2 system is designed to replace traditional GPU clusters for AI workloads, speeding up training and inference while simplifying the setup by eliminating the need for parallel programming, distributed training, and cluster management. The product works as a single, large processor-based accelerator with accompanying software and cloud services to run AI models efficiently, reducing latency and time to results. Compared with competitors, Cerebras differentiates itself with the largest processor in the industry and an integrated hardware-software stack that aims to streamline AI workflows rather than relying on multi-GPU clusters. The company’s goal is to help research labs, healthcare, finance, and other industries achieve faster, more cost-effective AI development and deployment by offering a turnkey high-performance AI compute solution.
Company Size
501-1,000
Company Stage
IPO
Headquarters
Sunnyvale, California
Founded
2016
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Professional Development Budget
Flexible Work Hours
Remote Work Options
401(k) Company Match
401(k) Retirement Plan
Mental Health Support
Wellness Program
Paid Sick Leave
Paid Holidays
Paid Vacation
Parental Leave
Family Planning Benefits
Fertility Treatment Support
Adoption Assistance
Childcare Support
Elder Care Support
Pet Insurance
Bereavement Leave
Employee Discounts
Company Social Events
Cerebras Systems unveiled its CS-4 server rack for AI inference on 18 August in San Francisco. The Sunnyvale, California-based chipmaker said the system uses three WSE-3 Turbo chips and new networking components. CS-4 is the first product based on Cerebras' Nexus architecture, a modular design for compute, power, and input/output. The system features programmable input/output and direct wafer links to connect wafers within and across racks with lower latency. The chips are manufactured using TSMC's 5-nanometre process. The launch follows Cerebras reporting an adjusted loss of $6.9 million on sales of $180.1 million last week. The system will be available in the third quarter. Cerebras competes with Nvidia in inference hardware for workloads such as chatbot response generation.
Cerebras Systems has launched its CS-4 AI system, promising up to twice the token-generation speed of its predecessor, six times higher system-level performance, and up to 10 times more tokens per watt in certain applications. The company announced partnerships with OpenAI and AMD to enhance inference capabilities. The AMD collaboration features a split architecture where GPUs handle model prefill whilst Cerebras manages token decoding. Cerebras is expanding its data-centre capacity, with 600 megawatts of power expected online or under contract by the end of next year. The company is targeting AI-agent, design, coding, and cybersecurity applications that benefit from lower latency. CEO Andrew Feldman emphasised that inference speed has become a product-level consideration. Cerebras hardware powers OpenAI's GPT-5.6 Sol Ultrafast mode, which runs models at up to 14 times standard speed.
Cerebras Systems has launched the CS-4, its latest AI accelerator, delivering up to 30 times faster performance than GPU-based solutions. The rack-scale system incorporates three Wafer Scale Engine 3 Turbo processors and offers 750 petaflops of AI compute. The CS-4 provides up to twice the speed of its predecessor whilst delivering 10 times more throughput per watt. On GPT-OSS-120B, the system achieved over 4,400 tokens per second per user in testing. Built on the new Cerebras Nexus Platform Architecture, the CS-4 features a modular "backpack" design that reduces deployment time from days to hours. The system supports models exceeding 50 trillion parameters with wafer-to-wafer latency as low as two microseconds. First shipments begin this quarter. Cerebras Systems trades on NASDAQ under the ticker CBRS.
Cerebras Systems reported $180.1 million in second-quarter revenue, missing Wall Street's $193.6 million estimate. However, the AI chipmaker's cloud business showed strong momentum, with core cloud and services revenue surging 287% to $127.7 million as AI inference demand accelerates. Despite the revenue miss, management raised its full-year 2026 core revenue outlook and expects improving margins as the company transitions to more company-owned data centre capacity. The stock fell sharply following the earnings report. Cerebras, a Sunnyvale, California-based AI semiconductor company, develops specialised computing systems using its Wafer-Scale Engine technology. The company went public on Nasdaq on 14 May 2026 at $185 per share. Shares closed their first trading session at $311.07, marking a 68.2% gain from the IPO price.
Cerebras Systems' hardware is powering OpenAI's new Ultrafast mode in its GPT-5.6 Sol API, delivering inference speeds up to 14 times faster than standard processing. The integration uses Cerebras' wafer-scale architecture to reduce latency bottlenecks in AI workloads. The partnership expands Cerebras' role in OpenAI's generative AI infrastructure and places its technology more prominently in mainstream AI application pipelines. With a market capitalisation of approximately $52 billion, Cerebras ranks among the larger pure-play AI infrastructure specialists serving frontier model providers. The collaboration supports the thesis that customers will pay premium prices for fast inference in latency-sensitive applications like coding and document drafting. However, the tight integration also highlights customer concentration risk and execution pressure, particularly as Nvidia and AMD pursue similar inference speed improvements.