Full-Time

Director of ASIC Design Verification

Updated on 9/5/2026

SambaNova Systems

SambaNova Systems

201-500 employees

Enterprise AI platform with hardware stack

Compensation Overview

$270k - $340k/yr

+ Equity

San Jose, CA, USA

In Person

Bachelor's, Master's, PhD

Category
Engineering Management (1)
Required Skills
Claude
Verilog
Python
Software Testing
Machine Learning
Computer Networking

Get referred to SambaNova Systems

See people who can refer or advise you

Requirements
  • A bachelor's degree in Electrical Engineering, Computer Science, or Computer Engineering, with 12 or more years of industry experience in ASIC design verification, or equivalent industry experience.
  • At least 5 years of experience managing design verification engineers, including hiring and performance management, and experience leading senior and principal-level engineers.
  • A track record of inheriting an established verification environment and evolving it with the team rather than replacing it wholesale.
  • Deep, current technical expertise in verification, including ownership of methodology, arbitration of testbench architecture, and personal leadership of debug on the hardest pre-silicon failures.
  • Full-cycle design verification experience from planning through tape-out and into production silicon on at least one complex ASIC or large subsystem.
  • Hands-on expertise in SystemVerilog, SystemVerilog Assertions, Python, and industry verification methodologies such as Universal Verification Methodology, Open Verification Methodology, or Verification Methodology Manual.
  • Experience through silicon bring-up, including lab functional validation and debugging first-silicon issues to root cause.
  • Working knowledge of machine learning systems-on-chip or microprocessors, including central processing units or graphics processing units, and modern computer architecture.
  • Experience communicating effectively and building consensus among stakeholders and team members.
Responsibilities
  • Own pre-silicon functional and performance verification for SambaNova ASICs, including quality, coverage closure, and on-time delivery to tape-out.
  • Define verification architecture, strategy, schedules, and design verification environments for complex, high-performance machine learning ASICs.
  • Serve as the technical authority on verification by arbitrating testbench architecture, setting coverage and sign-off criteria, and leading debug on the hardest failures.
  • Set and defend the design verification schedule and own tape-out readiness sign-off, including trade-offs between coverage depth and time to silicon; report risk and readiness to executive staff with a clear point of view.
  • Support lab functional validation and bring-up debug for first silicon, feeding findings back into the pre-silicon environment.
  • Own the direction of the established verification environment and decide what to evolve and what to replace.
  • Lead, coach, and retain principal-level verification engineers; remain accountable for their development and continued engagement.
  • Grow the team toward a multi-workstream organization by owning staffing requirements, leveling, interview-loop design, senior-candidate closing, and development of technical leads.
  • Own the design verification methodology end to end, including verification architecture, reuse strategy, and sign-off criteria.
  • Lead development of test benches, agents, checkers, and stimulus at block, full-chip, system, and gate levels.
  • Drive pre-silicon debug and closure for functionality and performance; own regression health and root-cause discipline.
  • Implement new design verification methodologies and processes, including artificial-intelligence-assisted verification workflows, to improve execution effectiveness and efficiency.
  • Direct emulation and field-programmable gate array prototyping strategy, and own third-party verification intellectual property evaluation and electronic design automation vendor relationships for design verification tooling.
  • Partner with architecture, design, and software teams to settle requirements and drive design verification planning.
  • Work with compiler and runtime teams to map machine-learning application flows into verification so silicon is validated against real workloads.
  • Partner with hardware and systems teams on lab validation of first silicon, and represent design verification in reviews with physical design, process and quality, and product.
Desired Qualifications
  • A master's degree or PhD in EECS with 15 or more years of experience in ASIC design verification.
  • Experience leading multiple large-scale, high-impact chips to silicon.
  • Experience scaling a design verification team through a significant increase in headcount or chip complexity.
  • Experience owning or directing post-silicon functional validation in a lab environment.
  • Deep understanding of performance verification methodology.
  • Experience with high-performance interconnect protocols such as PCI Express and Ethernet, and memory subsystems.
  • Experience with advanced TSMC nodes and modern multi-die or chiplet verification challenges.
  • Experience with emulation and field-programmable gate array prototyping platforms.
  • Experience managing cross-site or distributed teams.
  • Fluency with Claude or other artificial intelligence tools applied to verification and flow development.

Generating company summary.

Company Size

201-500

Company Stage

Series F

Total Funding

$2.5B

Headquarters

Palo Alto, California

Founded

2017

Get referred to SambaNova Systems

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • July 8, 2026 Series F raised $1 billion at an $11 billion valuation.
  • July 22, 2026 Genesis Mission Consortium adds U.S. government credibility and sovereign-AI access.
  • August 2026 Hot Chips coverage highlighted SN50 scaling to 256 chips and 763 tokens per second.

What critics are saying

  • April 2026 layoffs cut 77 employees, signaling still-tight economics and execution pressure.
  • Nvidia’s Blackwell platform and Groq’s inference chips keep squeezing SambaNova’s price-performance claims.
  • If SN50 adoption stalls, SambaNova becomes another capital-intensive chip startup, not a durable platform.

What makes SambaNova Systems unique

  • February 2026 SN50 targets decode-heavy inference with dataflow architecture, not training clusters.
  • Intel’s February 2026 collaboration routes Xeon servers and channels into SambaNova deployments.
  • July 2026 JPMorganChase chose SambaNova for secure on-prem inference, validating regulated-enterprise demand.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible PTO in US

Parental Leave

Benefits (medical, dental and vision)

Flexible Spending Accounts

401k/Pensions

Gym Access

Flexible Working Hours

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

1%
ServeTheHome
Aug 25th, 2026
SambaNova's SN50 RDU for AI at Hot Chips 2026.

SambaNova's SN50 RDU for AI at Hot Chips 2026. August 25, 2026 For as new as the dedicated AI accelerator field is, SambaNova is one of the older and more established hardware vendors. The company is now in the fifth generation of their reconfigurable dataflow unit (RDU) technology, with the SN50 that was launched earlier this year. As with the other major AI vendors at this year's Hot Chips conference, the company has come to present new technical details on SN50, and outline what makes it competitive in the burgeoning field of dedicated AI accelerators. SambaNova's hardware has taken on an increased prominence in the industry thanks to the company's connection to Intel. While Intel itself is still trying to catch up on AI accelerators, the company has become increasingly attached to the hip to SambaNova, whose RDUs provide the dedicated, high-efficiency and low-latency AI accelerators that round out Intel's hardware stack. Thus the company's progress with the SN50 (and future RDUs) is material not just for SambaNova, but for Intel as well. Setting the stage, agentic inference is all the rage right now. Where does all the execution time go? SambaNova has a breakdown of it. Most time is spent in decode, especially on DeepSeek V3 where it's 97% of the time, versus 3% for prefill. And decode, in turn, is bandwidth-bound. The FLOPS-per-byte ratio is quite low, even for large batch sizes. Bandwidth utilization is often misunderstood. SambaNova is laying out what they mean for this talk. In short, they aren't talking about just how much HBM bandwidth is being used, but rather the Model Bandwidth Utilization (MBU) model. And specifically, what fraction of that is being used to cache data and otherwise handle data usage. Looking at the current state of tech, GPUs offer low bandwidth usage, even with highly optimized GPU-friendly benchmarks. Things get worse for GPUs when you scale up the number of them; performance does go up, but MBU drops significantly. Meanwhile frontier models require being able to scale up. Again with a GPU example, a GPU can get to around 30TB/second of model bandwidth. But they can't get past that. Enter SambaNova's SN50 dataflow RBU. They have doubled-down on what worked well from SN40, such as the large on-chip SRAM. 5x as many FLOPS as SN40, and it is designed to scale-up to a much larger domain of 256+ chips. And there is a separate scale-out network using 400Gb networking. Notably, there are no I/O dies or similar here. Instead it is just two max reticle dies for the logic, and then HBM stacks for the memory. Though it is interesting that SambaNova's choice of HBM here is quite dated; SN50 still uses HBM2e here (which is going to be a problem in the future as production of the memory is already ramping down). Moving up to the SN50 rack architecture, there are 16 RDUs in a single air-cooled rack, split over two nodes. Diving a bit deeper into SN50 and the dataflow architecture. The core element of the SN50 is the sea of compute cores (PCUs) and memory cores (PMUs). There is no hardware memory management; this is all software managed. Every unit operates when it has input and sends it to the outputs. To better illustrate how the dataflow architecture works, here is an example of how it maps to a transformer. Here is what a GPU looks like. And how it look on the SN50. For compute, data from the HBM is fed into the AGCU portal that control off-chip access, and from there into the PCUs and PMUs. The SRAM amount used is not a function of the size of the model. That was one RDU. How do things scale up for multiple RDUs? SambaNova employs both scale-up and scale-out networking. The Scale-up network is based on 800GbE, while scale-out is 400GbE. And then there is a front-end network. SambaNova uses an all-to-all topology for an 8 socket configuration. To go above 8 sockets, then the scale-up network is employed using Ethernet switches. Links are ganged, and every node is connected to each of two switches in this 64 chip socket configuration. Then things can be scaled out further, in this case employing both scale-up and out for a 512 socket configuration. The key to performance on SN50 is overlap. SN50 supports all forms of model parallelism, and the collective communication forms that these models are built on. Here is a brief look at performance with GEMM benchmarking. The utilization is consistently 70% or higher even at 32 sockets. If you are able to overlap, you can do the compute and communications in parallel. That kind of overlap is not something GPUs can do. The building block for SambaNova is collective communication, which is the purple boxes in these diagrams. And the SRAMs can stream from one to another without having to go through a higher layer (e.g. HBM). Here's a look at parallelism with tensor parallel. Here is a look at the bandwidth utilization that SN has achieved with DeepSeek. Meanwhile they can also use expert parallel (EP) as an additional form of parallelism. This relies on broadcast-dispatch as well as all-to-all dispatch-combine. With all-to-all, one way is to dynamically send everything to the target RDUs. Alternatively, you can just blast everything to all of the RDUs and then filter out things afterwards. The all-to-all method requires a group-by operation at the end of the router to collect (group) the tokens before transmitting them SRAM-to-SRAM. All-to-all also means allowing dynamic traffic. Now here's the other method of broadcast + filter. That is still an SRAM-to-SRAM operation, but with a filter operation on the PCUs of the receiving RDU. This keeps the network traffic parallel; though it does increase it a bit. And by not depending on the router, the transfer can be started early. Here is another DeepSeek example, with SambaNova getting close to 80% bandwidth utilization for loading the experts in MoE. As a result of this, SN50 achieves a high MBU value even at scale, with MBU holding at 45% even with 256 SN50s. And this makes it possible to keep adding RDUs to scale up things even further. This, in turn, means that models don't have to give up bandwidth. Going back to SambaNova's original chart about power scaling, here is what SN50 clusters of different sizes look like. A 512 RDU configuration is able to scale up to an aggregate model bandwidth capacity of over 350 TB/second. The systems can strongly scale, with MBUs still in the 40% range at 512 sockets. Ultimately SambaNova is promoting a very similar picture as other dedicated inference chip firms, using one type of chips for prefill (and midfill), while using separate accelerators (i.e. SN50) for decode. Specifically, they've been using NVIDIA H200 + SN50, with RoCE for transferring between them. Finally, taking a look at that performance in action, based on an Artificial Analysis benchmark of SN50. The hardware achieves over 750 tokens-per-second in MiniMax M2.7.

Tech Funding News
Aug 20th, 2026
Callosum raises $100M led by Atomico in one of Europe's largest seed rounds.

Callosum raises $100M led by Atomico in one of Europe's largest seed rounds. August 20, 2026 * Having raised $10.25 million in a pre-seed round in February, Callosum has now obtained $100 million in seed funding. * Atomico led the seed round, with Plural, DCVC, and the UK's Sovereign AI Fund also joining. * According to the fund's website, the investment in Callosum was its first, made in April. Callosum, a London-based startup, has raised a $100 million seed round, one of the largest ever in Europe, led by Atomico. This follows its pre-seed round at a $10.25 million valuation led by Plural in February. UK's £500 million Sovereign AI Fund, along with some other unnamed investors and angels, took part. Since its foundation in 2025, Callosum has now raised approximately $110 million in funding. The company has not shared its valuation. From neuroscience PhDs to a nine-figure seed. Callosum was founded by Danyal Akarca and Jascha Achterberg, who first began working together while pursuing their PhDs at Cambridge in neuroscience, computing, and AI. Their research has appeared in Nature journals, and both have held positions at Intel and Google DeepMind. Callosum's method is based on the founders' academic backgrounds. They hold the view that, just as intelligence in nature arises from a variety of neurons, AI should not rely on identical chips. Instead, their software breaks down AI tasks into separate steps and sends each one to the most suitable model and chip rather than using the same hardware for all purposes. As a growing share of industry spending shifts from training to inference, AI companies frequently spend half or more of their revenue on inference. For instance, one step in an autonomous computing task may require fast pattern matching, while another may call for deeper reasoning. The company's system therefore allocates each step to different hardware, thus preventing competition for resources. It claims that its solution is twice as accurate, seven times faster, and four times cheaper for complex tasks compared with using uniform hardware, even though it has not made the technical benchmark figures available. Nvidia's CUDA ecosystem, with its ~85% share of the GPU market, remains Callosum's principal competitor. Cerebras, the wafer-scale chipmaker that went public in May 2026 at a valuation of nearly $56 billion, has now entered into a partnership with Callosum. Although Etched and SambaNova are also developing their own custom chips to improve inference efficiency, their methods are less extensive than Callosum's chip-agnostic approach. A 10x increase in funding since February. The company's new funding will enable it to become a "global heterogeneous integrator" and establish new computing partnerships. Its key partnership with Cerebras focuses on achieving low-latency inference at scale, and it is also entering into a collaboration with the Korean chipmaker Rebellions. "By integrating Cerebras into Callosum's platform, we're making ultra-low-latency inference available exactly where it creates the greatest impact, enabling customers to build AI systems that simply weren't practical before," said Andrew Feldman, Cerebras' chief executive. Sunghyun Park, chief executive of Rebellions, framed the deal as a structural stand against single-vendor lock-in: "Working with Callosum puts our architecture into systems alongside hardware chosen for different parts of the workload, instead of asking one chip to do every job. That's the difference between a partnership and a deployment that only works in one environment or geography." Kanishka Narayan, the UK's minister for artificial intelligence, tied the investment to national chip strategy: "AI is nothing without the chips that underpin it, and the eye-watering demand for them is only going to grow. In the race to develop and use AI, success will depend not just on having access to those chips, but on using them as efficiently as possible." Nvidia believes the total AI infrastructure market could reach at least $1 trillion by 2027, driven primarily by inference workloads. Callosum maintains that no single chip can meet this demand. The UK government has again supported this view since April, even if it has not stated so outright.

Mart Infomedia
Aug 20th, 2026
WSO2 appoints Harry Ault as Chief Executive Officer to drive global growth strategy.

WSO2 appoints Harry Ault as Chief Executive Officer to drive global growth strategy. WSO2 has announced the appointment of Harry Ault as its new Chief Executive Officer, marking a significant step in the company's growth journey as it expands its presence in the rapidly evolving enterprise technology and AI governance landscape. Ault assumes day-to-day leadership of WSO2 with immediate effect and will work closely with the executive team to strengthen the company's global market position, deepen customer engagement and accelerate innovation across its product portfolio. He brings more than two decades of experience in revenue growth, strategic partnerships, corporate development and global go-to-market leadership. Before joining WSO2, Mr Ault served as Chief Revenue Officer at SambaNova Systems, where he led worldwide sales, marketing, field engineering and customer support operations. Throughout his career, he has played a key role in scaling technology businesses, building high-performance teams and driving growth across international markets. The appointment comes during a period of continued expansion for WSO2. Over the past year, the company has strengthened its leadership team with key appointments in finance and marketing as it advances its vision of trusted AI governance and enterprise transformation. WSO2 is focused on helping organisations manage increasingly complex digital environments by combining API management, integration, identity, engineering and AI-driven platforms within its Agentic Enterprise Fabric (AEF). The platform is designed to provide enterprises with greater control, governance and security as they adopt AI-powered technologies and autonomous agents. Jonas Persson, Chairman of the Board, said Ault's experience in building global teams and driving business growth makes him well suited to lead WSO2 through its next phase of expansion. Harry Ault said WSO2 has built a strong reputation over the past two decades, serving more than 700 enterprise customers worldwide. He added that organisations are increasingly seeking trusted partners that can help them adopt AI technologies without adding complexity, and WSO2 is well positioned to support that transition. Mr Ault succeeds WSO2 Founder Dr Sanjiva Weerawarana, who stepped down as CEO in May 2026. During the transition period, Chief Revenue Officer Devaka Randeniya served as Acting CEO and will continue to work closely with Mr Ault. Founded in 2005, WSO2 provides open-source technology platforms that support API management, integration, identity and AI-driven enterprise applications. The company operates across multiple global markets, with offices in North America, Europe, Asia, the Middle East and Australia, serving enterprises across a wide range of industries.

Mylstingo
Aug 8th, 2026
AMD buys Taalas, the startup etching AI models into silicon.

AMD buys Taalas, the startup etching AI models into silicon. AMD just bought a company whose flagship chip can run exactly one AI model. That is not a design flaw. It is the whole point. The chipmaker announced on August 6 that it has signed a definitive agreement to acquire Taalas, a Toronto startup founded in 2023 with a radical approach to AI hardware. Instead of building general-purpose processors that can load any model, Taalas physically etches a specific model's weights into the silicon itself. Financial terms were not disclosed. A chip that cannot change its mind. Taalas builds what it calls model-specific silicon, and its first product shows how literal that phrase is. The HC1 chip runs Meta's Llama 3.1 8B and nothing else, because the model's parameters are hardwired into the chip during manufacturing. There is no loading weights from external memory, no shuffling data between the processor and storage. The model is the chip. That design attacks the biggest bottleneck in AI inference. Modern GPUs spend an enormous share of their time and energy moving model weights between memory and compute units rather than doing the actual math. By baking the weights into the transistors, Taalas removes that traffic entirely. The chips are fabricated by TSMC on its 6-nanometre process node, a mature and relatively affordable technology compared with the cutting-edge nodes that flagship GPUs demand. The claimed results are striking. Taalas says its hardware can generate more tokens per second per user than Nvidia's H200 and B200 accelerators, and outpaces specialist inference hardware from Groq, SambaNova, and Cerebras, while consuming one tenth of the power. Those are the company's own figures, and independent benchmarks will be worth watching. But even a fraction of that efficiency gain would matter enormously in an industry where electricity has become the scarcest resource. Why AMD wants frozen models. The obvious objection is that AI models change constantly. A chip that runs one model forever sounds like a liability in a field where labs ship new versions every few months. AMD is betting the economics say otherwise. Inference, the work of actually running trained models for users, now dominates AI computing demand. Training happens once. Serving happens billions of times a day, and the workloads are surprisingly stable. A company running a popular model at scale might serve the same version for months. For those customers, a hyper-efficient chip dedicated to that exact model could cut serving costs dramatically, even if the hardware needs replacing when the model does. AMD said it plans to fold Taalas technology into its accelerator roadmap and develop system-level products alongside its Instinct GPUs, EPYC processors, Helios rack-scale platform, and ROCm software stack. The picture that emerges is a portfolio play: flexible GPUs for training and fast-moving workloads, hardwired silicon for the stable, high-volume serving that drains most of the power budget. The inference race tightens. The acquisition lands in the middle of an intensifying contest for the inference market. Nvidia still dominates AI hardware overall, but inference is where challengers see an opening, because the requirements differ from training. Groq built its business on deterministic, low-latency inference chips. Cerebras sells wafer-scale monsters. SambaNova pitches reconfigurable dataflow architecture. Each is chasing the same insight: the hardware that trained the model is not necessarily the best hardware to serve it. AMD has been assembling inference capabilities piece by piece, and Taalas gives it something none of its rivals currently ship, a commercial chip with the model burned in. For Nvidia, the deal is one more sign that the competition has stopped trying to beat its GPUs at their own game and started changing the rules instead. There is also a Canadian angle worth noting. Taalas emerged from Toronto's deep pool of semiconductor and machine learning talent, and its acquisition continues a pattern of US chip giants shopping north of the border for specialised engineering teams. What happens next. The open questions are practical ones. How quickly can AMD turn a startup's first product into something hyperscalers will deploy by the rack? Which models get the hardwired treatment, and who decides? And can the efficiency claims survive contact with independent testing? Answers should arrive over the coming year as AMD integrates the team and slots the technology into its roadmap. If the numbers hold up, the industry may need to rethink one of its base assumptions: that AI hardware has to be flexible. Sometimes the fastest chip is the one that only knows a single trick. For more coverage of AI hardware and the chip industry, visit Mylistingo. Ramo is the editorial voice of Mylistingo - an AI and technology news platform based in The Hague, Netherlands. Covering artificial intelligence, machine learning, robotics, and the future of technology, Ramo delivers accurate, accessible reporting for both general audiences and industry professionals. Every article is fact-checked and written to meet Mylistingo's strict no-fabrication editorial standards.

Yahoo Finance
Aug 4th, 2026
SambaNova appoints Mohsen Moazami as vice chair following $11B Series F valuation

SambaNova, a leading AI inference company, has appointed Mohsen Moazami as vice chair of global strategy and partnerships. Reporting to CEO Rodrigo Liang, Moazami will drive expansion across global enterprise, government, and sovereign AI markets. The appointment follows SambaNova's Series F financing at an $11 billion valuation. Moazami brings extensive experience from senior roles at Groq, where he served in the Office of the CEO and led international operations through a $20 billion strategic transaction. He also serves as vice chair of DDN, an AI data platform provider. At SambaNova, Moazami will leverage his experience in enterprise, government, and sovereign markets to scale the company's AI infrastructure offerings globally.