Simplify Logo
d-Matrix

d-Matrix

Delivers memory-integrated AI compute platforms

Machine Learning Research Intern

Summer 2026, Fall 2026Posted on 1/27/2026
$30 - $59/hr
Internship
Master's, PhD
Santa Clara, CA, USA
Hybrid

Three days on-site per week in Santa Clara, CA.

About the job

Requirements
  • Pursuing Masters/PhD degree in Computer Science, Electrical and Computer Engineering, or a related scientific discipline
  • High proficiency with PyTorch is a must
  • High proficiency in algorithm analysis, data structure, and Python programming is a must
  • Current knowledge in machine learning and modern deep learning
  • Hands-on experience with modern neural network architectures such as MoEs and Diffusion models
Responsibilities
  • Design, implement and evaluate efficient deep neural network architectures and algorithms for d-Matrix's AI compute engine
  • Engage and collaborate with internal and external ML researchers to meet R&D goals
  • Engage and collaborate with Software team to meet stack development milestones
  • Conduct research to guide hardware design
  • Develop and maintain tools for high-level simulation and research
  • Port customer workloads, optimize them for deployment, generate reference implementations and evaluate performance
  • Report and present progress timely and effectively
  • Contribute to publications of papers and intellectual property
Desired Qualifications
  • Knowledge and experience with efficient deep learning is preferred: quantization, sparsity, distillation
  • Strong publication records in top machine learning conferences or journals is preferred
  • Proficiency with C/C++ programming is preferred
  • Proficiency with GPU CUDA programming is preferred
  • Experience with AutoML and meta learning is preferred
  • Experience with numerical analysis preferred
  • Experience with specialized HW accelerator systems for deep neural network is preferred

About the company

d-Matrix provides scalable, modular AI compute hardware and software for large datacenters, prioritizing energy efficiency and reduced data movement. Its core DIMC engine embeds compute directly into programmable memory, while a fabric of low-power chiplets delivers configurable compute resources and the accompanying software optimizes performance. This combination cuts data transfers and power use, aligning hardware design with memory-based computation for AI inference. The goal is to let large datacenters run AI workloads more efficiently at scale with customizable, modular compute platforms.

Company Size

201-500

Company Stage

Series C

Total Funding

$429M

Headquarters

Santa Clara, California

Founded

2019

Get referred to d-Matrix

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • d-Matrix raised $275 million in November 2025, valuing it at $2 billion.
  • Headcount reached 305 by March 2026, and hiring stayed active across 70 roles.
  • NVIDIA NVLink Fusion support in September 2026 broadens distribution through MGX ecosystems.

What critics are saying

  • Raptor's custom DRAM supplier remains undisclosed, threatening 2027 volume and pricing.
  • Raptor ships in 2027; any tape-out or yield slip hands Nvidia more time.
  • A 32GB Raptor card faces 192GB HBM4 rivals, forcing costly rack-scale deployments.

What makes d-Matrix unique

  • d-Matrix's Raptor stacks logic directly on custom DRAM, targeting 100 TB/s per card.
  • Parasail deployed Corsair with NVIDIA Hopper and Blackwell for heterogeneous inference in July 2026.
  • Wallaroo.ai and GigaIO acquisitions give d-Matrix silicon, software, and rack-scale integration.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Hybrid Work Options

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

-2%
Crypto Briefing
Sep 10th, 2026
D-Matrix targets Nvidia MGX rack integration for Raptor XPU by Q4 2027.

D-Matrix targets Nvidia MGX rack integration for Raptor XPU by Q4 2027. The AI inference startup plans to slot its custom accelerators into Nvidia's rack infrastructure, promising a nearly 5x throughput advantage over conventional HBM designs 2 hours ago Sponsored: CryptoSlots - Cryptoslots Play now! AI inference chip startup d-Matrix wants its upcoming Raptor XPU to live inside Nvidia's MGX rack architecture by the end of 2027, and the company has a fairly concrete roadmap to get there. The Santa Clara-based firm expects its first Raptor tape-out to land before the close of 2026, with full rack-scale deployment targeting Q4 2027. D-Matrix is integrating Raptor with Nvidia's NVLink Fusion interconnect, which means up to 144 Raptor XPUs could operate inside a single NVLink fabric by the time 2027 wraps up. What Raptor is actually built to do. Raptor debuted publicly at Hot Chips 2026 in August, and the technical specs are worth paying attention to. The platform combines a TSMC 4nm logic die with 3D-stacked DRAM, targeting over 100 TB/s of memory bandwidth with 32 GB of capacity per card. D-Matrix claims Raptor delivers approximately 4.7 times higher throughput per card compared to HBM-based alternatives for generative inference tasks. The Corsair foundation and the funding behind it. Raptor is d-Matrix's next act, but the company is not a pre-revenue concept. Its current platform, Corsair, is already in production and uses digital in-memory compute built on SRAM. The company has raised between $450 million and $500 million in total funding, including a $275 million Series C round that valued the company at roughly $2 billion. Partners in the broader ecosystem include Alchip and Andes, and the Raptor integration specifically involves Astera Labs, a connectivity chipmaker that sits inside Nvidia's technology ecosystem. Disclosure: This article was edited by Editorial Team. For more information on how Crypto Briefing create and review content, see its Editorial Policy.

The Register
Sep 10th, 2026
d-Matrix adopts Nvidia's NVLink Fusion for AI inference accelerators with 7.2PB/s bandwidth

AI infrastructure startup d-Matrix has licensed Nvidia's NVLink Fusion interconnect technology and MGX rack-scale reference designs for its upcoming Raptor inference accelerators. The integration allows d-Matrix to scale its chip architecture across large compute clusters using Nvidia's existing infrastructure. By end of 2026, d-Matrix expects to offer systems with up to 144 Raptor accelerators connected via NVLink fabric. Each Raptor card features 32GB of 3D-stacked DRAM delivering 100TB/s memory bandwidth — roughly 4.5 times that of Nvidia's Rubin GPU. The company joins MediaTek, Marvell, Qualcomm, Arm, Fujitsu, and Amazon Web Services in adopting NVLink Fusion. Beyond licensing revenues, Nvidia benefits as d-Matrix will also deploy its Vera CPUs, NVSwitch appliances, BlueField NICs, and SpectrumX Ethernet products.

HPCwire
Sep 10th, 2026
d-Matrix and NVIDIA Plan NVLink Fusion Rack System for AI Inference.

d-Matrix and NVIDIA Plan NVLink Fusion Rack System for AI Inference. September 10, 2026 Press play to listen to this content SANTA CLARA, Calif., Sept. 10, 2026 - d-Matrix today announced a new collaboration with NVIDIA that includes a multi-year product roadmap, providing d-Matrix XPUs entry into the widely deployed NVIDIA AI factory ecosystem. The centerpiece of the collaboration is an NVLink Fusion enabled rack-level system that AI labs, hyperscalers, and neoclouds can seamlessly deploy for ultra-low latency premium-level token services. As an NVIDIA NVLink Fusion partner, d-Matrix will work closely with NVIDIA to incorporate d-Matrix's next-gen inference XPUs, starting with Raptor, directly into the latest NVIDIA rack reference architecture design featuring NVIDIA Vera CPUs, NVIDIA NVLink switches, NVIDIA BlueField-4 DPUs, NVIDIA ConnectX-9 SuperNICs, and NVIDIA Spectrum-X Ethernet networking. d-Matrix is also partnering with Astera Labs, a connectivity solution leader within the NVIDIA NVLink Fusion ecosystem, to deliver custom solutions to ensure high-throughput, seamless data flow throughout the system. The d-Matrix rack, enabled by the MGX platform, will feature modular cable-free trays built with NVIDIA's mature, proven MGX ecosystem and supply chain for fast and seamless deployments. NVLink Fusion gives d-Matrix a mature, high-bandwidth, low-latency scale-up foundation for connecting d-Matrix XPUs to NVIDIA rack-scale infrastructure. With NVLink Fusion and the MGX ecosystem, d-Matrix can build around the same rack architecture, networking and supply-chain used across the NVIDIA platform. The first engagement point in the collaboration will be d-Matrix Raptor XPUs plugging into the NVIDIA MGX rack resulting in higher performance, greater deployment flexibility for customers, and a scalable architecture for expanding Raptor-based inference clusters as demand grows. "This collaboration with NVIDIA is a defining moment on our journey to infinite inference, accessible to all," said Sid Sheth, founder and CEO at d-Matrix. "Being integrated into NVIDIA's latest MGX rack-scale infrastructure with NVLink Fusion means our customers can deploy our inference XPUs alongside the broadly available NVIDIA AI factory platform. That's the future d-Matrix has been building toward - ultra-low latency, energy-efficient inference XPUs and GPUs working together, at rack scale, to deliver premium AI experiences." "NVLink Fusion enables partners to integrate custom silicon with NVIDIA's deep ecosystem of NVLink, advanced packaging, rack-scale systems and networking technologies," said Jensen Huang, founder and CEO of NVIDIA. "With NVIDIA AI infrastructure deployed across cloud and on-premises data centers worldwide, NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms - expanding accelerator choice for customers building the next generation of AI factories." "Purpose-built connectivity is what turns innovative compute into high-performing AI factories," said Jitendra Mohan, CEO of Astera Labs. "Our partnership with d-Matrix and NVIDIA brings this vision to life within the NVLink Fusion ecosystem, delivering high-throughput for low latency AI inference." The Rise of the Premium Token Economy As agentic AI workloads have spiked inference demand, AI service providers are increasingly seeking mixed-architecture systems to deliver inference economics their customers require. The d-Matrix MGX rack system is designed for latency-sensitive applications such as AI coding assistants, real-time chatbots, and voice agents where interactivity is paramount and customers are willing to pay a premium for speed. Built on NVIDIA's mature, proven MGX ecosystem and supply chain, the system extends a unified rack architecture that gives AI factories the flexibility to deploy the right compute for each workload. Using heterogeneous disaggregation, operators can split the workload between d-Matrix Raptor XPUs and NVIDIA Vera Rubin, allowing them to optimize each phase of inference. For one of today's most popular disaggregated applications, AI coding, GPUs can handle the compute-intensive prefill phase of a workload while d-Matrix inference XPUs speed up the latency-sensitive decode phase. Raptor: d-Matrix's Next-Gen Inference XPU A follow-on to the d-Matrix Corsair XPU platform currently in production, the d-Matrix Raptor platform extends d-Matrix's memory-centric architecture. Through a first-of-its-kind 3D DRAM stacking approach, Raptor brings a DRAM memory chip and an SRAM compute chip together to form a single "two-story" package. Technical details about the 3D DRAM technology were recently published by IEEE and previewed by d-Matrix co-founder and CTO Sudeep Bhoja at the 2026 Hot Chips conference. d-Matrix has designed Raptor, which is expected to tape-out before the end of the year, from the ground up for integration with NVIDIA NVLink Fusion and the NVIDIA MGX rack-scale ecosystem, reflecting d-Matrix's commitment to building purpose-built inference silicon that works seamlessly alongside the NVIDIA stack. Raptor is being actively evaluated at AI hyperscalers and frontier labs for its unique memory stacking solution and is backed by more than 100 patents. Availability Initial availability of d-Matrix Raptor XPUs integrated into the NVIDIA MGX rack is expected Q4 2027. To learn more, visit the d-Matrix booth at AI Infra Summit to see a demo or contact d-Matrix at www.d-matrix.ai/contact-sales. About d-Matrix d-Matrix is a leader in AI inference compute, delivering memory-centric compute chips, software and rack-scale systems for datacenters, making AI faster, more energy efficient, and more accessible. By bringing memory directly into compute, d-Matrix overcomes the cost, latency, and power limitations of traditional architectures - working independently or in partnership with GPUs and other compute platforms. Its multi-generation roadmap, from planar to 3D memory-compute substrates and rack-scale systems, is built to help hyperscalers, frontier labs, and AI clouds deploy advanced AI experiences at global scale. As demand for inference accelerates, d-Matrix is charting the path to infinite inference with no latency. For more information, visit www.d-matrix.ai. OpenAI says one of its AI systems has solved a math problem that has stumped... After more than two years in pilot mode, the National Science Foundation this week announced... Greg Kurtzer is biased. As the creator of CentOS, Rocky Linux, Apptainer (Singularity), and Warewulf,... The recently launched Genesis Mission includes a wide range of projects aimed at using advanced... Deep Origin this month announced that its AI drug discovery framework delivered nearly a 31%... AI models are getting better at a rapid pace. They are now able to reason,...

Tom's Hardware
Aug 26th, 2026
Hot Chips 2026: d-Matrix stacks AI accelerator directly on custom DRAM for 100 TB/s per card - TSMC 4nm compute die bonded face-to-face at a 36-micron pitch on top of a custom-designed die.

Hot Chips 2026: d-Matrix stacks AI accelerator directly on custom DRAM for 100 TB/s per card - TSMC 4nm compute die bonded face-to-face at a 36-micron pitch on top of a custom-designed die. Published 3 hours ago Heading into production during the tightest DRAM market in more than a decade. This Tom's Hardware Premium article is free to read with a Tom's Hardware account; no payment necessary. We're offering free access from August 23 to 26 so you can read all of our reporting from Hot Chips. d-Matrix presented Raptor, which it calls the first 3D DRAM accelerator for generative inference, at Hot Chips 2026 this week, showing a TSMC 4nm compute die bonded face-to-face at a 36-micron pitch on top of a custom-designed DRAM die that delivers 100 TB/s of bandwidth from 32GB per card. Co-founder and CTO Sudeep Bhoja put the vertical interface's energy cost at 0.37 pJ/bit against roughly 2.4 pJ/bit for moving data into an HBM4 base die, calling it "a measured number" from working silicon, and the accompanying ISCA 2026 paper, written with the University of British Columbia, projects around 4.7 times higher throughput per card than HBM-based designs. Bhoja, however, didn't disclose who manufactures the DRAM die. Latest Videos FromTom's Hardware How to check SSD health in Windows 10 and Windows 11 0 seconds of 1 minute, 25 seconds Volume 0% CEO Sid Sheth told CNBC in June that Raptor is slated to launch in 2027, but at Hot Chips the company gave no firm date or information on volume and pricing, and every performance figure shown, including 988 tokens per second per user on the 2.8-trillion-parameter Kimi K3 model at 1M-token context, is a d-Matrix projection built on early silicon. The custom DRAM die. Raptor inverts the usual 3D stacking arrangement by putting the logic die on top and the DRAM underneath, so a cold plate sits directly on the compute silicon and the DRAM die doubles as the interposer, carrying PCIe and die-to-die signals down through its TSVs. "For the same amount of power, you can drive the bandwidth up," Bhoja said during the session, "and so we were able to drive the bandwidth up here to 100 terabytes per second." The one-high stack achieves a power density of roughly 0.5W per square millimeter, which liquid cooling can handle, but the DRAM is designed for a junction temperature of 105°C, where retention collapses from a standard 32ms to 4ms, and the memory must refresh eight times more often. d-Matrix absorbed that penalty by shrinking each microbank to 1,366 rows and about 5.33MB, so a full refresh sweep costs only 1.37% of overall bandwidth. The die carries 840 banks per chiplet, of which 72 (around 9%) are spares wired into a two-level mux chain the company calls bank chaining, letting any two failed banks anywhere on the die be switched out while channels stay symmetric. A [132,128] Reed-Solomon code on the logic die corrects two symbol errors per 128 bytes, with a CRC behind it. The interface has no PHY, no burst structure, and no sideband pins, so conventional data-bus inversion was impossible; d-Matrix instead compares each 128-byte flit to the previous one and stores a 1-bit inversion tag alongside the ECC metadata, recovering roughly 20% of the I/O power DBI would have saved. At full tilt, the vertical interface still burns 296W of the 422W per-package budget the ISCA paper discloses. The bank geometry, spare-bank mux tree, refresh behavior, and interleaved ECC columns were all co-designed with the compute die. The 0.37 pJ/bit figure exists only because of that pairing, and no memory maker has anything like this die in its catalog. Who's supplying it? d-Matrix has named TSMC for the N4P logic die and Alchip as its ASIC design and 2.5D/3D packaging partner, but neither company operates a DRAM fab, and across the Hot Chips talk, the ISCA paper, and every public announcement since the Pavehawk 3DIMC test silicon came online last September, the firm has never identified who fabricates its custom DRAM. Only three companies make leading-edge DRAM at volume, and all three are allocating capacity to HBM4 lines that are effectively sold out through 2026. J.P. Morgan estimates DRAM prices will have risen more than 400% between the start of 2024 and the end of 2026. In addition, analysts have recorded contract price increases of 90% to 95% in Q1 2026 alone, and SK hynix CEO Kwak Noh-jung told Reuters in July that "customer demand will remain higher than our supply capacity even beyond 2030." Nvidia, the memory makers' largest and most leveraged customer, is reportedly testing Rubin Ultra configurations with as little as 192GB because it may not be able to source enough HBM4E. A startup with roughly $450 million raised, asking a memory maker to run a bespoke die with non-standard bank geometry on capacity that could otherwise print HBM, is negotiating from a far weaker position than that, and until the supplier is named, Raptor's 2027 volume plan rests entirely on an unknown, undisclosed dependency. The die's 11.4 MB/mm[2] density is roughly half of HBM4's 21.9 to 26.3 MB/mm[2], Bhoja acknowledged during the Q&A: "A lot of the drop for us was also because we used a not-so-advanced DRAM. And so if we used a more mainline DRAM, just like the HBM4 guys are doing, we would be able to push that up almost all the way to the HBM4 numbers." The density penalty therefore tracks back to whatever foundry arrangement d-Matrix currently has. 32GB per card. Raptor's 32GB per card stands against 192GB to 288GB for HBM4-equipped accelerators, so d-Matrix sizes deployments at rack scale instead: 72 cards carry 2.3TB, enough to hold Kimi K3's weights at 4-bit precision with headroom for around 54 concurrent users at 1M context, by its own calculations. "Even with 32 gigabytes of memory capacity, we are able to solve SOTA models in a scale-up network, so no bits are wasted," Bhoja said. KV cache growth works against that arithmetic over time, and co-presenter Aayush Ankit, who led Raptor's SoC architecture at d-Matrix before joining Meta's MTIA team, explained the failure mode while dismissing SRAM alternatives: "We are making this unit of compute blazingly fast. Communication becomes a bottleneck soon enough." Once models and context no longer fit a single rack, inference spills into the inter-card synchronization overhead that the vertical bandwidth was meant to eliminate, which the ISCA paper concedes. Cerebras claimed 969 tokens per second on Llama 3.1 405B in November 2024, with third-party firm Artificial Analysis verifying the figure on live hardware, and that remains the closest published reference point for Raptor's numbers. d-Matrix's 988 tokens per second per user on a model seven times larger would be a step change if it holds water, but no third party has measured Raptor, and the comparison points in the ISCA paper are simulations anchored to early silicon characterization. Asked by an Nvidia employee about multi-layer stacking plans, Bhoja said: "Our roadmap is still a work in progress, and we have a hard enough time trying to get one-high to work and work around all of the thermal issues of that." Samsung brought its own version of the idea to Hot Chips with zHBM, a concept that stacks HBM directly on the processor and carries a claimed 70% power-efficiency gain over an HBM4E setup, with no production timeline before HBM5. The problem for d-Matrix is that Samsung runs its own DRAM fabs; d-Matrix doesn't. Whether it can get wafers at volume remains to be seen. Full d-Matrix Hot Chips 2026 presentation. Image 1 of 30 Gain access to Tom's Hardware Premium's Hot Chips 2026 coverage between 23-26 Aug with a free account * | Get deeper news analysis * | Access detailed hardware roadmaps * | Explore performance data with our exclusive Bench tool Stay signed in for an improved experience across Tom's Hardware Contributor Luke James is a freelance writer and journalist. Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory.

Daily AI Brief
Aug 25th, 2026
d-Matrix raises $275M for AI inference chips at $2B valuation

d-Matrix, a Santa Clara-based AI chip company, raised $275 million in Series C funding at a $2 billion valuation, bringing total funding to $450 million. The round was led by Bullhound Capital, Triatomic Capital, and Temasek, with new investors including Qatar Investment Authority and EDBI. The company will use the funds to advance its product roadmap, expand globally, and support deployments of its data centre inference platform. d-Matrix's products include Corsair inference accelerators, JetStream networking accelerators, and Aviator software. The platform claims to produce up to 30,000 tokens per second at 2 milliseconds per token on a Llama 70B model and can run models with up to 100 billion parameters in a single rack.

INACTIVE