P

Persimmons.ai

Generative AI inference accelerator for edge

Compiler Engineer - Mid, Backend

Full-Time
No salary listed
Senior
San Jose, CA, USA
In Person
Company Does Not Provide H1B Sponsorship

About the job

Requirements
  • Solid understanding and experience with the underlying principles and methods of the MLIR framework, including static single assignment representation, interfaces, rewriting, and dialect hierarchy.
  • Hands-on experience developing MLIR-based compiler infrastructure, algorithms, and techniques for non-GPU or custom spatial hardware architectures.
  • Working experience lowering SIMD operations from PyTorch, Triton, xDSL, pyDSL, or similar Python-based frontends toward LLVM intermediate representation and then to a SIMD kernel library.
  • Extensive experience and understanding of loop optimization based on polyhedral principles.
  • Experience and understanding of SPMD-based distributed collective operations, specialized MLIR compiler dialects such as MESH and SHARDY, and collective-operation lowering in compilers for spatial hardware.
  • Experience with padding, bufferization, inlining, and other lowering techniques.
  • Knowledge of register allocation and instruction scheduling in spatial architectures.
  • Experience lowering and integrating custom operations and kernels at the compiler mid- and backend.
  • Familiarity with graph and tensor partitioning, mapping optimization algorithms, and their integration into compiler workflows.
  • At least five years of experience with C++ and an understanding of clean, maintainable code.
  • Demonstrated fluency with modern artificial intelligence tools and workflows, including using AI assistants for research, analysis, or productivity.
Responsibilities
  • Develop and enhance MLIR-based compiler pipelines targeting Persimmons' custom spatial accelerator hardware.
  • Design and optimize the Persimmons Compiler mid- and backend techniques for efficient lowering, graph-to-resources mapping, and code generation.
  • Implement transformations that convert Python, PyTorch, and similar kernel representations to LLVM intermediate representation and runtime-ready libraries.
  • Architect and implement support for SPMD-based distributed collective operations and lower them through specialized MLIR compiler dialects such as MESH and SHARDY.
  • Drive loop optimizations using polyhedral analysis, including loop tiling, fusion, interchange, and skewing.
  • Apply and optimize bufferization, padding, inlining, and integration of custom operations and kernels within the compilation workflow.
  • Work on register allocation and instruction scheduling for Persimmons' spatial hardware to ensure high resource utilization, throughput, and low latency.
  • Contribute to graph and tensor partitioning logic for optimal hardware-targeted execution.
  • Collaborate across hardware, systems, and software teams to deliver performant compilation flows from high-level machine-learning representations to low-level executable artifacts.
Desired Qualifications
  • Good knowledge of Python.

About the company

Persimmons.ai designs and builds a flexible, low-power generative AI inference system as a fabless semiconductor company. Its main product is a generative AI accelerator and accompanying software stack for AI inference, targeting high compute density and scalable performance across both edge devices and data centers. The system combines hardware (chiplet-based accelerator architecture), algorithms, and compiler technology to optimize generative AI workloads efficiently. Unlike traditional single-die accelerators, Persimmons emphasizes a scalable, full-stack approach that integrates hardware, software, and compilers to maximize inference throughput per watt from edge to high-performance computing (HPC) environments. The company aims to deliver a practical solution that enables dense, energy-efficient AI inference at scale for diverse deployment locations and workloads.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Jose, California

Founded

2023

Get referred to Persimmons.ai

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • August 2026 hiring for compiler and silicon roles signals active buildout and capital access.
  • SkyeChip’s November 2025 IP licensing deal validates Persimmons’ chiplet platform and partnerability.
  • Leadership pages list veterans from Nvidia, Intel, Microsoft, Amazon, and Apple, strengthening credibility.

What critics are saying

  • No verified funding or customer disclosures surfaced by August 2026, signaling weak commercial traction.
  • Public job listings for compiler and silicon engineers imply heavy pre-tapeout execution risk in 2026.
  • A failed first silicon or missed software stack would leave Persimmons.ai boxed out by Nvidia.

What makes Persimmons.ai unique

  • Persimmons.ai’s chiplet-based inference stack targets edge devices and data centers with one architecture.
  • June 2026 job posts show deep compiler, RTL, and physical-design integration under one roof.
  • The company markets full-stack control over hardware, compilers, and model optimization for inference.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Unlimited Paid Time Off

401(k) Retirement Plan

401(k) Company Match

Growth & Insights

Headcount

6 month growth

↑ 4%

1 year growth

↑ 0%

2 year growth

↑ 10%