Full-Time

AI/ML Compiler & Runtime Software Engineer

Posted on 9/11/2026

GlobalFoundries

GlobalFoundries

10,001+ employees

Semiconductor foundry delivering specialized chips

No salary listed

Pune, Maharashtra, India + 1 more

More locations: Bengaluru, Karnataka, India

In Person

Category
Software Engineering (1)
Required Skills
Python
PyTorch
C/C++
Linux/Unix

Get referred to GlobalFoundries

See people who can refer or advise you

Requirements
  • 3-12 years of hands-on software engineering experience, with strong experience in compiler, runtime, embedded software, or AI/ML systems.
  • Strong hands-on experience with IREE, LLVM, and MLIR compiler infrastructure.
  • Good understanding of IREE code generation flow, dispatch formation, executable generation, HAL/runtime concepts, and target-specific lowering.
  • Strong exposure to AI compiler/runtime stacks used for edge AI or accelerator-backed inference.
  • Experience with AI model formats and frameworks such as PyTorch, ONNX, TensorFlow Lite/TFLite, and related conversion or import flows.
  • Working knowledge of torch-mlir, TOSA, Linalg, tensor dialects, bufferization, quantization dialects, and MLIR-based model lowering concepts.
  • Strong understanding of neural network execution and optimization, including quantization, operator fusion, tensor layouts, memory planning, tiling, vectorization, and kernel selection.
  • Experience enabling or optimizing workloads for AI accelerators, NPUs, DSPs, vector processors, matrix engines, or custom SoC IP.
  • Strong C/C++ programming skills, with good Python scripting ability for compiler tooling, testing, automation, and model workflow integration.
  • Experience working in Linux development environments, including cross-compilation, debugging, profiling, build systems, and runtime bring-up. Strong debugging and problem-solving skills across compiler IR, generated code, runtime behavior, and hardware/software interaction.
  • Ability to work with architecture and hardware teams to understand accelerator capabilities and translate them into compiler/runtime enablement.
  • Proven ability to technically lead complex software modules, mentor engineers, and drive execution across cross-functional teams.
Responsibilities
  • Architect, design, and develop AI/ML compiler and runtime software for RISC-V based IP, NPU, and SoC platforms.
  • Develop and enhance IREE-based compiler flows, including MLIR lowering, code generation, runtime integration, and deployment paths for edge AI workloads.
  • Create and maintain custom MLIR dialects, compiler passes, lowering pipelines, and transformation flows to map AI workloads efficiently to custom NPU and accelerator hardware.
  • Work across AI framework import paths including PyTorch, ONNX, and TFLite, and enable lowering through torch-mlir, TOSA, Linalg, and related MLIR dialects.
  • Optimize neural network workloads for edge deployment, including operator fusion, tiling, memory planning, quantization, layout transformation, and accelerator-aware scheduling.
  • Enable efficient execution of AI models across CPU, vector, matrix, and NPU acceleration paths, balancing latency, throughput, memory footprint, and power efficiency.
  • Collaborate closely with architecture, hardware, firmware, FPGA, validation, and product teams to bring up AI workloads on simulators, FPGA platforms, emulation environments, and silicon.
  • Analyze model performance, identify compiler/runtime bottlenecks, and drive optimizations across graph-level, operator-level, and kernel-level execution paths.
  • Define software architecture and technical direction for AI SDK components, including compiler pipelines, runtime interfaces, model deployment flows, and accelerator integration.
  • Build test infrastructure, validation flows, benchmark suites, and CI pipelines for AI compiler/runtime correctness, performance, and regression tracking.
  • Provide technical leadership to engineers working on AI compiler, runtime, model deployment, and edge AI software development.
  • Work with internal and customer-facing teams to support software enablement, debugging, performance tuning, and deployment of AI workloads on target platforms.
Desired Qualifications
  • Experience working on RISC-V, ARM, x86, DSP, GPU, or custom accelerator software stacks.
  • Familiarity with RISC-V Vector, matrix acceleration concepts, custom instructions, or accelerator-specific code generation.
  • Experience with edge AI deployment on real devices, development boards, FPGA platforms, emulators, simulators, or early silicon.
  • Familiarity with FPGA prototyping, Linux bring-up, board-level debugging, or pre-silicon software validation.
  • Exposure to LLM and edge inference stacks such as llama.cpp, GGML/GGUF, ONNX Runtime, TensorFlow Lite, TVM, XNNPACK, or similar frameworks, with understanding of quantization, memory footprint optimization, kernel performance, and deployment constraints on resource-limited devices.
  • Experience with AI model benchmarking and optimization for vision, audio, transformers, GenAI, robotics, automotive, industrial, or real-time embedded workloads.
  • Understanding of hardware-software co-design, memory hierarchy, DMA, scratchpad memory, cache behavior, and accelerator data movement. Experience with runtime systems, kernel libraries, microkernels, custom dispatch flows, or accelerator runtime APIs.
  • Familiarity with CI/CD and agile tools such as Jenkins, Git, CMake, Bazel, Jira, or similar engineering infrastructure.
  • Experience working in customer-facing enablement, silicon bring-up, platform software, or SDK delivery environments.
  • Excellent communication and interpersonal skills, with the ability to explain complex compiler and AI runtime topics clearly to software, hardware, and product stakeholders.

GlobalFoundries is a semiconductor foundry that manufactures chips for a broad range of customers, focusing on feature-rich, reliable technologies for cars, IoT, 5G, and other high-growth markets. It operates by fabricating silicon wafers for customer designs, offering mature, specialized processes rather than chasing the most advanced nodes. This makes it different from peers that prize the latest process nodes; GlobalFoundries emphasizes diversification across a global manufacturing footprint (U.S., Europe, Singapore) and stable supply for essential applications. Its goal is to be a leading supplier of high-volume, specialty manufacturing that enables reliable chips for everyday devices and growing tech markets.

Company Size

10,001+

Company Stage

IPO

Headquarters

Town of Malta, New York

Founded

2009

Get referred to GlobalFoundries

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • GlobalFoundries won a $375 million CHIPS award on September 8, 2026.
  • Communications infrastructure and data-center revenue grew 62% year over year in Q2 2026.
  • Monolithic Power Systems and Navitas partnerships expand 2027 domestic manufacturing demand.

What critics are saying

  • Smart mobile remains 36% of revenue, and management expects low-teens decline in 2026.
  • GlobalFoundries sued Tower Semiconductor in March 2026, signaling costly IP warfare and distraction.
  • If silicon photonics and quantum programs miss milestones, CHIPS-linked growth narratives evaporate quickly.

What makes GlobalFoundries unique

  • GlobalFoundries owns geographically diverse fabs across U.S., Europe, Singapore for supply security.
  • Its specialty focus on silicon photonics, SiGe, RF, and power separates it from leading-edge foundries.
  • GF's 2026 UX, GCRAM, and advanced packaging launches deepen its niche platform portfolio.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Professional Development Budget

Growth & Insights and Company News

Headcount

6 month growth

-2%

1 year growth

-1%

2 year growth

-1%
Omega Technology Solutions Group, Inc.
Sep 9th, 2026
Kepler Computing raises $468M to boost memory density without EUV lithography

Kepler Computing has raised $468 million from investors including GlobalFoundries, Intel Capital, AMD Ventures, and Bill Gates to develop memory chips that bypass expensive extreme ultraviolet lithography. The San Jose startup uses ferroelectric materials and 3D stacking to increase high-bandwidth memory and SRAM density whilst working with existing semiconductor fabs. The company claims it can convert a fab to next-generation capability in eight months versus the typical 24-month timeline. However, the technology faces contamination challenges because its composite material includes iron, requiring dedicated equipment or encapsulation. Kepler has been testing at GlobalFoundries facilities in Singapore and Vermont, producing chips on approximately 2,000 wafers. The US Department of Commerce committed up to $245 million in July to support domestic development. First HBM samples are planned for late 2025, with US production targeted for 2028.

Crypto Briefing
Sep 9th, 2026
Kepler Computing claims new chip design can solve AI memory shortage

Kepler Computing lands $245 million in CHIPS Act funding for ferroelectric RAM technology it claims can solve the AI-driven memory shortage with

Associated Press
Sep 2nd, 2026
GlobalFoundries and RAAAM tape out GCRAM chip, achieving 40% area shrink and 60% power reduction vs SRAM

GlobalFoundries and RAAAM Memory Technologies have announced a collaboration to develop Gain-Cell RAM (GCRAM) technology on GlobalFoundries' FDX platform. The partnership has achieved a significant milestone with the successful tape-out of a GCRAM test vehicle. The technology addresses growing on-chip memory demands for edge AI, automotive, and IoT devices. Using RAAAM's patented GCRAM, the companies achieved a 40% memory area reduction and up to 60% power savings compared to SRAM. The firms have jointly designed and taped out a GCRAM test chip. They plan to provide GCRAM design access to lead customers in early 2027. NXP Semiconductors is evaluating the solution. Victor Wang, vice president of front end innovation at NXP, noted the technology's potential for increasing on-chip memory capacities and improving system-level efficiency.

PR Newswire
Sep 2nd, 2026
QuickLogic launches enhanced eFPGA Hard IP for GlobalFoundries 12LP with 250K+ LUT scalability

QuickLogic has released an enhanced embedded FPGA Hard IP for GlobalFoundries' 12LP process. The new offering includes a DSP Multiply-and-Accumulate unit with enhanced preadder, cascade connections, and SIMD functionality for applications such as phased array radar, electronic warfare, and satellite communications. The eFPGA fabric can now scale beyond 250,000 LUTs for designs requiring larger embedded programmable logic. An optional Configuration Bit Health Monitor detects and corrects errors induced by Single Event Upsets whilst the design is active. The technology enables programmable logic to be embedded directly within SoCs, allowing hardware functions and interfaces to be updated post-manufacturing. QuickLogic showcased the enhanced IP at the GlobalFoundries Technology Summit North America in Santa Clara, California, on 2 September 2026.

Associated Press
Sep 1st, 2026
Navitas ships Gen 5 GaNFast semiconductors from US factory in GlobalFoundries partnership for AI infrastructure

Navitas Semiconductor has begun shipping its fifth-generation GaNFast gallium nitride power semiconductors manufactured in the United States through a partnership with GlobalFoundries. The first shipments from GlobalFoundries' Burlington, Vermont facility are scheduled for September, with strategic customer samples expected before year-end. The collaboration, announced in November 2025, combines Navitas' proprietary GaN technology with GlobalFoundries' 200mm manufacturing capabilities. The partnership aims to strengthen the domestic supply chain for AI infrastructure and critical national security applications. The initial product family includes 650V GaN FETs with various resistance specifications. Navitas has pioneered GaN power semiconductor innovation since 2014 and holds more than 300 issued and pending patents across GaN and silicon carbide technologies. The partnership marks a significant milestone in establishing trusted US-based production of advanced power semiconductors for AI infrastructure, high-performance computing, and industrial applications.