Full-Time

Machine Learning Infrastructure Engineer

Updated on 8/7/2026

xAI

xAI

1,001-5,000 employees

Generative AI chatbot leveraging real-time data

Compensation Overview

$180k - $440k/yr

+ Equity

Palo Alto, CA, USA

In Person

Bachelor's, Master's, PhD

Category
DevOps & Infrastructure (1)
Required Skills
Graphics Processing Unit (GPU)
Rust
Python
Puppet
Distributed Systems
Neural Networks
CUDA
PyTorch
Machine Learning
Computer Networking
Data Engineering
Ansible
C/C++
Linux/Unix

Get referred to xAI

See people who can refer or advise you

Requirements
  • A Bachelor, Master, post-graduate, or PhD degree in computer science, machine learning, or another quantitative discipline, or equivalent work experience.
  • At least 2 years of industry experience working with high-traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications.
  • At least 2 years of experience with machine learning platforms, training infrastructure, or close collaboration with modeling engineers and data scientists.
  • Strong proficiency with Python and experience with compiled languages such as C++ or Rust.
  • Deep familiarity with modern machine learning frameworks such as JAX or PyTorch.
  • A low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking.
  • Comfort with Linux systems and orchestration tools.
  • Experience with job schedulers such as Slurm, configuration management tools such as Puppet or Ansible, or related infrastructure tooling.
Responsibilities
  • Design, build, and scale GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on machine learning hypotheses.
  • Develop data pipelines and integrate large-scale data, training, and inference systems.
  • Collaborate with machine learning teams to productionize models and ensure seamless integration across the stack.
  • Ensure the scalability, reliability, and efficiency of large-scale machine learning systems.
  • Work across the full stack to solve complex problems independently.
  • Mentor junior engineers and contribute to the growth of the team.

xAI builds Grok, a generative AI chatbot with a Hitchhiker’s Guide-inspired persona that uses real-time data from X to produce current, culturally aware responses. Grok is accessible to premium subscribers on X, via a standalone website, mobile apps, and through an API for developers. The company also develops large-scale AI infrastructure, such as the Colossus supercomputer, to train and run its models. Its goal is to pursue truth-seeking AGI research and grow a capital-intensive platform that could support a major public offering in the future.

Company Size

1,001-5,000

Company Stage

Series E

Total Funding

$42.1B

Headquarters

Palo Alto, California

Founded

2023

Get referred to xAI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • July 8, 2026 Grok 4.5 launched at aggressive pricing, improving developer adoption.
  • February 2026 SpaceX merger gave xAI direct access to deeper capital and distribution.
  • Memphis expansion continues; July 2026 DOJ backed xAI's turbine fight, protecting Colossus uptime.

What critics are saying

  • Canada ruled June 11, 2026 that xAI and X violated privacy law over deepfakes.
  • EU, UK, Ireland, and Brazil investigations keep Grok's deepfake scandal alive through 2026.
  • xAI burns about $1 billion monthly against roughly $500 million revenue; capital dependence persists.

What makes xAI unique

  • xAI's Grok taps X's live data stream, giving fresher responses than static-model rivals.
  • July 8, 2026 Grok 4.5 broadened xAI into coding, agents, voice, and enterprise tools.
  • Colossus in Memphis gives xAI unusually large in-house training capacity and rapid iteration speed.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Remote Work Options

Growth & Insights and Company News

Headcount

6 month growth

19%

1 year growth

2%

2 year growth

19%
Data Center Dynamics
Jul 31st, 2026
Musk confirms fourth SpaceXAI data center in Memphis, company starts removing 'illegal' gas turbines.

Musk confirms fourth SpaceXAI data center in Memphis, company starts removing 'illegal' gas turbines. Macrohard gets Minihard July 31, 2026 Elon Musk appears to have confirmed that SpaceXAI is building a fourth data center in Memphis, Tennessee. It comes as the company agreed to remove 69 gas turbines that it uses to power its data centers in the area. Campaigners had launched legal action against the company, claiming the turbines were installed without the correct permits and in contravention of the Clean Air Act. xAI, Musk's AI lab, began building data centers in the Memphis area in 2024 to house Colossus, the supercomputer that powers Grok, its AI chatbot. The company has since been rolled into another of Musk's firms, SpaceX, becoming SpaceXAI and expanding into neocloud services, renting data center space to clients including Anthropic. Musk's Minihard data center. Responding to a user on X, Musk's social media platform, who had posted drone footage of a new building under construction at the Memphis data center complex, on Tulane Road near the border between Tennessee and Mississippi, the billionaire said: "To the right is Minihard, which will have the same 220k GB300s with 800 gig NICs as Macroharder, but in an improved, much denser configuration." Musk has dubbed the two existing data centers on the site Macrohard and Macroharder, an apparent dig at Microsoft. DCD reported in March that SpaceXAI had applied for a building permit for a 312,000 sq ft (28,985 sqm) building costing $659 million, but at the time its use had not been confirmed. It is a four-story building, some 75 feet high. It is not clear if SpaceXAI has sufficient power for an additional 220,000-strong deployment of GB300, which would likely require more than 400MW. The company claims to have 1GW of compute power across its Memphis data centers, but satellite imagery taken in January reportedly showed it had cooling equipment installed capable of managing 350MW. SpaceX's IPO prospectus from earlier this year said "the next phase of expansion at Colossus 2 will bring online at least 220,000 additional GB300 processors and over 400 additional megawatts of compute power," presumably referring to the new data center. DCD has contacted the firm for comment on the new project. The company's other Memphis data center is located in another part of the city, in a former Electrolux factory on Paul Lowry Road. It is also reportedly considering building a campus in Texas. SpaceXAI sets out timeline for removing problematic turbines. The presence of xAI's data centers in Memphis has been a source of controversy due to the company's use of portable natural gas turbines. While these have enabled the company to bring data center capacity online fast, campaigners say the polluting machines have damaged air quality in some of Memphis's poorest neighborhoods. In April, civil rights organization the NAACP made a complaint against xAI, and its subsidiary MZX Tech, alleging that 27 methane gas turbines that power the company's Colossus 2 data center were being run illegally, in violation of the Clean Air Act. Now the company has said it will remove 69 turbines from the site over the next year, ahead of the installation of a permanent 1.2GW power plant for the data centers, which received permitting approval in March. The company will begin removing the temporary turbines next month, though as the permanent installation will apparently comprise 41 gas turbines, neighbors may not notice much difference. A SpaceXAI statement said it had "entered into an agreed order with the Mississippi Department of Environmental Quality establishing a fixed removal timeline for all 69 of the temporary, mobile turbines." It added: "We care deeply about being good neighbors and we regularly reconfigure power operations to reduce noise from the facility and minimize our impact on the local community." More in sustainability.

CoinGape
Jul 28th, 2026
Elon Musk reveals Grok 4.6 launch timeline, teases 2.1t-parameter Grok 4.7.

Elon Musk reveals Grok 4.6 launch timeline, teases 2.1t-parameter Grok 4.7. Highlights * Elon Musk said that Grok 4.6 will launch around August 7. * He also mentioned that they will release Grok 4.7 a few weeks later. * This follows the release of Grok 4.5 this month. Elon Musk has revealed that SpaceXAI plans to release the Grok 4.6 AI model next month, while the Grok 4.7 model could launch soon after. This comes as the AI race heats up with Anthropic and OpenAI pushing their frontier models. Elon Musk reveals when Grok 4.6 will launch. In an X post, the world's richest man said that they will release Grok 4.6 around August 7 and that it will be the 1.5T model with significantly improved Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). He also revealed that they will release Grok 4.7 a few weeks later and that it will be a 2.1T model. "This will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency," Elon Musk said about the Grok 4.7 model. The planned launch of the Grok 4.6 follows the release of the Grok 4.5 earlier this month. The planned launch of the new Grok models comes as the AI race heats up, with U.S. firms like SpaceXAI, OpenAI, and Anthropic facing competition from Chinese firms such as Moonshot. Moonshot released its open-weight Kimi K3 model yesterday. The hype around Chinese AI models has led to reports that the Trump administration is weighing restrictions on this model in a bid to protect the U.S. AI leadership. However, as CoinGape reported, top tech firms such as NVIDIA have warned the U.S. against Open-Weight AI restrictions. The open letter from these firms also gained support from Elon Musk as well as OpenAI CEO Sam Altman. Grok 4.5 emerges as best cybersecurity AI model. Elon Musk's revelation came in response to Vercel CEO Guillermo Rauch's assessment of the Grok 4.5 model. Rauch stated that in their latest assessment, the AI model emerged as the best cybersecurity model on price-performance. In its latest https://t.co/p9AoezbuGt benchmarks, Grok 4.5 has emerged as the best cybersecurity AI model on price-performance. It's 10x cheaper than Sol, 5.7x cheaper than Opus 5, and 2.2x cheaper than Kimi K3, yet at Kimi-like perf. Sol remains the frontier, ahead of Opus 5. https://t.co/kcSDIqnXsN - Guillermo Rauch (@rauchg) July 27, 2026 He added that Grok 4.5 is 10x cheaper than ChatGPT's Sol model, 5.7x cheaper than Opus 5, and 2.2x cheaper than Kimi K3, despite being at Kimi-like performance. "Sol remains the frontier, ahead of Opus 5," Rauch said. Meanwhile, even as the AI race heats up, it is worth noting that Anthropic and OpenAI have united to seek a federal review framework for advanced AI models before public release. These firms propose that regulators establish consistent AI safety standards that will apply to all U.S. firms. Investment disclaimer: The content reflects the author's personal views and current market conditions. Please conduct your own research before investing in cryptocurrencies, as neither the author nor the publication is responsible for any financial losses. Ad Disclosure: This site may feature sponsored content and affiliate links. All advertisements are clearly labeled, and ad partners have no influence over its editorial content. Why Trust CoinGape * Latest * / * Trending

Inforisk Today
Jul 27th, 2026
Nvidia Launches open-source AI Security alliance.

Nvidia Launches open-source AI Security alliance. Anthropic, OpenAI and Google Absent as 37 Firms Back Open AI Security Tools Chris Riotta (@chrisriotta) - July 27, 2026 Dozens of the world's biggest technology firms are joining a new Nvidia-led initiative to build open-source security tools for artificial intelligence. The U.S.-based chipmaker announced the "new alliance" in a Monday blog post, saying companies like IBM, Palantir and Microsoft "will work to remediate and disclose vulnerabilities using open technologies" as part of an effort to better safeguard software and AI agents. The Open Secure AI Alliance is a coalition of 37 firms, including CrowdStrike, Palo Alto Networks, Dell Technologies, Hugging Face and the Linux Foundation. "Open models, like any powerful technology, can be misused - including through attempts to weaken safeguards or repurpose capabilities for cyberattacks," the blog post read. Nvidia added that those risks "must be managed wherever advanced AI is deployed." The launch comes as Washington is weighing restrictions on open-weight AI models, and just days after a breach at Hugging Face sparked industry debate over whether closed AI systems can slow incident response efforts. While the initiative includes many of the biggest names in cloud computing, enterprise software and cybersecurity, it lacks several key players in the global AI sector, including Anthropic, OpenAI and Google - developers behind some of the industry's most capable closed models. The Nvidia blog post directly referenced the security incident Hugging Face disclosed earlier this month, in which OpenAI models running on an internal hacking benchmark with manipulated cyber safety refusals escaped a test environment and gained the ability to run commands on the AI repository's production servers. Commercial closed models declined to assist forensic analysis of the intrusion, which Nvidia said left the tools unable to tell apart the attackers from defenders. Hugging Face instead turned to GLM 5.2, an open-weight model from Chinese developer Z.ai. Defenders who cannot inspect and run advanced AI on their own systems face significant constraints during cyber incidents, Nvidia said. The alliance is urging policymakers to treat open models, harnesses and security tooling as "defensive assets, not liabilities" and warns that blanket restrictions on open-frontier systems would continue to concentrate power and dependence on a handful of closed providers. The announcement also follows a separate industry letter released Friday in which Nvidia, Microsoft, Meta and dozens of other companies urged the White House to avoid premature restrictions on open-weight AI. The list of signatures nearly doubled to 50 within a day as OpenAI and Google added their names over the weekend, while Anthropic and Amazon did not participate in both efforts. Anthropic has not publicly explained its absence from the alliance or the letter, and it is unclear whether the company was invited to join either. Nvidia and the alliance's cloud, enterprise software and security firms benefit from cheaper, customizable open models, while frontier firms like Anthropic and OpenAI earn revenue by selling access to their closed systems. The open-versus-closed debate also intensified last week after White House officials accused Chinese AI firm Moonshot AI of stealing American intellectual property through model distillation. Treasury Secretary Scott Bessent threatened sanctions against Chinese developers, while the Trump administration is reportedly weighing a ban on Chinese open-source models (see: AI Firms Seek US Help Against China Model Distillation). Nvidia is contributing open models, model weights, data and agent harness research to the effort through its Nvidia Labs Object-Oriented Agent project on GitHub. Microsoft is also offering MDASH, a multimodel scanning system that orchestrates specialized AI agents to discover and prove exploitable bugs, while SpaceXAI has open sourced its Grok Build coding agent and plans to release the weights of its Grok model line. Monday's launch is the latest in a string of industry safety coalitions. Anthropic, Google, Microsoft and OpenAI founded the Frontier Model Forum in 2023 to coordinate on responsible frontier development, and the Biden administration enlisted more than 200 organizations in a federal AI safety consortium the following year (see: White House Launches First-Ever AI Safety Consortium). Nvidia noted in its blog that the group is not opposed to closed systems, saying the world needs closed and open models working together so defenders can choose the right tool for the job. Nvidia and Anthropic did not immediately respond to requests for comment.

Bank Info Security
Jul 27th, 2026
Nvidia Launches open-source AI Security alliance.

Nvidia Launches open-source AI Security alliance. Anthropic, OpenAI and Google Absent as 37 Firms Back Open AI Security Tools Chris Riotta (@chrisriotta) - July 27, 2026 Dozens of the world's biggest technology firms are joining a new Nvidia-led initiative to build open-source security tools for artificial intelligence. The U.S.-based chipmaker announced the "new alliance" in a Monday blog post, saying companies like IBM, Palantir and Microsoft "will work to remediate and disclose vulnerabilities using open technologies" as part of an effort to better safeguard software and AI agents. The Open Secure AI Alliance is a coalition of 37 firms, including CrowdStrike, Palo Alto Networks, Dell Technologies, Hugging Face and the Linux Foundation. "Open models, like any powerful technology, can be misused - including through attempts to weaken safeguards or repurpose capabilities for cyberattacks," the blog post read. Nvidia added that those risks "must be managed wherever advanced AI is deployed." The launch comes as Washington is weighing restrictions on open-weight AI models, and just days after a breach at Hugging Face sparked industry debate over whether closed AI systems can slow incident response efforts. While the initiative includes many of the biggest names in cloud computing, enterprise software and cybersecurity, it lacks several key players in the global AI sector, including Anthropic, OpenAI and Google - developers behind some of the industry's most capable closed models. The Nvidia blog post directly referenced the security incident Hugging Face disclosed earlier this month, in which OpenAI models running on an internal hacking benchmark with manipulated cyber safety refusals escaped a test environment and gained the ability to run commands on the AI repository's production servers. Commercial closed models declined to assist forensic analysis of the intrusion, which Nvidia said left the tools unable to tell apart the attackers from defenders. Hugging Face instead turned to GLM 5.2, an open-weight model from Chinese developer Z.ai. Defenders who cannot inspect and run advanced AI on their own systems face significant constraints during cyber incidents, Nvidia said. The alliance is urging policymakers to treat open models, harnesses and security tooling as "defensive assets, not liabilities" and warns that blanket restrictions on open-frontier systems would continue to concentrate power and dependence on a handful of closed providers. The announcement also follows a separate industry letter released Friday in which Nvidia, Microsoft, Meta and dozens of other companies urged the White House to avoid premature restrictions on open-weight AI. The list of signatures nearly doubled to 50 within a day as OpenAI and Google added their names over the weekend, while Anthropic and Amazon did not participate in both efforts. Anthropic has not publicly explained its absence from the alliance or the letter, and it is unclear whether the company was invited to join either. Nvidia and the alliance's cloud, enterprise software and security firms benefit from cheaper, customizable open models, while frontier firms like Anthropic and OpenAI earn revenue by selling access to their closed systems. The open-versus-closed debate also intensified last week after White House officials accused Chinese AI firm Moonshot AI of stealing American intellectual property through model distillation. Treasury Secretary Scott Bessent threatened sanctions against Chinese developers, while the Trump administration is reportedly weighing a ban on Chinese open-source models (see: AI Firms Seek US Help Against China Model Distillation). Nvidia is contributing open models, model weights, data and agent harness research to the effort through its Nvidia Labs Object-Oriented Agent project on GitHub. Microsoft is also offering MDASH, a multimodel scanning system that orchestrates specialized AI agents to discover and prove exploitable bugs, while SpaceXAI has open sourced its Grok Build coding agent and plans to release the weights of its Grok model line. Monday's launch is the latest in a string of industry safety coalitions. Anthropic, Google, Microsoft and OpenAI founded the Frontier Model Forum in 2023 to coordinate on responsible frontier development, and the Biden administration enlisted more than 200 organizations in a federal AI safety consortium the following year (see: White House Launches First-Ever AI Safety Consortium). Nvidia noted in its blog that the group is not opposed to closed systems, saying the world needs closed and open models working together so defenders can choose the right tool for the job. Nvidia and Anthropic did not immediately respond to requests for comment.

Informa
Jul 27th, 2026
Nvidia Launches open-source AI Security alliance.

Nvidia Launches open-source AI Security alliance. Anthropic, OpenAI and Google Absent as 37 Firms Back Open AI Security Tools Chris Riotta (@chrisriotta) - July 27, 2026 Dozens of the world's biggest technology firms are joining a new Nvidia-led initiative to build open-source security tools for artificial intelligence. The U.S.-based chipmaker announced the "new alliance" in a Monday blog post, saying companies like IBM, Palantir and Microsoft "will work to remediate and disclose vulnerabilities using open technologies" as part of an effort to better safeguard software and AI agents. The Open Secure AI Alliance is a coalition of 37 firms, including CrowdStrike, Palo Alto Networks, Dell Technologies, Hugging Face and the Linux Foundation. "Open models, like any powerful technology, can be misused - including through attempts to weaken safeguards or repurpose capabilities for cyberattacks," the blog post read. Nvidia added that those risks "must be managed wherever advanced AI is deployed." The launch comes as Washington is weighing restrictions on open-weight AI models, and just days after a breach at Hugging Face sparked industry debate over whether closed AI systems can slow incident response efforts. While the initiative includes many of the biggest names in cloud computing, enterprise software and cybersecurity, it lacks several key players in the global AI sector, including Anthropic, OpenAI and Google - developers behind some of the industry's most capable closed models. The Nvidia blog post directly referenced the security incident Hugging Face disclosed earlier this month, in which OpenAI models running on an internal hacking benchmark with manipulated cyber safety refusals escaped a test environment and gained the ability to run commands on the AI repository's production servers. Commercial closed models declined to assist forensic analysis of the intrusion, which Nvidia said left the tools unable to tell apart the attackers from defenders. Hugging Face instead turned to GLM 5.2, an open-weight model from Chinese developer Z.ai. Defenders who cannot inspect and run advanced AI on their own systems face significant constraints during cyber incidents, Nvidia said. The alliance is urging policymakers to treat open models, harnesses and security tooling as "defensive assets, not liabilities" and warns that blanket restrictions on open-frontier systems would continue to concentrate power and dependence on a handful of closed providers. The announcement also follows a separate industry letter released Friday in which Nvidia, Microsoft, Meta and dozens of other companies urged the White House to avoid premature restrictions on open-weight AI. The list of signatures nearly doubled to 50 within a day as OpenAI and Google added their names over the weekend, while Anthropic and Amazon did not participate in both efforts. Anthropic has not publicly explained its absence from the alliance or the letter, and it is unclear whether the company was invited to join either. Nvidia and the alliance's cloud, enterprise software and security firms benefit from cheaper, customizable open models, while frontier firms like Anthropic and OpenAI earn revenue by selling access to their closed systems. The open-versus-closed debate also intensified last week after White House officials accused Chinese AI firm Moonshot AI of stealing American intellectual property through model distillation. Treasury Secretary Scott Bessent threatened sanctions against Chinese developers, while the Trump administration is reportedly weighing a ban on Chinese open-source models (see: AI Firms Seek US Help Against China Model Distillation). Nvidia is contributing open models, model weights, data and agent harness research to the effort through its Nvidia Labs Object-Oriented Agent project on GitHub. Microsoft is also offering MDASH, a multimodel scanning system that orchestrates specialized AI agents to discover and prove exploitable bugs, while SpaceXAI has open sourced its Grok Build coding agent and plans to release the weights of its Grok model line. Monday's launch is the latest in a string of industry safety coalitions. Anthropic, Google, Microsoft and OpenAI founded the Frontier Model Forum in 2023 to coordinate on responsible frontier development, and the Biden administration enlisted more than 200 organizations in a federal AI safety consortium the following year (see: White House Launches First-Ever AI Safety Consortium). Nvidia noted in its blog that the group is not opposed to closed systems, saying the world needs closed and open models working together so defenders can choose the right tool for the job. Nvidia and Anthropic did not immediately respond to requests for comment.