Full-Time

Technical Support Engineer

Inference, US Weekends

Together AI

Together AI

201-500 employees

Open-source AI research via decentralized cloud

Compensation Overview

$160k - $230k/yr

+ Equity

Remote in USA

Remote

US daytime hours are required, with weekend coverage and a transition to a four-day weekend shift after ramp-up.

Category
IT & Security (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
Microsoft Azure
Python
JavaScript
Grafana
High Performance Computing (HPC)
Postman
Git
Machine Learning
Infrastructure as Code (IaC)
TypeScript
AWS
Prometheus
Ansible
REST APIs
DevOps
Google Cloud Platform

Get referred to Together AI

See people who can refer or advise you

Requirements
  • 6+ years of experience in a customer-facing technical role, Site Reliability Engineering, DevOps, or infrastructure engineering, including at least 1 year supporting an artificial intelligence service.
  • Experience as a Site Reliability Engineer or DevOps Engineer working with Kubernetes.
  • Knowledge of artificial intelligence, machine learning, graphics processing unit technologies, and their integration into high-performance computing environments.
  • Production-level experience with infrastructure services such as Kubernetes and SLURM, infrastructure as code solutions such as Ansible, high-performance network fabrics, NFS-based storage management, and container infrastructure.
  • Familiarity with operating Vast and Weka storage systems in high-performance computing environments.
  • Ability to diagnose complex network-layer issues and read traces.
  • Strong knowledge of Python, TypeScript, and/or JavaScript, with testing and debugging experience using curl and Postman-like tools.
  • Expertise with observability tooling such as Prometheus and Grafana at scale.
  • Familiarity with REST API debugging and HTTP semantics.
  • Experience with large language model inference frameworks, LoRA fine-tuning, and common training failure modes.
  • Experience with infrastructure as code and Git-based workflows.
  • Background in graphics processing unit cluster management.
  • Experience with cloud platforms including AWS, GCP, and/or Azure.
  • Foundational understanding of installing, configuring, administering, troubleshooting, and securing compute clusters.
  • Ability to solve complex technical problems and troubleshoot issues proactively.
  • Ability to work cross-functionally with Sales, Engineering, Support, Product, and Research teams to drive customer success.
  • Ability to explain complex technical concepts to non-technical stakeholders.
  • Ability to operate in dynamic environments, manage multiple projects, and frequently switch context and priorities.
Responsibilities
  • Engage directly with customers to resolve complex technical challenges involving GPU clusters and inference and fine-tuning services.
  • Act as a customer-facing Site Reliability Engineer to keep customer inference endpoints running on Kubernetes healthy, stable, and performant.
  • Become a product expert in Together AI's generative artificial intelligence solutions and serve as the final technical defense before escalation to Engineering and Product.
  • Assist with hardware and platform migrations by validating system health and traffic routing, monitoring dashboards for anomalies, and escalating with data-backed analysis.
  • Manage customer-facing communications during incidents and degradations, translating technical findings into clear, evidence-backed updates without exposing platform internals.
  • Contribute infrastructure changes for model deployment, capacity rebalancing, and cluster configuration through pull requests and infrastructure as code.
  • Execute infrastructure changes for endpoint configuration, model bring-up and bring-down, and capacity scaling.
  • Flag engine-level bugs for Engineering with logs and reproduction steps.
  • Collaborate with Engineering, Research, and Product teams and with senior internal and external leaders to address customer concerns and support customer satisfaction.
  • Identify patterns in support cases and work with Engineering and Go-To-Market teams to inform the product roadmap.
  • Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and frequently asked questions.
  • Provide support coverage during holidays, nights, and weekends as required.

Together AI provides open-source AI tools and decentralized cloud services to train, fine-tune, and deploy generative models for researchers, developers, and organizations. It runs tasks in the cloud where users run training jobs, manage model versions, and deploy applications via subscriptions and usage fees. It differentiates itself by prioritizing open-source, transparency, and a decentralized cloud approach instead of a proprietary stack. Its goal is to broaden access to powerful AI and build open, verifiable AI systems that benefit society through shared technology.

Company Size

201-500

Company Stage

Series C

Total Funding

$1.3B

Headquarters

Menlo Park, California

Founded

2022

Get referred to Together AI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Together AI reported 400 trillion tokens monthly on inference in August 2026.
  • L&T's Chennai project adds 10,000 NVIDIA B300 GPUs for expansion.
  • Aramco Ventures, Nvidia, and Salesforce Ventures backed the $800 million Series C.

What critics are saying

  • The 2024 privacy class action in Northern California still hangs over 2026.
  • OpenAI, Anthropic, and Meta keep squeezing pricing and model quality.
  • GPU dependence on IBM and L&T creates supply-chain choke points into 2027.

What makes Together AI unique

  • July 2026 Series C valued Together AI at $8.3 billion.
  • Together AI serves open-source inference, training, fine-tuning, and agentic workflows.
  • IBM chose Together AI for a first dedicated B300 inference cluster.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

0%
Dairy Dimension
Aug 19th, 2026
L&T's transformation: from dairy machines to AI factories and advanced technology.

L&T's transformation: from dairy machines to AI factories and advanced technology. * August 19, 2026 * 3 minutes read Larsen & Toubro (L&T) is entering a new phase of transformation, moving beyond its traditional engineering and infrastructure businesses into artificial intelligence, semiconductors, electronics, green energy, defence technology and other advanced industries. The shift comes after L&T secured a major contract to develop an AI infrastructure facility in Chennai for US-based AI cloud company Together AI. The project is expected to deploy around 10,000 NVIDIA B300 GPUs to support large-scale AI workloads. The project represents L&T's entry into the AI Factory business and demonstrates how the company is applying its traditional strengths in engineering, power infrastructure, cooling, procurement and large-scale project execution to the rapidly growing AI economy. From dairy equipment to digital infrastructure. L&T's transformation has deep historical roots. The company's origins date back to 1938, when Danish engineers Henning Holck-Larsen and Søren Kristian Toubro established a business in Bombay that initially imported dairy equipment. Over the decades, the company expanded into engineering, construction, manufacturing, infrastructure, energy and defence. It later exited several businesses, including cement, glass and tractors, to concentrate capital around its engineering capabilities. Today, L&T is once again broadening its portfolio - but this time toward technology-led businesses. A ₹5,000 crore electronics investment. L&T has announced plans to invest around ₹5,000 crore over five years in a new Electronic Products & Systems business. The business will focus on: * Power electronics * Industrial robotics and automation * Mobility * Strategic electronics * Product development * Electronics manufacturing The company is developing manufacturing capabilities in Coimbatore, alongside engineering and R&D operations in Bengaluru and Coimbatore. This represents a shift from simply executing large projects for customers toward developing products and technologies that can be manufactured and sold repeatedly. Expanding into semiconductors and AI. L&T's technology ambitions extend beyond AI infrastructure. The company has been developing its semiconductor business through L&T Semiconductor Technologies and has expanded its exposure to AI and cloud infrastructure. Its AI infrastructure business is developing capabilities spanning hyperscale AI data centres, sovereign cloud platforms, GPU-as-a-Service and managed AI platforms. The Chennai AI Factory is expected to combine high-performance computing, networking, storage and specialised AI infrastructure to support demanding workloads. Green energy and energy storage. L&T is also targeting technologies supporting the energy transition. Its businesses include electrolyser manufacturing for green hydrogen and battery energy-storage systems. These investments allow L&T to participate not only in building renewable-energy infrastructure but also in manufacturing equipment and systems required by the emerging clean-energy economy. Defence becomes increasingly technology-driven. Defence is another important part of the transformation. L&T's Precision Engineering & Systems business is developing indigenous systems across areas such as strategic electronics, radars, drones, launch systems, land platforms and underwater systems. This reflects a broader move from conventional defence contracting toward technology development, advanced manufacturing and intellectual-property-led products. Exploring rare-earth magnets. L&T has also shown interest in rare-earth permanent magnets, which are important components in electric vehicles, renewable-energy equipment and several advanced industrial applications. The company was among the companies that bid under the Indian government's scheme aimed at establishing domestic rare-earth permanent-magnet manufacturing capacity. While the opportunity is still developing, it fits into L&T's broader strategy of building capabilities around technologies that could become strategically important to India. The traditional L&T remains the foundation. Despite these new investments, infrastructure and energy continue to dominate L&T's business. The company's traditional engineering and construction operations provide the scale, cash flow, engineering expertise and customer relationships needed to invest in newer businesses. The strategy is therefore not about abandoning L&T's core operations. Instead, the company is attempting to use its existing engineering capabilities to move into higher-value technology and manufacturing segments. A new growth model. L&T's latest moves suggest a broader change in its business model - from primarily building infrastructure for customers to increasingly owning and developing the technologies behind that infrastructure. The NVIDIA B300 AI Factory is a strong example of this transition. Although the project remains an infrastructure contract, it places L&T directly within the computing infrastructure supporting the global AI economy. With investments spanning AI, semiconductors, electronics, green hydrogen, energy storage, defence and potentially rare-earth magnets, L&T is positioning its next phase of growth around advanced technology and manufacturing.

Dalal Street Investment Journal
Aug 13th, 2026
Rs 7,79,000 crore Order Book: this heavy civil infrastructure company Secures Mega order to build India's largest NVIDIA B300 AI Factory.

Rs 7,79,000 crore Order Book: this heavy civil infrastructure company Secures Mega order to build India's largest NVIDIA B300 AI Factory. Larsen & Toubro, through Vyoma.AI and its subsidiary LTN Compute, will deploy 10,000 NVIDIA B300 GPUs at its Chennai data centre campus for US-based AI cloud company Together AI. Key takeaways. On Thursday, Indian equity benchmark indices traded lower, with the benchmark Nifty 50 index falling 40.10 points, or 0.16 per cent, to 24,395.85. Amid the market movement, Larsen & Toubro share price rose 1.26 per cent to Rs 4,070.70 after the company announced that it had secured a mega order as part of a strategic partnership with US-based AI cloud platform Together AI to build India's largest single-cluster NVIDIA B300 AI Factory. The development marks L&T's foray into the AI Factory business. L&T also has a strong Order Book of Rs 7,79,000 crore, equivalent to Rs 7.79 lakh crore, providing significant revenue visibility for the company L&T Secures Mega AI Infrastructure Order Looking for stable blue chip investment opportunities? Explore DSIJ's Large Rhino - a research-driven service focused on fundamentally strong Large-Cap companies known for stability, consistent growth, and long-term wealth creation potential. Larsen & Toubro, through its AI infrastructure subsidiary LTN Compute under Vyoma.AI, has secured the project to establish an NVIDIA B300 AI Factory at its Chennai data centre campus. The facility will support Together AI's AI-native cloud platform for large-scale inference, fine-tuning and training workloads. The project involves an integrated AI Factory with capacity for 10,000 NVIDIA B300 GPUs. The platform will combine hyperscale data centre infrastructure, accelerated computing, high-performance networking, ultra-low-latency interconnects, high-throughput parallel storage and AI infrastructure operations. The company's filing classifies the project as a Mega order, with its order-classification framework defining the Mega category as orders valued between Rs 10,000 crore and Rs 15,000 crore. Chennai Campus to Support AI Infrastructure Expansion The AI Factory will be hosted at Vyoma's Chennai data centre campus, which is described as a gigawatt-scale AI infrastructure site. Phase 1 of the campus has been designed for 250 MW, with power infrastructure readiness of 150 MVA, providing a scalable base for future AI Factory expansion. The facility is expected to strengthen India's AI infrastructure ecosystem while supporting global AI workloads. The platform will enable customers to deploy and scale AI workloads through an integrated infrastructure stack covering computing, networking, storage and AI infrastructure operations. Also Read - Low PE, High ROCE Solar Pumps Stock Bags Rs 78 Crore Rooftop Solar Project in Bihar Under PM Surya Ghar Scheme; Check Details L&T Enters AI Factory Business The latest project marks L&T's entry into the AI Factory business through its digital infrastructure operations. Vyoma.AI is L&T's sovereign, secure and integrated AI cloud and hyperscale data centre business, while LTN Compute is building AI-ready digital infrastructure across India through hyperscale AI data centres, sovereign cloud platforms, AI Factory services, GPU-as-a-Service and managed AI platforms. The company's latest partnership with Together AI comes as demand for high-performance computing infrastructure continues to increase with the expansion of artificial intelligence and generative AI applications. Reuters also reported that the project is worth up to Rs 15,000 crore. Strategic Partnership with Together AI Together AI is a US-based full-stack AI cloud platform that provides accelerated infrastructure, open foundation models and developer services for organisations building, training and deploying generative AI applications. L&T Chairman and Managing Director S N Subrahmanyan said the deployment represents a significant milestone in the company's Gigawatt AI Infrastructure Mission and its objective of supporting next-generation AI infrastructure in India. Together AI Co-founder and CEO Vipul Ved Prakash said the company partnered with L&T for its scale, resilience and engineering capabilities.

Yahoo Finance
Aug 13th, 2026
L&T wins $1.57B order from Together AI to build India's largest AI data centre with 10,000 Nvidia chips

India's Larsen and Toubro has secured an order worth up to 150 billion rupees ($1.57 billion) from US-based cloud platform Together AI to host an AI data centre using Nvidia's high-performance chips. The company's shares rose as much as 1.2% in midday trade following the announcement. The order, secured by L&T's AI infrastructure subsidiary LTN Compute, ranges between 100 billion and 150 billion rupees and marks L&T's entry into the AI factory business. The centre will be hosted at the firm's Vyoma.AI unit in Chennai and will have a capacity of 10,000 Nvidia B300 chips. L&T described this as the largest such deployment in India. The facility will support Together AI's cloud platform for AI inference, fine-tuning, and training.

Yahoo Finance
Aug 12th, 2026
IBM lands $240M AI deal with Together AI for Nvidia-powered inference cluster

IBM signed a $240 million deal with Together AI to build an Nvidia-powered AI inference cluster on IBM Cloud, announced 11 August. The cluster will use roughly 2,000 Nvidia Blackwell 300 chips and is expected to sell out two to three months before going live. The deal comes as IBM faces challenges turning scattered growth into sustained momentum. Software revenue rose 5% year-over-year in the most recent quarter, with Red Hat up 11%. However, overall revenue grew just 1% to $17.2 billion in the quarter ended 30 June, prompting IBM to cut full-year guidance from more than 5% growth to 4-5%. IBM shares have fallen 33% from their 52-week high and remain down 19% for the year.

Analytics India Magazine
Aug 11th, 2026
IBM and Together AI sign $240 mn deal to build NVIDIA B300 inference cluster.

IBM and Together AI sign $240 mn deal to build NVIDIA B300 inference cluster. Together AI will use the infrastructure to increase inference capacity and bring down the cost of serving open-source models to enterprise customers. AUGUST 11, 2026, 7:23 PM IBM and Together AI have signed a multi-year $240 million agreement to deploy a large cluster of NVIDIA HGX B300 systems on IBM Cloud, with the infrastructure expected to be available in the first quarter of 2027. Together AI will use the cluster to provide inference for open-source AI models. The deployment will be the first dedicated large-scale inference cluster on IBM Cloud using NVIDIA HGX B300 systems and NVIDIA Spectrum-X Ethernet networking. The companies said the infrastructure will help Together AI increase inference capacity while reducing the cost of serving AI models to enterprise customers. "Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale," said Vipul Ved Prakash, CEO of Together AI. "Working alongside IBM with NVIDIA gives us that foundation. This cluster lets us bring production-grade inference to more companies, faster, and it's a big step in our push to make open-source AI the obvious choice for enterprises," he added. Together AI said its inference business now processes 400 trillion tokens a month. The company raised $800 million in a Series C round at an $8.3 billion valuation, with the funding going towards its AI Native Cloud. Its platform covers inference, model training, fine-tuning and agentic workflows. The company has positioned open-source models as an alternative to closed AI systems, with developers able to build applications using modular model and infrastructure stacks. Together AI selected IBM and NVIDIA based on their infrastructure roadmaps and ability to provide GPU capacity at the pace required for its expansion. The companies also pointed to the cost of serving tokens as a key consideration. The cluster will use NVIDIA HGX B300 systems connected through NVIDIA Spectrum-X Ethernet networking. NVIDIA said its B300 platform can deliver 30 times more AI factory output than prior generations. IBM said the infrastructure will also give Together AI a path into the enterprise market by combining its inference platform with IBM Cloud infrastructure. "Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes," said Alan Peacock, General Manager of IBM Cloud. "IBM and NVIDIA are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure." The deal also expands an existing collaboration between IBM and NVIDIA. The companies have been working across GPU-native data analytics, unstructured data extraction, on-premises and cloud infrastructure, and consulting services. Dion Harris, Senior Director of HPC and AI Infrastructure Solutions at NVIDIA, said the infrastructure would help enterprises deploy open-source AI for real-time services. "AI factories are becoming essential enterprise infrastructure - like electricity and telecommunications - turning compute and data into intelligence," Harris said. Your reaction Discussion. No comments yet - be the first to share your view.