Full-Time

AI Infrastructure Engineer

Posted on 9/9/2026

Bright Vision Technologies

Bright Vision Technologies

51-200 employees

AI-powered talent intelligence and automation platform

Compensation Overview

$100k - $160k/yr

No H1B Sponsorship

Remote in USA

Remote

Bachelor's, Master's

Category
DevOps & Infrastructure (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
Python
PyTorch
Computer Networking
Data Engineering
Go
Observability
C/C++
DevOps
Linux/Unix

Get referred to Bright Vision Technologies

See people who can refer or advise you

Requirements
  • A Bachelor's or Master's degree in Computer Science or a related field is required.
  • Ten or more years of experience in infrastructure, platform, or high-performance computing engineering is required.
  • Hands-on experience operating GPU clusters or large-scale machine learning training infrastructure is required.
  • Strong proficiency in Python and at least one systems language such as Go or C++ is required.
  • Deep understanding of distributed training, accelerator architectures, and collective communication is required.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for machine learning workloads is required.
  • Strong understanding of Linux internals, networking, and high-performance storage is required.
  • Experience with at least one major cloud provider's machine learning infrastructure offerings is required.
  • Strong software engineering practices, including testing, continuous integration and continuous delivery, and code review, are required.
  • Excellent communication and cross-functional collaboration skills are required.
Responsibilities
  • Design and operate GPU and accelerator infrastructure for training and inference across on-premises clusters, cloud-managed services, and hybrid configurations.
  • Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.
  • Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering.
  • Operate high-performance storage systems and data pipelines that keep accelerators supplied with training data at near-line rate.
  • Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication.
  • Build observability for artificial intelligence workloads, including utilization, throughput, training stability, and failure-mode analytics.
  • Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale.
  • Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing.
  • Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently.
  • Partner with research and applied machine learning teams to plan capacity for upcoming training runs.
  • Implement security controls, isolation, and access management for multi-tenant artificial intelligence infrastructure.
  • Drive automation across cluster provisioning, lifecycle management, and configuration enforcement.
  • Maintain runbooks, capacity dashboards, and operational documentation for the artificial intelligence platform.
  • Stay current with artificial intelligence infrastructure research, accelerator hardware, and emerging open-source artificial intelligence tooling.
Desired Qualifications
  • Experience operating InfiniBand or RDMA networking at scale.
  • Contributions to open-source machine learning infrastructure projects.
  • Familiarity with custom orchestrators or research-grade training stacks.
  • Exposure to frontier model training operations.
  • Experience with FinOps for artificial intelligence workloads.
Bright Vision Technologies

Bright Vision Technologies

View

Bright Vision Technologies offers Lumina, an AI-powered platform for talent intelligence and enterprise automation, along with consulting and staffing services. Lumina analyzes unstructured data with generative AI, Large Language Models orchestrated via LangChain, and Retrieval-Augmented Generation to support sourcing and screening candidates, integrated with CRM, ERP, and HRMS on a cloud-native microservices architecture across AWS, Azure, and Google Cloud. The company differentiates itself by combining a sophisticated AI product with hands-on consulting and staffing expertise, plus its minority-owned status and a dual model that helps clients implement technology and solve broader business challenges. Its goal is to streamline and automate complex workflows in IT talent acquisition and management to enable digital transformation and efficient enterprise operations.

Company Size

51-200

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2020

Get referred to Bright Vision Technologies

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • April 2026 LinkedIn launch signals active product commercialization and go-to-market acceleration.
  • August 2026 site updates show hiring across Bridgewater, Princeton, and Chennai.
  • The company’s minority-owned status and U.S.-India footprint strengthen enterprise procurement positioning.

What critics are saying

  • No disclosed funding, customers, or revenue makes Lumina's traction impossible to verify.
  • The June 2026 pivot from staffing to product risks channel conflict and execution drag.
  • Enterprise AI recruiting faces entrenched rivals like Workday, LinkedIn, and Eightfold by 2027.

What makes Bright Vision Technologies unique

  • April 2026 launch of Lumina combines talent intelligence, automation, and hybrid cloud.
  • Lumina integrates RAG, LLM orchestration, semantic matching, and blockchain credential verification.
  • Bright Vision serves staffing and consulting clients, enabling implementation alongside the product.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Remote Work Options