Full-Time

AI Systems Performance Specialist

Posted on 9/8/2026

Bright Vision Technologies

Bright Vision Technologies

51-200 employees

AI-powered talent intelligence and automation platform

Compensation Overview

$130k - $180k/yr

No H1B Sponsorship

Remote in USA

Remote

Bachelor's, Master's

Category
AI & Machine Learning (1)
Required Skills
LLM
Microsoft Azure
Python
Distributed Systems
Software Testing
CUDA
Computer Networking
AWS
C/C++
Google Cloud Platform

Get referred to Bright Vision Technologies

See people who can refer or advise you

Requirements
  • A Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline.
  • At least 10 years of professional experience in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing, or distributed computing.
  • Expert-level programming skills in Python and C++.
  • Extensive experience optimizing GPU-accelerated AI workloads using CUDA, distributed training frameworks, and modern deep learning libraries.
  • Strong knowledge of Large Language Models, deep learning frameworks, model serving, and production AI inference.
  • Hands-on experience with profiling tools such as NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar performance analysis tools.
  • Experience deploying and optimizing AI workloads on AWS, Microsoft Azure, or Google Cloud Platform.
  • Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture.
  • Excellent analytical, troubleshooting, communication, and technical leadership skills.
Responsibilities
  • Optimize AI training and inference pipelines for maximum throughput, low latency, scalability, and infrastructure efficiency.
  • Analyze and improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production AI workloads.
  • Design and implement optimization techniques including quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism.
  • Profile AI applications using industry-standard performance analysis tools and identify bottlenecks across compute, memory, networking, and storage.
  • Optimize distributed training and inference using NCCL, DeepSpeed, PyTorch Distributed, Ray, MPI, or similar distributed computing frameworks.
  • Collaborate with AI researchers, machine learning engineers, platform engineers, and infrastructure teams to improve model performance and production reliability.
  • Build automated benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines.
  • Evaluate emerging AI hardware, GPU architectures, inference frameworks, and optimization technologies to improve enterprise AI capabilities.
  • Drive AI infrastructure cost optimization through efficient resource utilization, cloud optimization, and FinOps best practices.
  • Mentor engineering teams and provide technical leadership on AI systems architecture, GPU optimization, and performance engineering.
Desired Qualifications
  • Experience optimizing production-scale Large Language Model inference and serving large foundation models.
  • Hands-on experience with vLLM, TensorRT-LLM, DeepSpeed, Triton Inference Server, CUTLASS, FasterTransformer, or similar AI optimization frameworks.
  • Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference optimization techniques.
  • Experience implementing FinOps strategies for AI infrastructure cost optimization and resource management.
  • Contributions to AI systems research, open-source AI infrastructure projects, patents, or technical publications.
  • Familiarity with emerging AI accelerator technologies, including AMD ROCm, Intel oneAPI, or custom AI hardware.
Bright Vision Technologies

Bright Vision Technologies

View

Bright Vision Technologies offers Lumina, an AI-powered platform for talent intelligence and enterprise automation, along with consulting and staffing services. Lumina analyzes unstructured data with generative AI, Large Language Models orchestrated via LangChain, and Retrieval-Augmented Generation to support sourcing and screening candidates, integrated with CRM, ERP, and HRMS on a cloud-native microservices architecture across AWS, Azure, and Google Cloud. The company differentiates itself by combining a sophisticated AI product with hands-on consulting and staffing expertise, plus its minority-owned status and a dual model that helps clients implement technology and solve broader business challenges. Its goal is to streamline and automate complex workflows in IT talent acquisition and management to enable digital transformation and efficient enterprise operations.

Company Size

51-200

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2020

Get referred to Bright Vision Technologies

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • April 2026 LinkedIn launch signals active product commercialization and go-to-market acceleration.
  • August 2026 site updates show hiring across Bridgewater, Princeton, and Chennai.
  • The company’s minority-owned status and U.S.-India footprint strengthen enterprise procurement positioning.

What critics are saying

  • No disclosed funding, customers, or revenue makes Lumina's traction impossible to verify.
  • The June 2026 pivot from staffing to product risks channel conflict and execution drag.
  • Enterprise AI recruiting faces entrenched rivals like Workday, LinkedIn, and Eightfold by 2027.

What makes Bright Vision Technologies unique

  • April 2026 launch of Lumina combines talent intelligence, automation, and hybrid cloud.
  • Lumina integrates RAG, LLM orchestration, semantic matching, and blockchain credential verification.
  • Bright Vision serves staffing and consulting clients, enabling implementation alongside the product.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Remote Work Options