Full-Time

LLM / Agentic Evaluation Rig Engineer

Posted on 8/22/2026

Phizenix

Phizenix

No salary listed

Hyderabad, Telangana, India

Hybrid

Category
AI & Machine Learning (1)
Required Skills
LLM
Python
GitHub Actions
Data Visualization
RAG
AWS
LangGraph
DevOps

Get referred to Phizenix

See people who can refer or advise you

Requirements
  • At least 4 years of experience in software or machine learning engineering, including hands-on work building large language model evaluation or quality tooling.
  • A strong understanding of grounding, faithfulness, and hallucination, including how to measure them rigorously.
  • Strong Python skills and solid engineering practices involving reproducibility and continuous integration/continuous delivery.
  • Ability to design evaluations for non-deterministic systems without producing flaky or meaningless metrics.
  • Familiarity with large language model evaluation frameworks and large-language-model-as-judge patterns.
Responsibilities
  • Build and curate evaluation datasets, including adversarial and edge-case sets with ground-truth labels.
  • Build scorers for grounding, faithfulness, hallucination, factual consistency, and structured-output validity.
  • Combine rule-based checks, reference-based metrics, and large-language-model-as-judge methods where appropriate.
  • Verify that generated claims map to verified source data and contain no unsupported statements.
  • Build reproducible harnesses that run evaluations across model, prompt, and agent versions.
  • Integrate evaluation into continuous integration so grounding and faithfulness regressions block releases.
  • Track quality over time using dashboards and clear pass/fail thresholds.
  • Evaluate multi-step and agentic flows, including routing, tool use, verification, and confirmation.
  • Build trace capture and step-level scoring for agent runs.
  • Detect where an agentic flow silently degrades.
  • Partner with the Staff AI Engineer to turn findings into model, prompt, and orchestration improvements.
  • Partner with QA to integrate AI evaluation into the broader release process.
Desired Qualifications
  • Experience evaluating agentic or multi-step large language model systems.
  • Familiarity with retrieval-augmented generation, structured output, and managed large language models in a virtual private cloud.
  • Experience in the financial technology or financial-services domain, or in another high-stakes, correctness-critical AI environment.
  • A background in statistics or measurement and metrics design.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A

Get referred to Phizenix

See people who can refer or advise you