Full-Time

Agentic AI Evaluation Engineer

Ernst & Young

Ernst & Young

10,001+ employees

Provides consulting, assurance, and tax services

No salary listed

Kolkata, West Bengal, India

In Person

Bachelor's, Master's

Category
Software Engineering (1)
Required Skills
LLM
Microsoft Azure
Python
Software Testing
Data Visualization
Git
OpenTelemetry
A/B Testing
Docker
RAG
Quality Assurance (QA)
Observability
DevOps
Data Analysis

Get referred to Ernst & Young

See people who can refer or advise you

Requirements
  • A bachelor's or master's degree in data science, Statistics, Engineering, Operational Research, or another related field with a strong focus on modern data architectures, processes, and environments.
  • 4–7+ years of relevant experience in ML, AI, GenAI, agentic engineering, NLP/LLMs, evaluation engineering, applied research, security testing, red teaming, or building evaluation harnesses for safety, reliability, and robustness.
  • Strong hands-on Python skills for building evaluation harnesses, including data processing, metric computation, orchestration, and reporting pipelines.
  • Practical understanding of GenAI architectures, including retrieval-augmented generation, embeddings, vector search, prompt orchestration, tool calling, multi-agent systems, memory, and routing.
  • Experience designing metrics and evaluation methods, including rubrics, automated scoring, sampling strategies, and regression design.
  • Familiarity with LLM risks and mitigations, including data leakage, hallucinations, prompt injection, unsafe content, and bias.
  • Understanding of the OWASP Top 10 for LLM Applications and its translation into test cases and controls.
  • Experience with adversarial testing approaches, including jailbreak prompts, injection patterns, tool misuse scenarios, and retrieval poisoning patterns.
  • Familiarity with secure-by-design practices for LLM applications, including least privilege, safe tool invocation, output encoding and validation, and monitoring.
  • Experience with evaluation frameworks and tooling such as RAGAS, DeepEval, LangSmith, Phoenix/Arize, and custom evaluation harnesses.
  • Experimentation practices including an A/B testing mindset, baseline comparisons, and statistical rigor for sample sizes.
  • Observability and tracing experience with structured logging, OpenTelemetry, Langfuse-style traces, and dashboards.
  • Basic DevOps practices including Git, continuous integration and continuous delivery, containerization with Docker, and reproducible environments.
  • Strong written communication for producing clear evaluation plans and reports for technical and non-technical stakeholders.
  • Ability to challenge assumptions constructively and influence engineering teams toward remediation.
  • Ability to operate in ambiguity with fast-evolving GenAI tooling and risk landscapes.
  • Excellent written, oral, presentation, and facilitation skills.
  • Ability to coordinate multiple projects and initiatives through prioritization, organization, flexibility, and self-discipline.
  • Demonstrated project management experience.
  • Knowledge of the firm's reporting tools and processes.
  • Ability to analyze complex or unusual problems and deliver insightful, pragmatic solutions.
  • Ability to create, gather, and analyze data from varied sources.
  • A robust and resilient disposition that encourages disciplined team behaviors.
Responsibilities
  • Define and operationalize evaluation strategies for GenAI systems across Q&A assistants, summarization, extraction, drafting, agentic systems, and multi-step workflows.
  • Translate business use cases into structured evaluation plans covering scope, assumptions, success criteria, datasets, metrics, red-team scenarios, thresholds, and reporting requirements.
  • Create reusable evaluation templates, test case libraries, scoring rubrics, and reporting formats across product teams.
  • Design dataset requirements and ensure coverage of core user journeys, business intents, edge cases, adversarial cases, bias and fairness cases, sensitive demographic proxies, protected attributes, and stereotyping patterns.
  • Define guidance for dataset sufficiency and statistical coverage, including minimum samples, distribution balance, scenario matrices, and stratification by intent and risk.
  • Build reusable evaluation pipelines for answer quality, grounding and faithfulness, agentic behavior, and operational quality.
  • Combine LLM-as-judge and human evaluation using calibrated rubric design, sampling plans, and agreement checks.
  • Implement automated Python evaluation harnesses supporting batch scenario-suite runs, configurable metrics, reproducible run IDs and artifacts, and auditable trace and output storage.
  • Execute structured red teaming aligned with the OWASP Top 10 for LLM Applications, including prompt injection, tool hijacking, sensitive data disclosure, PII leakage, insecure output handling, training data leakage, memorization probes, model denial-of-service, and denial-of-wallet patterns.
  • Integrate evaluations into the development lifecycle through pre-release regression gates, CI checks, and benchmark comparisons across model versions, prompts, tools, and retrievers.
  • Perform adversarial testing of agentic workflows for tool misuse, over-permissioned access, unauthorized action execution, exfiltration through tools and connectors, and prompt injection through retrieved documents.
  • Recommend mitigations including input validation, retrieval filtering, tool sandboxing, least-privilege permissions, guardrails, policy prompting, refusal logic, output encoding, and monitoring alerts.
  • Produce auditable, decision-ready evaluation reports containing methodology, datasets, metrics, thresholds, quantitative results, qualitative results, risk assessments, and recommended control actions.
  • Present findings to stakeholders, explain residual risk and limitations, and provide rationale for go/no-go decisions.
  • Coordinate with product teams, solution architects, risk and compliance stakeholders, and assurance leadership to define evaluation requirements, obtain test datasets, execute evaluations, and recommend controls.
Desired Qualifications
  • Experience in assurance, finance, or regulatory environments, including model validation, risk acceptance workflows, or an audit-evidence mindset.
  • Familiarity with responsible AI frameworks such as the NIST AI Risk Management Framework, ISO/IEC 42001, and EU AI Act concepts.
  • Experience evaluating multilingual systems or domain-heavy enterprise assistants.
  • Hands-on experience with the Azure ecosystem, including Azure OpenAI, AI Search, Function Apps, App Insights, and Key Vault.
  • Security and red-teaming experience.

EY (Ernst & Young) is a global professional services firm that provides consulting, assurance, tax, and transaction advisory services across industries such as energy, healthcare, financial services, and real estate. Its work centers on helping clients solve critical business challenges by offering high-value advisory support in areas like supply chain, cybersecurity, sustainability, and digital transformation; revenue comes from fees for consulting, audit, and advisory services. What sets EY apart is its breadth of services across multiple disciplines, deep industry knowledge, and emphasis on thought leadership and research to inform clients’ strategic decisions. The firm aims to help organizations improve operational efficiency, navigate regulatory environments, and stay ahead of market trends by delivering practical, evidence-based guidance and execution support.

Company Size

10,001+

Company Stage

N/A

Total Funding

N/A

Headquarters

London, United Kingdom

Founded

1991

Get referred to Ernst & Young

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • April 7, 2026, agentic AI rollout targets all end-to-end audit activities by 2028.
  • February 10, 2026, Snowflake Innovation Center widens data-cloud transformation sales.
  • EY.ai Agentic for Sales launched March 2, 2026, with Snowflake and Canva.

What critics are saying

  • July 28, 2026, FRC sanctioned EY £1.197 million for Made.com audit failures.
  • July 2026 class action targets EY over tax-team breach exposing Social Security numbers.
  • Repeated sanctions break trust and drive clients from EY audits.

What makes Ernst & Young unique

  • EY Canvas and agentic AI power 160,000 audits globally.
  • Microsoft committed $1 billion with EY on May 21, 2026.
  • EY bundles tax, assurance, consulting, and EY-Parthenon inside one platform.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Professional Development Budget

Flexible Work Hours

Remote Work Options

Company News

Yahoo Finance
Aug 31st, 2026
EY commits $100M to bonuses rewarding human skills alongside AI adoption

Ernst & Young's US division is allocating $100 million this fiscal year to bonus payments rewarding employees who demonstrate adaptability, innovation, and judgement. Individual spot awards reach $500, whilst employees or teams making significant contributions can receive between $10,000 and $25,000, five times the previous programme's ceiling. The initiative also covers AI experimentation. "What we recognise signals what we value," said Ginnie Carlier, chief talent and culture officer for EY Americas. The bonuses form part of a broader strategy to reshape employee development across all career levels. Other professional services firms are pursuing similar approaches. KPMG restructured its audit internship this summer to emphasise critical thinking, whilst PwC US introduced training combining AI proficiency with human qualities like empathy. EY reported AI-related revenue grew 30% year-over-year in 2025.

Consultancy.eu
Aug 28th, 2026
EY-Parthenon acquires Dutch digital strategy consultancy SparkOptimus

EY-Parthenon has acquired SparkOptimus, a Dutch digital strategy consultancy founded in 2010. The Amsterdam-based firm employs around 50 consultants and specialises in AI transformation, digital business model redesign and technology-driven transformation. SparkOptimus founders Alexandra Jankovich and Tom Voskes, both former McKinsey consultants, said joining EY-Parthenon will create significant value for clients and staff whilst providing access to broader capabilities. The acquisition marks EY-Parthenon's first European deal since 2021. EY-Parthenon, established in 2014, is EY's strategy consulting and transactions advisory business with around 25,000 professionals globally. Mark Reich, Partner at EY-Parthenon Netherlands, said SparkOptimus' expertise will complement existing capabilities in the Dutch market. The deal closes on 1 September 2026. Financial terms were not disclosed.

Yahoo Finance
Aug 18th, 2026
KPMG and EY win $579M UK civil servants training contract despite consultancy spending pledge

The UK Government has awarded a contract worth up to £456 million to KPMG and EY to train civil servants, the Financial Times reported, citing government procurement tracker Tussell. Under the arrangement, the firms will train officials across various skills areas, including AI, between 2026 and 2028. KPMG's share is capped at £319 million, representing almost a quarter of its total UK advisory net sales from last year. EY's portion is worth £137 million, equivalent to around 13% of its UK consulting revenue. The deal is the largest single contract awarded to Big Four companies since Tussell started tracking records in 2012. The previous record was a £322 million deal between the Foreign Office and PricewaterhouseCoopers in 2012.

Business Insider
Jul 30th, 2026
EY's 'invisible' AI router cuts token costs by 60% by directing queries to cheaper models

EY has introduced an "invisible" AI router to manage internal AI spending, helping cut token consumption by up to 60% since its April rollout. The router sits behind specialised AI tools and directs employee queries to the most appropriate model for each task, rather than defaulting to the most powerful option. Token costs have become a growing concern as AI providers increasingly charge based on usage. EY's AI Pulse survey found that 82% of senior leaders at companies investing in AI were worried about token usage. The router has been deployed on department-specific platforms, including tax and risk functions, though not on the general Microsoft Copilot chatbot available to all staff. EY has also implemented token budgets based on employees' roles and departments. The firm's global consulting AI leader Dan Diasio said companies should focus AI investment on areas with the deepest impact rather than spreading resources thinly.

Yahoo Finance
Jul 29th, 2026
FRC fines EY $1.5M over Made.com audit failures before 2022 collapse

The UK's Financial Reporting Council has fined EY nearly £1.2m and audit partner Julie Carlyle £49,000 for failings in their 2021 audit of Made.com. Both received severe reprimands for breaching international auditing standards regarding going concern and deferred tax assets. The FRC said the auditors relied on management forecasts without sufficient challenge or adequate testing. Made.com, an online furniture retailer that listed on the London Stock Exchange in June 2021, entered administration in November 2022 after reporting a £35.3m loss. EY later disclaimed its opinion on Made.com's 2022 interim statements due to material uncertainty about the company's ability to continue operating.