Full-Time

Applied Data Scientist

Health AI Evaluation & Datasets

Innodata

Innodata

1,001-5,000 employees

Delivers AI data engineering services

Compensation Overview

$210k - $240k/yr

Remote in Canada

Remote

Bachelor's

Category
Data & Analytics (1)
Required Skills
Scikit-learn
Python
SQL
RAG
Pandas

Get referred to Innodata

See people who can refer or advise you

Requirements
  • 5+ years of data science experience, including at least 2+ years with healthcare, clinical, biomedical, payer, provider, pharma, life sciences, or comparable regulated health data
  • Working knowledge of healthcare data and standards: EHR structure, clinical documentation conventions, ICD-10, CPT, SNOMED Clinical Terms, LOINC, RxNorm, and at least passing familiarity with FHIR, HL7, or equivalent interoperability concepts
  • Hands-on experience designing ML datasets, not just consuming them: writing annotation guidelines, sizing cohorts, setting quality thresholds, designing QA checks, and shipping data that downstream teams can train or evaluate on
  • Familiarity with LLM-based health AI workflows, including prompt design, rubric-based evaluation, retrieval-augmented generation, LLM-as-judge methods, model comparison, and the limitations of automated evaluation in clinical contexts
  • Strong Python and SQL; comfort with pandas, scikit-learn, statsmodels or equivalent tools; working familiarity with modern LLM tooling such as Hugging Face, evaluation frameworks, prompt development tools, or model APIs
  • Statistical literacy across sampling design, bias and fairness analysis, inter-annotator agreement metrics (Cohen or Fleiss kappa, Krippendorff alpha), confidence intervals, significance testing where appropriate, error analysis
  • Solid grasp of healthcare privacy, compliance, and governance: HIPAA, de-identification standards (Safe Harbor and Expert Determination), practical mechanics of working with PHI safely, auditability, access control, and documentation fit for high-stakes or regulated AI programs
  • Ability to work credibly with clinicians, biomedical SMEs, research scientists, engineers, technical solutions teams, annotators, and customer stakeholders
  • A bias toward clinical realism: you would rather build a smaller dataset that reflects what clinicians, reviewers, patients, or care teams actually see than a larger dataset that looks impressive on paper but fails in practice
  • Degree in a relevant field such as biostatistics, epidemiology, computational biology, health informatics, computer science with a health focus, statistics, a clinical degree with quantitative training, or equivalent demonstrated experience
  • Clinical credentials are not required, but candidates must be able to work credibly with clinicians, biomedical SMEs, and health AI customers; candidates with MD, RN, PharmD, MPH, PhD, or health informatics backgrounds are especially encouraged
Responsibilities
  • Translate customer goals — such as improving differential diagnosis, evaluating a clinical note summarizer, testing a RAG-based medical literature assistant, or creating preference data for patient-facing chatbots — into dataset specifications, taxonomies, rubrics, sampling plans, and acceptance criteria
  • Make multimodal health AI a core focus: design training and evaluation datasets across clinical text, medical images, waveforms, structured EHR data, claims, trial data, medical literature, patient communications, payer policies, drug information, and other clinical artifacts, as well as use cases such as clinical reasoning, medical QA, note summarization, medical coding, patient communication, utilization management, and literature synthesis
  • Design evaluations for retrieval-augmented and source-grounded health AI systems, including evidence citation, faithfulness, contraindication handling, guideline adherence, source freshness, and failure modes caused by incomplete, conflicting, or stale context
  • Define sampling strategies, label schemas, inter-annotator agreement targets, adjudication workflows, SME review patterns, and quality thresholds in partnership with Language Data Scientists, clinicians, biomedical experts, and quality teams
  • Build statistical and ML checks that make healthcare datasets trustworthy: stratified sampling across specialties and patient subgroups, bias and representation analysis, leakage detection, distribution shift checks, uncertainty estimates, reliability metrics, and subgroup performance analysis
  • Partner with Applied Research Scientists and AI/ML Research Engineers to instrument datasets into evaluation and post-training pipelines, including rubric-grounded LLM-as-judge prompts, regression suites, model comparison workflows, experiment tracking, and model-improvement feedback loops
  • Evaluate health AI behavior beyond surface accuracy: calibration, hallucination on safety-critical content, refusal appropriateness, robustness under ambiguity, equity across patient subgroups, and safe handoff in agentic or workflow-integrated systems. Reason concretely about clinical workflow fit: where outputs enter care delivery, what evidence a clinician or reviewer would need to trust them, when uncertainty must be surfaced, and how patient-facing, clinician-facing, payer, pharma, and operational use cases differ in risk
  • Own data quality from source intake through delivery, including de-identified clinical text, medical literature, synthetic cases, structured records, client policies, and knowledge bases, with attention to PHI/PII handling, provenance, audit trails, versioning, and compliance documentation
  • Stay current on the health AI landscape — regulatory developments such as FDA guidance on AI/ML-enabled medical devices and EU AI Act health provisions, benchmark releases such as MedQA, MedMCQA, and HealthBench, and emerging clinical evaluation methodology
  • Support customer discovery and proposal work by scoping dataset programs, sizing annotation and SME review effort, identifying regulatory or data-access constraints, and explaining methodology choices to client clinical and ML leadership
  • Contribute to Innodata internal IP: reusable health-domain taxonomies, evaluation rubrics, golden datasets, clinical review playbooks, dataset quality checks, and methodology templates

Innodata is a global data engineering company that provides AI-enabled software platforms and managed services to create high-quality training data and data pipelines for AI. It combines proprietary software with a global network of over 5,000 subject-matter experts to collect, create, annotate, and validate data, and it also offers synthetic data generation and end-to-end AI lifecycle services. The DDS segment drives most revenue, with Synodex and Agility supporting healthcare data and media monitoring. Its goal is to deliver reliable, safe, and effective datasets and AI lifecycle solutions across industries including technology, finance, insurance, and government.

Company Size

1,001-5,000

Company Stage

IPO

Headquarters

Hackensack, New Jersey

Founded

1988

Get referred to Innodata

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Q2 2026 revenue jumped 58% to $92.1 million, with 49% adjusted gross margin.
  • Cash reached $250.4 million in June 2026, funding hiring, product development, and selective acquisitions.
  • Management raised 2026 revenue growth guidance to 40% plus, while adding a frontier-lab customer.

What critics are saying

  • A single customer drove 37% of Q2 2026 revenue; hyperscalers can insource fast.
  • The $300 million at-the-market program threatens 2026 dilution and signals management expects needs ahead.
  • April 27, 2026 securities dismissal did not erase Philippine judgment exposure and recurring litigation overhang.

What makes Innodata unique

  • August 2026 Cyber Training Suite trains coding agents on 12 reconstructed vulnerability datasets.
  • Innodata sold secure-code evaluation, agentic reinforcement learning, and public benchmarks, not generic labeling.
  • Rahul Singhal's September 30, 2026 promotion signals a stronger operator-led growth machine.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Remote Work Options

Company News

Yahoo Finance
Aug 8th, 2026
Innodata launches AI Cyber Training Suite for software vulnerability detection

Innodata has launched its AI Cyber Training Suite, a cybersecurity product targeting AI coding agents and enterprise software security. The suite introduces a methodology for training and evaluating AI agents on software vulnerability detection and patching. The product addresses concerns about trust, code quality, and security when deploying AI agents in large software environments. It targets enterprises relying on AI-generated code and legacy system modernization. The stock closed at $62.33, up 17.6% year to date and 43.2% over the past year, though down 9.7% over the past month. The launch reinforces Innodata's strategy of building proprietary tools that integrate into model builders' workflows. However, the company faces challenges from concentrated client reliance and competition from larger security vendors like CrowdStrike.

Yahoo Finance
Aug 7th, 2026
Innodata grows revenue 58% to $92.1M, reduces top client reliance to 37%

Innodata reported second-quarter 2026 revenues of $92.1 million, up 58% year-over-year, beating the Zacks Consensus Estimate of $86.3 million. Earnings of $0.41 per share exceeded the $0.21 consensus estimate. The company achieved a 49% adjusted gross margin. Chairman and CEO Jack Abuhoff said the largest customer represented 37% of quarterly revenues, down from 56% in the first quarter, whilst a Big Tech customer rose to 34% from 17%. The company added a new frontier-lab customer. President Rahul Singhal highlighted new programmes spanning agentic AI, model evaluation, cybersecurity, and physical AI. Innodata is developing motion-capture capabilities and released two public benchmarks. The company reiterated full-year 2026 revenue growth guidance of 40% or more.

CoinCentral
Aug 7th, 2026
Innodata stock falls 5% as $300M equity program raises dilution fears despite 58% revenue surge

Innodata shares fell 5% on Friday following the announcement of a $300 million at-the-market equity programme, sparking dilution concerns amongst investors. The decline overshadowed strong second-quarter results that beat analyst expectations. The AI data and digital services company reported Q2 revenue of $92.1 million, up 58% year-on-year. Diluted earnings per share rose to $0.41 from $0.20, whilst adjusted EBITDA climbed 92% to $25.4 million. The equity programme could result in approximately 4.15 million new shares at recent prices, representing a potential 12% increase to the existing share count of 34.4 million. Management reaffirmed guidance for more than 40% revenue growth in 2026, implying annual revenue of roughly $352 million.

Associated Press
Aug 6th, 2026
Innodata posts $92.1M revenue in Q2, up 58% YoY, beats consensus by 7%; CEO transition announced

Innodata reported record second quarter 2026 results with revenue of $92.1 million, up 58% year-over-year and beating consensus by 7%. Adjusted EBITDA reached $25.4 million, exceeding consensus by 50%. The company posted net income of $14.4 million, or $0.41 per diluted share, compared to $7.2 million in the prior year period. Adjusted gross margin expanded to 49%, nine points above the company's 40% target. Cash and short-term investments totalled $250.4 million at quarter end. Innodata announced a leadership transition effective 30 September 2026. Rahul Singhal will become president and chief executive officer, whilst founder Jack Abuhoff will transition to executive chairman. The company reiterated full-year 2026 revenue growth guidance of 40% or more year-over-year.

Yahoo Finance
Aug 4th, 2026
Innodata set to report Q2 results: Analysts expect 48% revenue growth to $86M

Innodata is scheduled to release second-quarter 2026 results on 6 August after market close. The Zacks Consensus Estimate for quarterly earnings per share stands at 21 cents, indicating 5% year-over-year growth, whilst revenue is expected to reach $86.3 million, up 47.8% from the prior year. The company has beaten earnings estimates in each of the past four quarters, with an average surprise of 98.9%. In the first quarter, adjusted earnings and revenues exceeded estimates by 223.1% and 17.8%, respectively, whilst growing 90.9% and 54.4% year over year. Zacks' model predicts a likely earnings beat, as Innodata carries an Earnings ESP of +17.65% and a Zacks Rank #1. The company raised its full-year 2026 revenue growth outlook to approximately 40% or more during the first-quarter earnings call.