Full-Time

Model Evaluations

Member of Technical Staff

Updated on 9/13/2026

Simile

Simile

51-200 employees

Enterprise AI platform simulating human behavior

Compensation Overview

$200k - $400k/yr

+ Equity

San Francisco, CA, USA + 1 more

More locations: New York, NY, USA

In Person

Category
AI & Machine Learning (1)
Required Skills
LLM
Python
Data Visualization
R
SQL
A/B Testing

Get referred to Simile

See people who can refer or advise you

Requirements
  • Strong intuition for what makes an evaluation meaningful, robust, and decision-relevant, including the ability to explain what an evaluation measures, what it does not measure, how it can be gamed, and why it should or should not affect a model or product decision.
  • Understanding of modern large language model training, post-training, model evaluation, and hill-climbing.
  • Ability to reason about noisy data, uncertainty, sampling, distributions, calibration, confidence intervals, measurement validity, bias, variance, and the difference between an observed result and the underlying population quantity it estimates.
  • Ability to build internal tools, scripts, dashboards, labeling workflows, analyses, or automated evaluation pipelines quickly using Python, SQL, R, notebooks, large language model application programming interfaces, and agentic coding tools.
  • Ability to independently drive a workstream while doing the hands-on work, including building initial versions, inspecting data, debugging workflows, writing rubrics, and revising metrics.
Responsibilities
  • Design evaluations, metrics, rubrics, datasets, dashboards, and workflows that measure whether models accurately predict human behavior across customer use cases, populations, question types, and decision contexts.
  • Evaluate new model versions, diagnose regressions, identify priority areas for model-improvement cycles, and maintain stable evaluation suites representing capabilities customers care about.
  • Build evaluations for qualitative responses, retrieval, survey generation, artificial-intelligence-generated research reports, and other product surfaces where model quality affects customer trust.
  • Turn subjective quality concerns into concrete rubrics, labeled data, automated graders, release criteria, and model-improvement signals.
  • Develop rigorous ways to compare simulated responses against human data, customer studies, company-collected ground truth, and behavioral datasets.
  • Help the company reason about sampling error, uncertainty, calibration, margin of error, representativeness, and the meaning of ground truth when human behavior is noisy.
  • Use modern agentic coding tools to build internal tools, inspect model outputs, create labeling workflows, validate evaluations, and turn ambiguous evaluation questions into working systems.
  • Prototype ways to evaluate behavioral predictions using transaction or purchase behavior, product interactions, intervention response, first-party experiments, and multi-agent group settings.
Desired Qualifications
  • Experience building model-evaluation dashboards, regression suites, release gates, benchmark sets, model-comparison workflows, or systems that help machine-learning teams prioritize work and decide when to ship.
  • Experience designing rubrics, automated graders, pairwise comparisons, expert-review workflows, labeling interfaces, grader calibration, or human/model hybrid evaluation systems.
  • Experience with sampling, weighting, margin of error, power analysis, uncertainty quantification, Bayesian modeling, causal inference, psychometrics, polling, or measurement theory.
  • Experience evaluating behavioral predictions using transaction data, purchase behavior, mobility data, product interactions, or other passively collected behavioral signals.
  • Experience designing randomized controlled trials, A/B tests, survey experiments, vignette studies, field experiments, behavioral games, or intervention studies.
  • Interest or experience in modeling group conversation, deliberation, focus groups, juries, committees, polarization, collective decision-making, or social influence.
  • Experience in large language model evaluations, applied machine-learning research, data science, research engineering, human data, market research, user-experience research, polling, behavioral science, computational social science, or behavioral economics.
  • Being a recent graduate or self-directed builder with unusually strong evaluation judgment, statistical ability, and artificial-intelligence tool fluency.

Simile builds AI models that simulate human behavior to forecast how individuals and groups will respond to different scenarios, using generative agents grounded in interviews and behavioral data. Its foundation-model approach powers these agents to mimic real decision-making for tasks like testing product launches, policy changes, or corporate announcements. The platform targets enterprises seeking decision intelligence, enabling rehearsals of earnings calls, litigation outcomes, and market strategies. By grounding agents in real data and behavioral science, Simile differentiates from general AI tools and focuses on helping organizations test decisions to reduce uncertainty before acting.

Company Size

51-200

Company Stage

Series B

Total Funding

$300M

Headquarters

Palo Alto, California

Founded

2025

Get referred to Simile

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Simile raised over $200 million in July 2026, accelerating product and hiring.
  • Customers include CVS, Telstra, Deloitte, Wealthfront, and Gallup, spanning healthcare, telecom, finance.
  • The company says revenue quintupled since February 2026 and tens of millions of simulations ran.

What critics are saying

  • Simile discloses no revenue, retention, or contract values, so valuation discipline stays weak.
  • Synthetic-user outputs face trust risk if earnings-call, litigation, or pricing forecasts miss reality.
  • If regulators reject simulated human data or customers revoke trust, Simile's core product collapses.

What makes Simile unique

  • Stanford founders Joon Sung Park, Percy Liang, and Michael Bernstein anchor unusual research credibility.
  • Simile sells human-behavior simulations, not generic copilots, targeting enterprise decisions before execution.
  • Its $2 billion Series B and CVS Health Ventures backing signal rare strategic validation.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

-17%

1 year growth

-17%

2 year growth

-17%
IntelPro
Sep 10th, 2026
AI research startup Listen Labs scrubbed a $1.5B funding round for Salesforce talks.

AI research startup Listen Labs scrubbed a $1.5B funding round for Salesforce talks. Listen Labs walked away from a signed Series C term sheet from Menlo Ventures, sources say. Listen Labs, a market research startup that uses voice AI to conduct customer interviews, recently signed a term sheet for a $125 million Series C at a $1.5 billion valuation, with Menlo Ventures set to lead the round, according to several people with knowledge of the matter. But that round never closed, the people said. Listen Labs walked away from the signed term sheet, a rare occurrence in the venture world and one that is generally frowned upon, according to VCs. The financing likely collapsed because of acquisition talks with Salesforce. The CRM giant has recently held talks to buy Listen Labs for around $2 billion, Business Insider reported. The discussions are not finalized, however, and may not result in a deal, the outlet notes. Listen Labs is one of the leading startups in the rapidly growing field of automating customer research with AI. The three-year-old startup has about $30 million in annualized revenue, about three times more than Simile, a competing startup that predicts human behavior, according to two people familiar with the companies' financials. In late July, Simile announced that it had closed a $200 million Series B at a $2 billion valuation led by Greenoaks - likely setting a new valuation benchmark for Listen Labs, one person said. If talks with Salesforce collapse, several VCs told TechCrunch that they expect Listen Labs to return to market and target a valuation of $2 billion or higher. While acquiring Listen Labs could strengthen Salesforce's AI capabilities by using the startup's AI to help predict customer needs, the CRM giant may ultimately decide that paying a 67 times revenue multiple is too steep a valuation, according to a person with experience negotiating exits to Salesforce. Listen Labs, Salesforce, Menlo Ventures, and Simile did not immediately respond to requests for comment. Listen Labs was co-founded in 2023 by Florian Jüngermann, a former German national champion in competitive computer programming, and Alfred Wahlforss, who previously founded a staffing startup called Bemlo. The two met while pursuing master's degrees at Harvard. Listen Labs' AI develops survey questions and interviews customers over audio or video. The resulting conversations are then packaged into reports and PowerPoint presentations, similar to those traditionally produced by human market researchers. Fortune 500 companies rely on this type of research to gauge customer needs and satisfaction with their brands and products, but traditional market research is expensive and can take weeks to complete. Listen Labs' technology helps reduce the time and cost of these projects, enabling companies to quickly understand how customers are reacting to product changes and iterate on them more efficiently. The startup's customers include Microsoft, Canva, Anthropic, and Sweetgreen. Listen Labs and Simile aren't the only startups using AI to disrupt the customer research market. Besides Simile, competitors in the space include Outset, Keplar, and Aaru. While some platforms automate interviews with real humans, others startups - like Aaru and Simile - take a synthetic approach, using AI to simulate human behavior and predict responses without interviewing anyone at all. Listen Labs was previously valued at $500 million when it announced a $69 million Series B round in late January led by Ribbit Capital, with participation from returning backers Sequoia, Conviction, and Pear VC.

MLQ AI
Jul 31st, 2026
Simile raises more than $200 million at a $2 billion post-money valuation.

Simile raises more than $200 million at a $2 billion post-money valuation. Key points * Greenoaks led the Series B, with Index Ventures, Hanabi, Bain Capital Ventures, A*, Factory, CVS Health Ventures and new investor Definition participating.[[1]] * The $2 billion figure is a post-money valuation; Simile did not disclose its pre-money valuation, share price or whether the round included secondary sales.[[1]] * Simile says revenue has increased fivefold since its February launch, but it provided no revenue figure, customer count or contract values.[[1]] Simile has raised more than $200 million in a Series B led by Greenoaks, giving the artificial-intelligence startup a $2 billion post-money valuation just five months after its public launch. Returning investors Index Ventures, Hanabi, Bain Capital Ventures, A*, Factory and CVS Health Ventures participated, while Definition joined as a new investor.[[1]] The Palo Alto company builds AI models intended to simulate how people and groups respond to products, marketing, pricing and other decisions. Simile says customers include CVS Health, Wealthfront, Deloitte and Gallup, and that Fortune 100 enterprises have run tens of millions of simulations through its technology.[[1]] Those product and usage figures were reported by the company and have not been independently audited. Rapid growth, limited financial disclosure. Simile said revenue has grown fivefold since it launched in February and that its workforce now exceeds 50 people. It did not provide starting or current revenue, annual recurring revenue, customer totals, pricing or retention figures, making the valuation difficult to assess against conventional software metrics.[[1]] The company also did not explain whether the $2 billion valuation resulted solely from the price paid for newly issued shares or whether the financing contained secondary transactions. The company says its models begin with data from real people and are checked against human responses. Its website claims more than 7,000 evaluations across demographic groups and enterprise use cases, along with a confidence model designed to estimate the accuracy of each simulation.[[2]] Simile has not released enough underlying commercial validation data for outsiders to test those claims broadly. Simile said the proceeds would accelerate its work and support hiring researchers, engineers, designers and operators, without providing a spending breakdown.[[1]] Index Ventures led its $100 million Series A announced on February 12, alongside Bain Capital Ventures, A* and Hanabi Capital and several individual AI investors.[[3]] TechCrunch first reported the new financing as a $200 million Series B.[[4]] Discover more Computer Science Market intelligence reports AI market analysis Discover more AI infrastructure solutions Dictionaries & Encyclopedias Web Apps & Online Tools At the intersection of AI, tech, and markets. The stories that matter, in one email. Free - unsubscribe anytime.

Simile
Jul 30th, 2026
Announcing Our Series B

Simile has raised over $200 million at a $2 billion post-money valuation, led by Greenoaks, to accelerate our mission: simulating all eight billion people.

Tech Funding News
Feb 13th, 2026
$100M for Stanford spinout Simile: AI that simulates human decisions — TFN

Artificial intelligence startup Simile has secured $100 million in fresh funding led by Index Ventures, included participation from others.

MR Web
Feb 13th, 2026
Daily Research News Online

Daily research News online. The global MR industry's daily paper since 2000. Follow DRNO on... Prediction tech company Simile launches with $100m. February 13 2026 AI firm Simile has secured $100m in funding, to develop its prediction technology, whose applications range from forecasting consumer behavior to suggesting likely questions from analysts during company earnings calls. Simile says it is building 'a foundation model that predicts human behavior in any situation, and a product that deploys it at scale.' According to Bloomberg News, the company emerged from stealth this week after training its model on a combination of consumer interviews, transaction data and behavioral experiment findings from scientific journals. For the first of these, it taps the probability-based, nationally representative panel of partner firm Gallup. The firm was founded by Stanford researchers Joon Park (CEO, pictured), Percy Liang and Michael Bernstein, and is backed by researcher Fei-Fei Li, co-director of Stanford's Human-Centered A.I. Institute; and OpenAI co-founder Andrej Karpathy. The $100m round was led by Index Ventures, with participation from Hanabi, A* and Bain Capital Ventures. Web site: www.simile.ai. All articles 2006-23 written and edited by Mel Crowther and/or Nick Thomas, 2024- by Nick Thomas, unless otherwise stated. Most viewed items in the last week... Each (*) indicates > 1,000 views. Select a region below...