Halluminate

Halluminate

AI agent evaluation with sandbox environments

Overview

Halluminate builds evaluation infrastructure for AI agents focusing on data, environments, and benchmarking to improve computer and browser-use agents. It offers a platform with fully managed, parallel sandbox environments that mimic enterprise software like Salesforce for safe, scalable testing. It also provides an evaluation service that combines proprietary datasets with expert annotations to identify and prioritize AI agent failure modes. The company targets AI engineers at Series A+ startups and enterprise ML teams, backed by Y Combinator and Alchemist Accelerator, aiming to reduce unreliability and accelerate safe real-world deployment.

YC Company

About Halluminate

Simplify's Rating
Why Halluminate is rated
C+
Rated C on Competitive Edge
Rated B on Growth Potential
Rated C on Differentiation

Industries

Data & Analytics

Enterprise Software

AI & Machine Learning

Financial Services

Company Size

11-50

Company Stage

Seed

Total Funding

$130K

Headquarters

San Francisco, California

Founded

2024

Get referred to Halluminate

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Westworld launched August 20, 2026, giving Halluminate fresh relevance inside frontier-model training loops.
  • WebBench’s 2,454 live-site tasks and 2025 MIT release built visible distribution and credibility.
  • Deals with finance, consulting, and model labs monetize proprietary operational data at six-figure licensing prices.

What critics are saying

  • Open benchmarks like WebBench invite commoditization; labs can replicate Halluminate’s moat quickly.
  • Westworld’s 51% best score on August 13, 2026 exposes weak agent performance and limited product readiness.
  • Anthropic, OpenAI, or browser-agent vendors can internalize similar eval stacks, collapsing Halluminate’s standalone demand.

What makes Halluminate unique

  • Halluminate pairs realistic sandboxes with finance-specific benchmarks like Westworld, not generic evals.
  • It tests full due-diligence workflows using real private-deal data and hundreds-step trajectories.
  • The company sells both the Halluminate harness and benchmark datasets, tightening workflow control.

Help us improve and share your feedback! Did you find this helpful?

Funding

Total Funding

$130k

Below

Industry Average

Funded Over

1 Rounds

Notable Investors:
Seed funding is usually the first official round after pre-seed, when a startup has a prototype or concept. It’s typically used to develop the product, test the market, and start building the team. Investors here are often angel investors or early-stage venture capitalists.
Seed Funding Comparison
Below Average

Industry standards

$3.3M
$130k
Halluminate
$1.5M
Slack
$2M
Netflix
$2.3M
Instacart
$3M
Robinhood

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Relocation Assistance

401(k) Company Match

Meal Benefits

Gym Membership

Commuter Benefits

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%
Halluminate
Aug 11th, 2026
Westworld Finance Diligence Bench.

Westworld Finance Diligence Bench. Evaluating AI Agents on a Complete company acquisition due-diligence process Halluminate, Inc. present Westworld Finance Diligence Bench, a benchmark of 88 problems that evaluates AI agents on a complete company acquisition due-diligence process. It draws on anonymized data from real private transactions, runs inside dynamic desktop environments, and spans trajectories that reach hundreds of steps. Every problem is written and reviewed by practicing finance deal professionals and verified by a mixed pipeline of agentic, discrete, and binary verifiers. Halluminate, Inc. evaluate a range of frontier models, each across its native harness and its internal Halluminate harness, and present every model-harness pair three ways. The histogram shows each configuration's mean score with average cost per run inside the bar. The Pareto 2D plot trades that cost against mean score and the summary table adds average tool usage, cost, mean score, and pass rate. Mean score is the per-problem pass@1 verifier grade averaged over the 88 problems, always between 0 and 1. Pass rate is the share of those problems whose single graded run clears the 0.50 bar, meaning it scores at least 0.50; the 0.50 threshold is a reporting choice and does not make the underlying reward binary, as Halluminate, Inc. explain below. Cost is estimated from each run's token usage and the model's published input/output/cached pricing, amortized across the 88 tasks. Key innovations of its benchmark. * 1Real private-deal data: anonymized documents from actual private equity transactions. To the best of its knowledge, no published benchmark evaluates a complete diligence process on private deal data of this breadth. * 2A complete process: tasks that together span one deal's diligence, from early analysis through to close, rather than isolated tasks tied to a single job title. * 3Mixed verifier pipeline: each problem is scored by a decomposed mix of agentic rubric grading and deterministic checks, weighted by expert authors and audited in an independent QA pass. * 4Computer use and tool use together: a single task requires operating a real desktop environment, with files, office applications, and a data room, while also calling structured tools for email and chat, all in one trajectory. A few benchmarks combine these two modes, but it remains rare, and rarer still on finance work. The rest of this post walks through the benchmark in detail. Halluminate, Inc. first place the benchmark in the context of related work. Then, walk through a sample of a single task, following it from the prompt through the tools, verifiers, and graders. Halluminate, Inc. then describe the quality process Halluminate, Inc. use to keep environments healthy. Next, Halluminate, Inc. present detailed results from its runs, including a per-trajectory error analysis. Halluminate, Inc. close with its takeaways and future work. 1. Previous works. Benchmarks for professional domains have grown narrow by construction: accounting tasks in one suite, investment banking tasks in another, each targeted at a single task type or a single job title. For instance, some of these are: * FinanceBench - 10,231 open-book question-answer pairs about public companies, testing whether a model can pull the right figure from a filing. * FinQA - expert-written numerical-reasoning questions over single earnings reports, each paired with an executable reasoning program. * ConvFinQA - a conversational extension of FinQA that chains multi-step numerical questions across the turns of a dialogue. * BizBench - eight quantitative-reasoning tasks that grade financial problem solving through program synthesis over structured data. * FinanceQA - hedge-fund, private-equity, and investment-banking analysis questions covering hand-spreading, valuation conventions, and reasoning under incomplete information. However, this decomposition no longer matches how AI systems enter these domains. Organizations now deploy agents against whole processes. This complete-process way of using AI has no frontier-quality benchmark, only piecewise tests that cannot be combined into a full picture. Some works that have attempted to address this are: * DealTrace - its own private-equity deal review across ten real deals and five sequential stages (extract, reconcile, forecast, market, recommend); a first step toward process-level evaluation, but narrower than a full diligence suite. * Finance Agent Benchmark - 537 expert questions across nine categories answered by an agent with web-search and EDGAR access; agentic, but still discrete research questions over public filings rather than one connected deal. * TheAgentCompany - long-horizon agent tasks inside a simulated software company that combine computer use (a real browser and desktop applications) with tool use (code execution, chat, and internal services) in the same task; the closest general analogue to whole-process work precisely because it exercises both modes together, though with almost no finance content. * OSWorld - 369 open-ended computer-use tasks across real operating systems and applications, establishing the stateful multi-app desktop setting but without any deal or finance domain. * tau-bench - multi-turn tool-agent-user interactions graded against database state and policy rules, testing sustained tool use, but on customer support rather than analyst deal work. Westworld Finance Diligence Bench targets this gap, simulating the full body of work that a team of investment bankers, consultants, and accountants would produce together over several weeks. Unlike pure computer-use benchmarks such as OSWorld, its agents operate in the environment through code, structured tools, and mouse-driven GUI control. Trajectories regularly run to hundreds of steps, and the longest model and harness pairings average close to nine hundred iterations per run, as the trajectory analysis below shows. 2. Coverage. The benchmark spans the full arc of a deal, from the first model to the signed close.

Teten
Jul 6th, 2026
How investors & consultancies can monetize their IP.

How investors & consultancies can monetize their IP. Halluminate.ai, an Orange Collective portfolio company, works with Anthropic/OpenAI/other leading AI labs to build RL (Reinforcement Learning) training environments that teach models how to do work in financial services, consulting, and general knowledge work. They are looking to license the internal operational data of financial institutions, consulting firms, or similar knowledge work institutions, e.g., old deal rooms, client projects, etc. They can pay up to 6 or 7 digits. They also will discuss partnering in a forward deployed project capacity for custom advisory or product implementation, and as part of that tradeoff getting access to real world environments and workflow understanding Any data you license will be fully anonymized and legally cleared before use. In addition, they can offer AI advisory/consulting services for interested parties as part of this exchange. If you'd like to learn more, contact Teten Advisors, LLC.

Skyvern
May 29th, 2025
Web Bench - A new way to compare AI Browser Agents

As a result, Skyvern partnered with Halluminate and created a new benchmark to better quantify these failures.

Recently Posted Jobs

Sign up to get curated job recommendations

Halluminate is Hiring for 8 Jobs on Simplify!

Find jobs on Simplify and start your career today

Don't see your dream role? Check out thousands of other roles on Simplify. Browse all jobs →