
Work Here?
Halluminate builds evaluation infrastructure for AI agents focusing on data, environments, and benchmarking to improve computer and browser-use agents. It offers a platform with fully managed, parallel sandbox environments that mimic enterprise software like Salesforce for safe, scalable testing. It also provides an evaluation service that combines proprietary datasets with expert annotations to identify and prioritize AI agent failure modes. The company targets AI engineers at Series A+ startups and enterprise ML teams, backed by Y Combinator and Alchemist Accelerator, aiming to reduce unreliability and accelerate safe real-world deployment.
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Financial Services
Company Size
11-50
Company Stage
Seed
Total Funding
$130K
Headquarters
San Francisco, California
Founded
2024
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$130k
Below
Industry Average
Funded Over
1 Rounds
Industry standards
Health Insurance
Dental Insurance
Vision Insurance
Relocation Assistance
401(k) Company Match
Meal Benefits
Gym Membership
Commuter Benefits
Westworld Finance Diligence Bench. Evaluating AI Agents on a Complete company acquisition due-diligence process Halluminate, Inc. present Westworld Finance Diligence Bench, a benchmark of 88 problems that evaluates AI agents on a complete company acquisition due-diligence process. It draws on anonymized data from real private transactions, runs inside dynamic desktop environments, and spans trajectories that reach hundreds of steps. Every problem is written and reviewed by practicing finance deal professionals and verified by a mixed pipeline of agentic, discrete, and binary verifiers. Halluminate, Inc. evaluate a range of frontier models, each across its native harness and its internal Halluminate harness, and present every model-harness pair three ways. The histogram shows each configuration's mean score with average cost per run inside the bar. The Pareto 2D plot trades that cost against mean score and the summary table adds average tool usage, cost, mean score, and pass rate. Mean score is the per-problem pass@1 verifier grade averaged over the 88 problems, always between 0 and 1. Pass rate is the share of those problems whose single graded run clears the 0.50 bar, meaning it scores at least 0.50; the 0.50 threshold is a reporting choice and does not make the underlying reward binary, as Halluminate, Inc. explain below. Cost is estimated from each run's token usage and the model's published input/output/cached pricing, amortized across the 88 tasks. Key innovations of its benchmark. * 1Real private-deal data: anonymized documents from actual private equity transactions. To the best of its knowledge, no published benchmark evaluates a complete diligence process on private deal data of this breadth. * 2A complete process: tasks that together span one deal's diligence, from early analysis through to close, rather than isolated tasks tied to a single job title. * 3Mixed verifier pipeline: each problem is scored by a decomposed mix of agentic rubric grading and deterministic checks, weighted by expert authors and audited in an independent QA pass. * 4Computer use and tool use together: a single task requires operating a real desktop environment, with files, office applications, and a data room, while also calling structured tools for email and chat, all in one trajectory. A few benchmarks combine these two modes, but it remains rare, and rarer still on finance work. The rest of this post walks through the benchmark in detail. Halluminate, Inc. first place the benchmark in the context of related work. Then, walk through a sample of a single task, following it from the prompt through the tools, verifiers, and graders. Halluminate, Inc. then describe the quality process Halluminate, Inc. use to keep environments healthy. Next, Halluminate, Inc. present detailed results from its runs, including a per-trajectory error analysis. Halluminate, Inc. close with its takeaways and future work. 1. Previous works. Benchmarks for professional domains have grown narrow by construction: accounting tasks in one suite, investment banking tasks in another, each targeted at a single task type or a single job title. For instance, some of these are: * FinanceBench - 10,231 open-book question-answer pairs about public companies, testing whether a model can pull the right figure from a filing. * FinQA - expert-written numerical-reasoning questions over single earnings reports, each paired with an executable reasoning program. * ConvFinQA - a conversational extension of FinQA that chains multi-step numerical questions across the turns of a dialogue. * BizBench - eight quantitative-reasoning tasks that grade financial problem solving through program synthesis over structured data. * FinanceQA - hedge-fund, private-equity, and investment-banking analysis questions covering hand-spreading, valuation conventions, and reasoning under incomplete information. However, this decomposition no longer matches how AI systems enter these domains. Organizations now deploy agents against whole processes. This complete-process way of using AI has no frontier-quality benchmark, only piecewise tests that cannot be combined into a full picture. Some works that have attempted to address this are: * DealTrace - its own private-equity deal review across ten real deals and five sequential stages (extract, reconcile, forecast, market, recommend); a first step toward process-level evaluation, but narrower than a full diligence suite. * Finance Agent Benchmark - 537 expert questions across nine categories answered by an agent with web-search and EDGAR access; agentic, but still discrete research questions over public filings rather than one connected deal. * TheAgentCompany - long-horizon agent tasks inside a simulated software company that combine computer use (a real browser and desktop applications) with tool use (code execution, chat, and internal services) in the same task; the closest general analogue to whole-process work precisely because it exercises both modes together, though with almost no finance content. * OSWorld - 369 open-ended computer-use tasks across real operating systems and applications, establishing the stateful multi-app desktop setting but without any deal or finance domain. * tau-bench - multi-turn tool-agent-user interactions graded against database state and policy rules, testing sustained tool use, but on customer support rather than analyst deal work. Westworld Finance Diligence Bench targets this gap, simulating the full body of work that a team of investment bankers, consultants, and accountants would produce together over several weeks. Unlike pure computer-use benchmarks such as OSWorld, its agents operate in the environment through code, structured tools, and mouse-driven GUI control. Trajectories regularly run to hundreds of steps, and the longest model and harness pairings average close to nine hundred iterations per run, as the trajectory analysis below shows. 2. Coverage. The benchmark spans the full arc of a deal, from the first model to the signed close.
How investors & consultancies can monetize their IP. Halluminate.ai, an Orange Collective portfolio company, works with Anthropic/OpenAI/other leading AI labs to build RL (Reinforcement Learning) training environments that teach models how to do work in financial services, consulting, and general knowledge work. They are looking to license the internal operational data of financial institutions, consulting firms, or similar knowledge work institutions, e.g., old deal rooms, client projects, etc. They can pay up to 6 or 7 digits. They also will discuss partnering in a forward deployed project capacity for custom advisory or product implementation, and as part of that tradeoff getting access to real world environments and workflow understanding Any data you license will be fully anonymized and legally cleared before use. In addition, they can offer AI advisory/consulting services for interested parties as part of this exchange. If you'd like to learn more, contact Teten Advisors, LLC.
As a result, Skyvern partnered with Halluminate and created a new benchmark to better quantify these failures.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Financial Services
Company Size
11-50
Company Stage
Seed
Total Funding
$130K
Headquarters
San Francisco, California
Founded
2024
Find jobs on Simplify and start your career today