Elorian

Elorian

Creates native multi-modal AI models

Overview

Elorian develops native multi-modal AI models that understand text, images, video, and audio to interpret the physical world. Its core capability, visual reasoning, lets AI agents perform tasks by understanding real-world dynamics rather than simply adding vision to a text model, and it can be applied to robotics and GUI automation without direct API integration. The company emphasizes end-to-end multi-modal understanding and GUI-enabled automation built from the ground up, instead of layering vision onto an existing language foundation. The goal is to move toward artificial general intelligence by enabling AI to comprehend and act within the physical world through vision and multi-modal perception for complex tasks across domains.

Launched Recently

About Elorian

Simplify's Rating
Why Elorian is rated
C+
Rated C on Competitive Edge
Rated B on Growth Potential
Rated C on Differentiation

Industries

Data & Analytics

Robotics & Automation

Enterprise Software

AI & Machine Learning

Company Size

11-50

Company Stage

Seed

Total Funding

$55M

Headquarters

null

Founded

2025

Get referred to Elorian

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • April 2026 stealth exit and $55M seed funded compute, hiring, and pilots.
  • TechCrunch reported potential 2026 API release, creating an early commercialization path.
  • Nvidia, Menlo, Altimeter, and Jeff Dean validation attracts customers and recruits.

What critics are saying

  • OpenAI o3 and Google Gemini Omni already ship native visual reasoning.
  • Elorian has no product or revenue; TechCrunch reported $55M seed at $300M valuation.
  • A failed 2026 API launch leaves Elorian trapped as a research lab and talent magnet.

What makes Elorian unique

  • Andrew Dai and Yinfei Yang bring DeepMind, Google, and Apple multimodal pedigree.
  • Elorian builds visual reasoning natively, not text-first adapters, for physical-world tasks.
  • Its dataset-and-model co-design targets GUI automation, robotics, and spatial understanding.

Help us improve and share your feedback! Did you find this helpful?

Funding

Total Funding

$55M

Above

Industry Average

Funded Over

1 Rounds

Notable Investors:
Seed funding is usually the first official round after pre-seed, when a startup has a prototype or concept. It’s typically used to develop the product, test the market, and start building the team. Investors here are often angel investors or early-stage venture capitalists.
Seed Funding Comparison
Above Average

Industry standards

$3.3M
$2M
Netflix
$2.3M
Instacart
$3M
Robinhood
$55M
Elorian

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Unlimited Paid Time Off

Parental Leave

Relocation Assistance

Company News

WhalesBook Private Limited
Jul 16th, 2026
Elorian AI raises $55M seed funding at $300M valuation with no product or revenue

Elorian AI has raised $55 million in seed funding at a $300 million valuation, despite having no product or revenue. The startup, led by former Google DeepMind researcher Andrew Dai, focuses on visual AI technology that enables machines to interpret visual information. The company's strategy centres on hiring top-tier engineering talent, competing directly with established tech giants. However, this approach requires substantial capital and poses execution risks. Investors face significant uncertainties, as the company operates in a competitive sector where larger firms could develop rival solutions more quickly. Success will depend on Elorian's ability to transition from research to a revenue-generating product whilst managing its cash effectively. The investment reflects broader trends in AI funding, where high valuations are driven by leadership expertise rather than proven commercial track records.

Unicorner
Jun 22nd, 2026
Moving AI beyond language

Elorian. Moving AI beyond language. Jun 22, 2026 Unicorner is about to pull off its most ambitious event yet in SF this Wednesday. If you're coming to Figma Config, join Unicorner, Bolt.new, and Notion for its experiential festival turned afterparty: Welcome to the Stratosphere. Food, drinks, and robots. Immediately followed by a live performance by Cheat Codes, the award-winning DJ trio behind top songs like No Promises. Spots are limited. Vibes are immaculate. See you there. * Tues. June 23 - AI Champions Dinner (San Francisco) - If you're spearheading AI usage at your company, this curated, three-course seated dinner is for you. * Wed. June 24 - Welcome to the Stratosphere (San Francisco) Elorian is an AI research and product lab building models that reason through visual information. The company's thesis is that current AI systems are still too dependent on language. Today's vision-language models often convert visual inputs into text before reasoning over them. Elorian argues that this approach starts to break down when the task depends on spatial relationships, physical constraints, or details that are hard to compress into words. Elorian is building what its founder and CEO Andrew Dai calls "visual thinking models." These are AI systems designed to reason in the visual domain rather than convert images into text first. The goal is to bring the kind of reasoning progress seen in coding and math into visual tasks that require the ability to interpret real-world scenes. If models can reason visually instead of only describing what they see, Elorian believes they could improve work across engineering, robotics, medicine, science, weather monitoring, disaster response, and precision agriculture. Elorian is pre-product, with plans to release a model API by the end of the year, focusing on technical teams building products where visual understanding is core. The first commercial layer is expected to center on access to its visual thinking model. The company is already speaking with potential customers about pilots across video understanding, robotics, engineering, and other vision-heavy workflows. Long term, Elorian may build a platform around the model, with tools that make visual reasoning easier to use inside existing systems. * Raised $55 million in seed funding from Striker Ventures, Menlo Ventures, and Altimeter, with participation from 49 Palms and prominent AI researchers including Jeff Dean * Emerged from stealth in April 2026 as a multimodal reasoning research and product lab focused on visual intelligence In the summer of 2025, Andrew Dai started testing frontier models on something familiar: board games. He often played with colleagues and friends, and after one game, he took photos of the board and asked the models simple questions like how many points a player had and how many resources were on the board. The models struggled with questions a person could answer simply by looking. Dai kept testing. He tried a New York Times crossword, then took a picture of a bar and asked what drinks were there, how many of each, and what needed to be refilled. The pattern held across tasks that required careful visual understanding rather than language. That was the mismatch Dai kept coming back to. The industry was already talking about AGI, and frontier models could write, code, and generally reason through text. But when the task required basic visual understanding, they were still missing things that felt obvious to humans. Dai had spent nearly 14 years at Google, including Google Brain and DeepMind, researching large-scale AI systems. Around the same time, Gemini and other frontier models were focusing heavily on coding and text-based reasoning, and Dai saw visual reasoning as an underdeveloped domain. A coworker connected him with Yinfei Yang, who had worked as a research scientist at Google and Apple and had specialized in multimodal foundation models. The fit made sense quickly. Dai left around Thanksgiving, Yang followed in December, and they incorporated Elorian that same month. AI has made enormous progress in language, but it still struggles with the visual world. That gap matters because much of human intelligence is not text-first. People understand space, motion, physical constraints, and visual relationships before they can explain them in words. A model that describes an image is not necessarily a model that reasons through what is happening inside it. Current vision-language models often rely on a translation step. They convert visual information into language, then reason over the text. Elorian argues that this chain is fragile because many visual tasks cannot be compressed cleanly into words. That blind spot shows up in surprisingly ordinary places. Dai tested frontier models on board games, photos, crosswords, bar shelves, maps, mazes, counting tasks, and other problems where the answer was already visible. These models could see everything they needed to. Where they struggled was to reason through visual structure. Recent research points in the same direction. BabyVision, a benchmark designed to test visual reasoning beyond language, found that state-of-the-art multimodal models still fail on basic visual tasks that even 3-year-olds can solve easily. Its results showed leading multimodal models performing well below human baselines, reinforcing that visual reasoning remains underdeveloped. Generating an image and reasoning visually are different problems. A model can create a convincing picture without understanding whether a design works, whether a robot can move through a room, whether a metro map connects correctly, or whether a count is off by one in a safety-critical setting. Elorian is designed to reason in the visual domain, building the model and the missing visual-reasoning dataset together from scratch. The timing matters. Text-based reasoning models like OpenAI's o1 unlocked major progress in coding and math, but that same reasoning shift has not fully reached the visual domain. The team makes that argument more credible. Dai and Yang are leading researchers who pushed language, data, and multimodal models to the frontier, giving them a clear view of both the progress and the limitations from the inside. Frontier labs are already pushing multimodal systems, and many companies will add stronger visual capabilities to existing models. The question is whether visual reasoning becomes a feature inside language-first systems or a separate foundation that needs to be built differently. Elorian is oriented around the second view: the next major step in AI will require models built to reason visually from the start. The visual thinking data it needs is not sitting online, waiting to be scraped. If it were, the major labs would already be using it. Elorian has to build the model and the data layer together, and that is what makes its wedge sharper. The opportunity is large for a different reason. The AI market is over-indexing on coding and math because those are the places where reasoning models first showed obvious progress. However, much of the real work is done visually too. Engineers design physical products visually. Traders read charts visually. Insurance teams evaluate damage visually. Similarly, robots need to understand the world visually before they can act inside it. No one designs the next iPhone entirely in code or builds rockets through text alone. If AI is going to move deeper into the physical world, visual reasoning cannot remain a secondary capability. The bullish case for Elorian is simple: the world is visual, and most AI still reasons as if it is not.

FinSMEs
Apr 10th, 2026
Elorian Raises $55M in Seed Funding at $300M Valuation

Elorian, a Palo Alto, CA-based AI research and product lab focused on advancing visual reasoning for AI, raised $55M in a seed funding round at a $300M valuation

Tech in Asia
Apr 10th, 2026
Ex-Google DeepMind researchers' startup Elorian raises $55M for visual AI reasoning

Elorian, a Palo Alto-based AI startup co-founded by former Google DeepMind researcher Andrew Dai, has emerged from stealth with $55 million in funding at a $300 million valuation. Menlo Ventures, Altimeter Capital, Striker Venture Partners, Nvidia and AI researcher Jeff Dean backed the round, raised across two tranches at $120 million and $300 million valuations. The pre-revenue company is building models that reason about images and visual data, focusing on physical AI for robotics and autonomous systems. Co-founders include former Google and Apple researcher Yinfei Yang and former Harvard professor Seth Neel. Elorian plans to release its first public reasoning model in approximately 12 months and is currently in discussions with potential customers.

Bloomberg
Apr 9th, 2026
Ex-Google DeepMind Researchers Debut Startup Called Elorian Focused on Visual AI

Former Google DeepMind researcher Andrew Dai believes that the artificial intelligence models at big labs have the intelligence of a 3-year-old kid, at least when it comes to making sense of visual prompts.

Recently Posted Jobs

Sign up to get curated job recommendations

Elorian is Hiring for 3 Jobs on Simplify!

Find jobs on Simplify and start your career today

Don't see your dream role? Check out thousands of other roles on Simplify. Browse all jobs →