Guide Labs

Guide Labs

Develops interpretable AI systems via API

Overview

Guide Labs builds interpretable AI systems and makes them available via an API. The company designs models and training pipelines that prioritize understandability and debuggability, redefining architecture, loss functions, and workflow so model behavior can be inspected and fixed by humans and domain experts. Their product works by delivering AI models whose properties are constrained and made transparent during training, enabling easier error identification and alignment, with an API-based access for developers to integrate. Compared with typical ML approaches that optimize only for performance on monolithic models, Guide Labs focuses on steerable, auditable systems whose decisions can be traced and adjusted. The goal is to create AI models that humans can reliably understand, debug, and align with specific objectives, improving trust and controllability in AI deployments.

YC Company
Significant Headcount Growth

About Guide Labs

Simplify's Rating
Why Guide Labs is rated
B-
Rated B on Competitive Edge
Rated B on Growth Potential
Rated C on Differentiation

Industries

Data & Analytics

Enterprise Software

AI & Machine Learning

Company Size

11-50

Company Stage

Seed

Total Funding

$130K

Headquarters

San Francisco, California

Founded

2023

Get referred to Guide Labs

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Guide Labs shipped Steerling-8B on February 23, 2026, proving execution.
  • The company says Steerling-8B reaches 90% capability with less training data.
  • Open-sourcing weights and code accelerates adoption among researchers and startups.

What critics are saying

  • Steerling-8B is only 8B parameters, while frontier labs outspend and outscale rapidly.
  • The 4,096-token context ceiling blocks many 2026 enterprise RAG and agent workflows.
  • Apache-licensed weights plus synthetic-data licensing issues invite downstream commercial disputes.

What makes Guide Labs unique

  • Steerling-8B traces tokens to training data, concepts, and prompt context.
  • Guide Labs builds interpretability into the architecture, not post-hoc analysis.
  • The model emphasizes controlled concept steering for regulated finance and scientific workflows.

Help us improve and share your feedback! Did you find this helpful?

Funding

Total Funding

$130k

Below

Industry Average

Funded Over

1 Rounds

Notable Investors:
Seed funding is usually the first official round after pre-seed, when a startup has a prototype or concept. It’s typically used to develop the product, test the market, and start building the team. Investors here are often angel investors or early-stage venture capitalists.
Seed Funding Comparison
Below Average

Industry standards

$3.3M
$130k
Guide Labs
$1.5M
Slack
$2M
Netflix
$2.3M
Instacart
$3M
Robinhood

Growth & Insights and Company News

Headcount

6 month growth

14%

1 year growth

14%

2 year growth

23%
TechAmerica
Feb 24th, 2026
Guide Labs unveils a new generation of interpretable large language model

Guide Labs unveils a new generation of interpretable large language model. Guide Labs introduces an interpretable large language model designed to improve transparency, explainability, and trust in AI systems for enterprise and research use. Feb 24, 2026 - 17:39 Feb 24, 2026 - 17:40 One of the hardest parts of working with deep learning systems is figuring out why they behave the way they do. Whether it's xAI repeatedly running into trouble while fine-tuning Grok's political edge, ChatGPT drawing criticism for sycophantic behaviour, or everyday hallucinations that appear in many models, it remains difficult to "look inside" a neural network with billions of parameters and clearly explain what caused a specific output. Guide Labs, a San Francisco startup led by CEO Julius Adebayo and chief science officer Aya Abdelsalam Ismail, says it has a solution. On Monday, the company open-sourced an 8-billion-parameter large language model called Steerling-8B, trained using a new architecture designed for interpretability. The key promise: every token generated by the model can be traced back to its origins in the model's training data. In practice, that traceability could mean identifying the exact reference materials behind the facts the model states, or delving much deeper into how it forms ideas about complex concepts such as humour, identity, or gender. "If I have a trillion ways to encode gender, and I encode it in 1 billion of the 1 trillion things that I have, you have to make sure you find all those 1 billion things that I've encoded, and then you have to be able to turn that on, turn them off reliably," Adebayo said. "You can do it with current models, but it's very fragile... It's sort of one of the holy grail questions." Adebayo began this line of research while working on his PhD at MIT. He co-authored a widely cited 2018 paper showing that existing methods for understanding deep learning systems were not reliable. That work eventually fed into a different approach for building LLMs. Rather than trying to interpret a black-box model after it has been trained, developers design interpretability into the architecture from the beginning. Guide Labs' method introduces a concept layer within the model. This layer "buckets" information into traceable categories, allowing specific outputs to be linked back to organised, labelled sources. The tradeoff is that the approach requires more up-front data annotation. But the company says it can reduce the burden by using other AI systems to assist with labelling, enabling it to scale up training. Steerling-8B is the company's largest proof-of-concept so far. "The kind of interpretability people do is... neuroscience on a model, and we flip that," Adebayo said. "What we do is actually engineer the model from the ground up so that you don't need to do neuroscience." A natural concern with any more controlled, structured architecture is that it might dampen the emergent behaviour that makes LLMs valuable - especially their ability to generalise to new situations or reason about concepts that were not explicitly taught during training. Adebayo argues that this kind of generalisation still appears in Sterling-8 B. His team tracks what it calls "discovered concepts," ideas the model appears to generate on its own, such as quantum computing. Adebayo believes interpretability will become essential across the industry. For consumer-facing LLMs, he says this approach could help model builders block or limit the use of copyrighted materials and provide stronger, more reliable control over sensitive outputs involving topics like violence or drug abuse. He also points to regulated environments - such as finance - where a model assessing loan applicants should consider legitimate factors, such as financial records, while excluding protected attributes, such as race. In these settings, controllable and transparent systems will matter more. The company also sees interpretability as a growing need in scientific applications. Deep learning has produced breakthroughs in areas like protein folding. Still, researchers often want more clarity on why a model generated a particular prediction or why it concluded that a certain combination is promising. Guide Labs says it has also developed technology aimed at this scientific interpretability problem. "This model demonstrates that training interpretable models is no longer a sort of science; it's now an engineering problem," Adebayo said. "We figured out the science, and we can scale them, and there is no reason why this kind of model wouldn't match the performance of the frontier-level models," even though those frontier systems typically have far more parameters. Guide Labs claims that Sterling-8 B can achieve 90% of the capability of existing models with less training data, attributing this to the model's architecture. The company's next step is to build a larger version and begin offering API access along with more agent-focused capabilities for users. Guide Labs emerged from Y Combinator and raised a $9 million seed round led by Initialised Capital in November 2024. Adebayo says the bigger mission is to make interpretability a standard feature of advanced AI systems rather than a specialised research effort. "The way we're currently training models is super primitive, and so democratising inherent interpretability is actually going to be a long-term good thing for our role within the human race," Adebayo said. "As we're going after these models that are going to be super intelligent, you don't want something to be making decisions on your behalf that's sort of mysterious to you."

HiTech.Expert
Feb 24th, 2026
Guide Labs introduces a new type of interpretive LLM

Guide Labs introduces a new type of interpretive LLM. 24.02.2026 The challenge with developing a deep learning model is often understanding why it does what it does: whether it's xAI's repeated attempts to fix Grok's weird policy, ChatGPT's fight against flattery, or plain old hallucinations, understanding a neural network with billions of parameters is not easy. Guide Labs, a San Francisco startup founded by CEO Julius Adebayo and Chief Scientific Officer Aya Abdelsalam Ismail, is now offering an answer to that problem. On Monday, the company has made it public The 8-billion-parameter LLM, Steerling-8B, is trained using a new architecture designed to make its actions easily interpretable: each token produced by the model can be traced back to its origin in the LLM training data. This can be as simple as identifying reference materials for the facts the model references, or as complex as understanding a humor or gender model. "If I have a trillion ways to encode gender, and I encode it in 1 billion of the 1 trillion things that I have, you have to make sure you find all of those 1 billion things that I encoded, and then you have to be able to reliably turn them on and off," Adebayo told TechCrunch. "You can do that with current models, but it's very unstable... It's kind of one of those Holy Grail questions." Adebayo began this work while pursuing his PhD at MIT, co-authoring a widely cited 2018 paper that showed that existing methods for understanding deep learning models are unreliable. That work ultimately led to a new way to build LLMs: Developers embed a conceptual layer into the model that breaks down data into observable categories. This requires more pre-annotation of the data, but with the help of other AI models, they were able to train the model as the largest proof-of-concept to date. "The interpretation that people make is... neurobiology on a model, and we're flipping that around," Adebayo said. "We're actually building a model from scratch so you don't have to do the neurobiology." One concern about this approach is that it could eliminate some of the novel behaviors that make LLMs so interesting: their ability to generalize in new ways about things they haven't been trained on. Adebayo says this is still happening in his company's model: his team is tracking what they call "open concepts" that the model has discovered on its own, such as quantum computing. Adebayo argues that such interpretable architecture will be needed by everyone. For consumer-facing LLMs, these techniques should allow modelers to block the use of copyrighted material or better control the results on topics like violence or drug abuse. Regulated industries will need more controlled LLMs - for example, in finance - where a model evaluating loan applicants must consider factors like financial records but not race. Interpretability is also needed in scientific work - another area where Guide Labs has developed technology. Protein assembly has been a big success for deep learning models, but scientists need a deeper understanding of why their software identified promising combinations. "This model demonstrates that training interpretive models is no longer a science; it's now an engineering problem," Adebayo said. "We've figured out the science and we can scale them, and there's no reason why this model can't match the performance of top-tier models" that have many more parameters. Guide Labs claims that the Steerling-8B can achieve 90% of the capabilities of existing models, but uses less data for training thanks to its new architecture. The next step for the company, which emerged from Y Combinator and raised $9 million in seed funding from Initialized Capital in November 2024, is to build a larger model and start providing API and agent access to users. "The way we train models right now is extremely primitive, so democratizing innate interpretability is actually going to be a long-term benefit to our role as humanity," Adebayo told TechCrunch. "As we strive to build highly intelligent models, you don't want something that is a mystery to you making decisions on your behalf."

TechCrunch
Feb 23rd, 2026
Guide Labs open-sources 8B-parameter LLM with traceable AI architecture after $9M seed round

Guide Labs, a San Francisco-based startup founded by CEO Julius Adebayo and chief science officer Aya Abdelsalam Ismail, has open-sourced Steerling-8B, an 8 billion parameter large language model with a novel interpretable architecture. Every token produced can be traced back to its origins in training data. The company inserts a concept layer that buckets data into traceable categories, allowing developers to understand and control model behaviour without post-hoc analysis. Guide Labs claims Steerling-8B achieves 90% of existing models' capability whilst using less training data. The startup emerged from Y Combinator and raised a $9 million seed round from Initialized Capital in November 2024. It plans to build larger models and offer API access, targeting regulated industries like finance and scientific applications requiring transparent AI decisions.

AI Finder Guru
Feb 23rd, 2026
Guide Labs Launches AI Transparency Platform to Decode Complex Neural Networks

Guide Labs launches AI transparency platform to decode complex neural networks. Understanding why a deep learning model behaves the way it does remains a significant challenge. This is evident in the ongoing efforts to adjust Grok's unconventional political responses, ChatGPT's tendencies toward sycophancy, and the common issue of AI hallucinations. Navigating a neural network with billions of parameters is inherently complex. A San Francisco startup called Guide Labs, founded by CEO Julius Adebayo and chief science officer Aya Abdelsalam Ismail, is now presenting a potential solution. The company recently open-sourced an 8 billion parameter large language model named Steerling-8B. This model features a novel architecture specifically designed for interpretability, allowing every token it generates to be traced back to its source within the training data. This capability ranges from simply identifying the references for factual statements to analyzing the model's comprehension of nuanced concepts like humor or gender. "You can attempt this with current models, but the process is fragile... It represents one of the holy grail questions in the field," Adebayo noted. Adebayo initiated this research during his PhD studies at MIT. He co-authored an influential 2020 paper demonstrating that existing methods for interpreting deep learning models were unreliable. This foundational work eventually led to a new approach for constructing LLMs, which involves inserting a conceptual layer that organizes data into traceable categories. While this method demands more extensive initial data annotation, the team leveraged other AI models to assist, enabling them to train Steerling-8B as their largest proof of concept to date. "Typical interpretability work is akin to performing neuroscience on a model. We invert that paradigm," Adebayo explained. "We engineer the model from its foundation so that such intensive analysis becomes unnecessary." A potential concern with this structured approach is that it might suppress the emergent behaviors that make LLMs fascinating - their ability to generalize and reason about topics beyond their explicit training. Adebayo asserts that his company's model still exhibits this quality; his team monitors what they term "discovered concepts," such as quantum computing, which the model identifies independently. He contends that this interpretable architecture will become essential across the board. For consumer-facing LLMs, these techniques could enable developers to block the use of copyrighted material or exert finer control over outputs related to sensitive subjects like violence or drug abuse. Regulated industries, such as finance, will require highly controllable models - for instance, an LLM evaluating loan applications must consider financial history while rigorously excluding factors like race. Interpretability is also critical in scientific research, another area where Guide Labs has developed technology. While deep learning has achieved breakthroughs in fields like protein folding, scientists need greater insight into why their software arrives at successful solutions. "What this model demonstrates is that training interpretable models is no longer purely a scientific pursuit; it's an engineering challenge," Adebayo stated. "We have solved the core science and can now scale these models. There's no inherent reason why this approach cannot match the performance of leading frontier models," which often possess far more parameters. Guide Labs reports that Steerling-8B achieves approximately 90% of the capability of existing models while utilizing less training data, thanks to its innovative design. The company, which graduated from Y Combinator and secured a $9 million seed round from Initialized Capital in November 2024, plans to develop a larger model next and begin offering API and agentic access to users. "As we pursue the development of super-intelligent models, it's crucial that systems making decisions on your behalf are not mysterious black boxes," Adebayo emphasized.

GuideLabs
Feb 23rd, 2026
Steerling-8B: The First Inherently Interpretable Language Model

Steerling-8B: the first inherently interpretable language model. Published: February 23, 2026 Guide Labs Inc. is releasing Steerling-8B, the first interpretable model that can trace any token it generates to its input context, concepts a human can understand, and its training data. Trained on 1.35 trillion tokens, the model achieves downstream performance within range of models trained on 2-7x more data. Steerling-8B unlocks several capabilities which include suppressing or amplifying specific concepts at inference time without retraining, training data provenance for any generated chunk, and inference-time alignment via concept control, replacing thousands of safety training examples with explicit concept-level steering. Overview. For the first time, a language model, at the 8-billion-parameter scale, can explain every token it produces in three key ways. More specifically, for any group of output tokens that Steerling generates, Guide Labs Inc. can trace these tokens to: * [Input context] the prompt tokens, * [Concepts] human-understandable topics in the model's representations, and * [Training data] the training data drove the output. Artifacts. Guide Labs Inc. is releasing the weights of a base model trained on 1.35T tokens as well as companion code to interact and play with the model. Steerling-8B in action. Below Guide Labs Inc. show Steerling-8B generating text from a prompt across various categories. You can select an example, then click on any highlighted chunk of the output. The panel below will update to show: * Input Feature attribution: which tokens in the input prompt strongly influenced that chunk. * Concept attribution: the ranked list of concepts, both tone (e.g. analytical, clinical) and content (e.g. Genetic alteration methodologies), that the model routed through to produce that chunk. * Training data attribution: how the concepts in that chunk distribute across training sources (ArXiv, Wikipedia, FLAN, etc.), showing where in the training data the model's knowledge originates. Steerling is built on a causal discrete diffusion model backbone, which lets Guide Labs Inc. steer generation across multi-token tokens rather than only at the next-token. The key design choice is decomposing the model's embeddings into three explicit pathways: ~33K supervised "known" concepts, ~100K "discovered" concepts the model learns on its own, and a residual that captures whatever remains. Guide Labs Inc. then constrain the model with training loss functions that ensure the model routes signal through concepts without a fundamental tradeoff with performance. The concepts feed into logits through a linear path, every prediction decomposes exactly into per-concept contributions, and Guide Labs Inc. can edit those contributions at inference time without retraining. For the full architecture, training objectives, and scaling analysis, see Scaling Interpretable Models to 8B. Performance. Despite being trained on significantly fewer compute than comparable models, Steerling-8B achieves competitive performance across standard benchmarks. The figure below shows average performance (across 7 benchmarks) versus approximate training FLOPs on a log scale, with vertical lines marking multiples of Steerling's compute budget. Interpretability. In the previous update, Guide Labs Inc. shared several ways that assess how interpretable a model's representations are. Here Guide Labs Inc. provide another metric that gives insight into the model's use of its concepts. On a held-out validation set, over 84% of token-level contribution comes from the concept module: the model is not just using the residual to make its predictions. This matters for control: if the model's predictions genuinely flow through concepts, then editing those concepts at inference time actually changes what the model does rather than nudging a side channel while the real work happens elsewhere. A useful check is what happens when Guide Labs Inc. remove the residual pathway. On several LM Harness tasks, dropping the residual has only a small effect, which suggests the model's predictive signal is largely routed through concepts rather than hidden "everything-else" channels. Finally, Steerling can detect known concepts in text with 96.2% AUC on a held-out validation dataset. What this unlocks. In the coming weeks, Guide Labs Inc.'ll be releasing deep dives on each of these capabilities: * Concept steering: precise control via intervention; * Concept discovery: what did Steerling learn that Guide Labs Inc. didn't teach it? Guide Labs Inc.'ll open up the discovered concept space and show structure that surprised Guide Labs Inc.. * Alignment without fine-tuning: replace thousands of safety training examples with a handful of concept-level interventions. * Memorization & training data valuation: trace any generation back to the training data that produced it, and assign value to individual data sources. * The case for inherent interpretability: what do you gain when interpretability is designed in from the start, and what do you miss when it's bolted on after the fact? Guide Labs Inc.'ll cover each of these in detail in upcoming posts, with quantitative evaluations and deployment-oriented case studies.

Recently Posted Jobs

Sign up to get curated job recommendations

There are no jobs for Guide Labs right now.

Find jobs on Simplify and start your career today

We update Guide Labs's jobs every few hours, so check again soon! Browse all jobs →