Full-Time

Member of Technical Staff

Research

The Token Company

The Token Company

1-10 employees

AI middleware for semantic prompt compression

No salary listed

H1B Sponsorship Available

San Francisco, CA, USA

In Person

Category
AI & Machine Learning (1)
Required Skills
LLM
Machine Learning
RAG
Reinforcement Learning

Get referred to The Token Company

See people who can refer or advise you

Requirements
  • Trained models from scratch and independently owned the data, architecture, and training loop.
  • Demonstrated strong machine learning fundamentals, including fluency with transformers and the ability to turn a paper or rough idea into a real training run.
  • Demonstrated the ability to learn quickly and iterate on new ideas independently.
  • Demonstrated a production focus and interest in deploying models for real users.
  • Experience primarily involving retrieval-augmented generation, agents, prompt engineering, or fine-tuning existing models through an application programming interface is not a fit for this role.
  • Preference for shipping models over publishing papers is expected for this role.
Responsibilities
  • Own the model training stack end to end, including data, architecture, training, evaluation, and shipping compression models into production.
  • Design and train models from scratch on NVIDIA B200 graphics processing units.
Desired Qualifications
  • Pretrained a transformer, or completed serious post-training or reinforcement learning on one.
  • Built a novel architecture or training method with results to support it.
  • Shipped a trained model into something people actually use.

The Token Company (Otsofy) provides an AI middleware API that semantically compresses prompts for large language models to reduce costs and latency. It uses a fast preprocesser, bear-1, to shorten input prompts before sending them to LLMs like GPT and Claude, preserving meaning while removing redundant tokens and supporting up to 100,000 tokens in under 100 ms. Unlike providers tied to a single model, it aims to be a neutral efficiency layer that works across multiple LLMs to improve speed and accuracy. Its goal is to become an essential component in the AI development stack, maximizing efficiency for every LLM request.

Company Size

1-10

Company Stage

Seed

Total Funding

$130K

Headquarters

Helsinki, Finland

Founded

2025

Get referred to The Token Company

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • March 2026 funding from Y Combinator and Pioneer Fund validated the category.
  • April 2026 site showed 66% token reduction and onboarding users.
  • September 2026 SOC 2 Type II and HIPAA BAA unlock enterprise deals.

What critics are saying

  • OpenAI and Anthropic prompt caching attacks the same cost problem immediately.
  • LLMLingua and other OSS tools undercut pricing and compressions by 2026.
  • Without fast customer adoption, 2026 runway runs out before Type II SOC 2.

What makes The Token Company unique

  • YC W26-backed middleware compresses prompts before GPT and Claude calls.
  • Bear-1 compresses 100,000 tokens in under 100 milliseconds.
  • Only standalone commercial prompt-compression API in 2026, unlike OSS LLMLingua.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity