Together AI provides open-source AI tools and decentralized cloud services to train, fine-tune, and deploy generative models for researchers, developers, and organizations. It runs tasks in the cloud where users run training jobs, manage model versions, and deploy applications via subscriptions and usage fees. It differentiates itself by prioritizing open-source, transparency, and a decentralized cloud approach instead of a proprietary stack. Its goal is to broaden access to powerful AI and build open, verifiable AI systems that benefit society through shared technology.
Company Size
201-500
Company Stage
Series C
Total Funding
$1.3B
Headquarters
San Francisco, California
Founded
2022
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Health Insurance
Company Equity
Strategic partnership between "Mubadala" and "Together AI" to support the AI infrastructure ecosystem in the UAE * Tuesday, September 29, 2026 10:39 AM * 2 min read * Summary * A+ Abu Dhabi, September 29 / WAM / Mubadala Investment Company "Mubadala", the sovereign investment company in Abu Dhabi, and "Together AI", the leading provider of cloud platform and solutions for building and operating generative AI using open-source models, announced a strategic partnership to explore opportunities in the field of AI infrastructure and its ecosystem in the UAE. Under the partnership, "Together AI" will establish its headquarters in Abu Dhabi, contributing to expanding its business scope, making available its inference and research capabilities to meet regional demand, and supporting developers working on building AI applications in the UAE and the region. "Mubadala" and "Together AI" will work to identify strategic opportunities that support the development of the AI ecosystem and its practical applications in the UAE, in line with the objectives of the Emirate of Abu Dhabi aimed at developing sovereign capabilities across the various stages of the AI value chain, enhancing its position as a global center for applied innovation, and consolidating specialized technical expertise and competencies within the country. This strategic cooperation between Mubadala's UAE Investment Division and "Together AI" comes after Mubadala's participation with a $100 million investment in "Together AI"'s Series C funding round, announced on July 1, 2026, alongside a group of global investors in the technology and financial services sectors. "Together AI" was founded in 2022 and is headquartered in San Francisco, United States, and provides an integrated AI platform combining inference services, model fine-tuning, and accelerated computing. Its platform also enables developers and enterprises to build and run applications using open and customized AI models. This partnership falls within Mubadala's ongoing investments in technological infrastructure supporting the UAE's transformation into a knowledge-based economy, and within its strategic focus on AI as a long-term driver for achieving sustainable economic value. Ali Eid AlMheiri, Executive Director of the Diversified Assets Unit in Mubadala's UAE Investment Division, affirmed Mubadala's keenness to attract leading technology companies, enhance national skills and capabilities, and provide an enabling environment in which scientific research is transformed into practical and tangible applications. He added that cooperation with "Together AI" is consistent with Mubadala's strategic approach to supporting the core pillars of the AI ecosystem and expanding access to its advanced infrastructure, noting that the company's establishment of a headquarters in Abu Dhabi will help deepen its cooperation with the country's growing technology sector, achieving long-term sustainable economic value for the Emirate of Abu Dhabi and the UAE. For his part, Vipul Ved Prakash, Founder and CEO of "Together AI", said that global demand for open-source AI solutions is witnessing accelerating growth that exceeds available computing capacity. Developers today are turning to open-source models because of the flexibility they offer in ownership, customization, and scalability, which requires ultra-fast, efficient, and on-demand infrastructure. He added that within this framework, the company is expanding this infrastructure globally, and Mubadala is an ideal strategic partner given its expertise and capabilities that help accelerate "Together AI"'s growth and achieve its expansion goals. * Economy * Science and technology * United Arab Emirates (ARE) * United States of America * Mubadala Investment Company The news is also available in the following languages:
Modal Labs closing $750M round at $15.75B valuation as AI inference demand soars. Tl;dr. * Modal Labs is reportedly closing a $750M round at a $15.75B valuation, more than tripling its valuation from just four months ago amid explosive demand for AI inference. * The surge reflects a broader market shift from model training to inference, where developers are paying a premium for serverless, on-demand GPU infrastructure that scales instantly. * The raise would position Modal as one of the most valuable AI clouds, putting it in direct competition with CoreWeave, Nebius, Lambda, Together AI, and Fireworks AI. A whale-sized round for serverless GPUs. Modal Labs, the New York-based AI infrastructure startup known for its developer-first serverless cloud, is reportedly nearing a massive $750 million funding round that would value the company at $15.75 billion. If finalized, the deal would mark one of the largest private raises for an AI infrastructure company this year and cement Modal as a breakout winner in the second wave of the AI cloud boom. The new valuation represents a stunning leap - more than 3x higher than where the company was valued just four months ago - underscoring how quickly investor appetite for inference infrastructure has accelerated. Sources familiar with the matter say the round is expected to close in the coming weeks, though terms could still shift. Modal has not publicly commented on the raise. From developer darling to decacorn contender. Founded in 2021 by Erik Bernhardsson, former CTO at Better.com and early Spotify engineer behind its music recommendation system, Modal started as a radically simpler alternative to AWS and Kubernetes for running code in the cloud. Its pitch: developers can run any AI workload - from batch jobs to large language model inference - with a single line of Python, no DevOps required. Containers spin up in seconds, scale to thousands of GPUs, and spin down to zero, with customers paying only by the second. That serverless model has struck a chord with AI startups and enterprises drowning in GPU complexity and costs. The company says it now powers workloads for thousands of teams in generative media, voice AI, biotech, and autonomous agents, with usage and revenue growing multiple-fold year-over-year. The reported $15.75 billion valuation is a dramatic jump from Modal's prior valuation earlier this summer, when the company was said to be worth around $5 billion. That prior round itself was already a sharp step-up from its $1 billion-plus valuation in late 2025, charting one of the steepest valuation curves in enterprise tech. Why inference providers are commanding premium valuations. Modal's meteoric rise isn't happening in a vacuum. Investors are pouring tens of billions into what many now call the inference economy. For the past three years, most AI infrastructure capital went toward training - massive clusters of Nvidia H100s and GB200s to build foundation models. Now, the money is shifting to serving those models to hundreds of millions of users. Inference is recurring, high-volume, and increasingly where the margins are. Several forces are driving the premium: First, demand is exploding. The rise of AI agents, real-time voice, video generation, and reasoning models that use 10x to 100x more compute per query has created insatiable need for low-latency GPU capacity. Second, developers want simplicity. Unlike hyperscalers that sell raw capacity by the hour, serverless inference platforms like Modal abstract away autoscaling, cold starts, and GPU orchestration. That ease-of-use commands higher gross margins and fierce loyalty. Third, scarcity equals pricing power. Despite easing shortages, access to latest-generation Nvidia Blackwell chips remains constrained. Neoclouds that secured supply early and can deliver it efficiently are able to grow revenue at triple-digit rates. As one venture investor put it recently, training was a one-time capex boom, inference is a forever software margin. How Modal stacks up against rivals. The AI cloud market has bifurcated into two camps, and Modal is trying to bridge both. On one side are the heavy-asset neoclouds like CoreWeave, now public and valued at over $40 billion, Nebius, Crusoe, and Lambda. These companies raise billions in debt to build massive data centers and lease GPU capacity to hyperscalers and labs on multi-year contracts. On the other side are developer-centric inference platforms like Together AI, last valued at over $3.3 billion, Fireworks AI, valued at $4 billion, and Baseten. These companies compete on speed, model library, and API experience rather than raw megawatts. Modal sits somewhat apart. Unlike pure model-API providers, it lets customers bring any custom model, container, or workflow - not just open-source LLMs. And unlike CoreWeave, it owns no massive data centers itself, instead orchestrating capacity across partners with a software-first layer that optimizes utilization to the second. That asset-light approach has been both its superpower and the bear case against it. Bulls argue Modal can scale faster with far less debt and achieve software-like 70%-plus margins. Skeptics question whether it can guarantee supply of cutting-edge chips at scale without owning the underlying iron, especially as rivals lock up power and Nvidia allocations for years ahead. With $750 million in fresh capital, Modal would have the war chest to answer those doubts - pre-purchasing compute, expanding global regions, and potentially striking dedicated capacity deals while continuing to pour into engineering talent. What comes next. If the round closes as reported, Modal will join an elite tier of private AI infrastructure companies worth over $10 billion, alongside Databricks, OpenAI's infrastructure affiliates, and CoreWeave before its IPO. The key questions now are burn versus growth, path to profitability, and whether an IPO is on the horizon. At $15.75 billion, public-market expectations will be unforgiving - investors will want to see sustained triple-digit growth, enterprise traction beyond startups, and defensibility against both AWS and Nvidia-backed rivals. For now, though, the message from the market is clear: in 2026, inference is king, and Modal Labs is one of its crown princes. AndroGuider Team Articles written by the AndroGuider team. Androguider try to make them thorough and informational while being easy to read.
LLM cascade: re-ask 8% of tasks, get 98.1% at a ninth of the cost. An LLM cascade lets a cheap model answer first and re-asks a large one only on risky answers. Ours got 98.1% right for $86 per million tasks, against $809. Summarize this article with AI An LLM cascade lets a cheap model answer every task and re-asks a larger model on a few. Together AI released Tev on 23 September 2026; two days later XY Space Inc. tested it against Jev and GLM-5.3. Its best cascade put Jev first and re-asked GLM-5.3 on 8.2% of tasks. It got 98.1% right for $86 per million tasks. GLM-5.3 alone got 99.0% for $809. * The rule: Jev answers first. If its answer is one of three it often gets wrong, GLM-5.3 answers the same prompt and its answer is used. * Accuracy: 98.1% on tasks held out from tuning, against 97.2% for Jev alone and 99.0% for GLM-5.3 alone. * Cost: $86 per million tasks, 9.4 times cheaper than GLM-5.3 alone. Re-asked tasks pay for both calls. * Speed: a median of 479 milliseconds per task, because only 8.2% of tasks wait for the second model. * The surprise: starting with Tev, the cheaper model, made a worse cascade. Its best version got 97.1% and re-asked 39% of tasks. * The test: 400 new pick-one tasks, run on 25 September 2026. Every number here comes from its public results file. What is an LLM cascade? An LLM cascade is a setup where a cheap, fast model answers every task and a larger, dearer model re-answers only some of them. You pay the large model's price on a small share of the work, not all of it. In its test the cheap model was Jev, TypeSafe's decision model, and the large one was GLM-5.3, a chat model from Z.ai. The idea has a name in the research. FrugalGPT, a 2023 research paper, called it an LLM cascade. It reported matching GPT-4 with up to 98% lower cost on its tasks. The same pattern also goes by model routing, or "cheap model first, then LLM". The hard part is choosing which tasks to re-ask. There are two common ways: * By confidence. The cheap model reports how sure it is. Anything below a bar goes to the large model. * By answer. You look at which answer the cheap model gave. Some answers are ones it often gets wrong, so those go to the large model. XY Space Inc. has used both. Its operations inbox uses the confidence version. This post measures the answer version, because it needs nothing from the cheap model except its answer. How much does it save? The cascade cost $86 per million tasks, against $809.08 for GLM-5.3 alone. That is 89% less for 0.9 points of accuracy. XY Space Inc. tested on 400 pick-one tasks. Each gives the model an input, a question and three to six options, and the model returns one option. Examples include ruling on a return request, routing an agent to a tool and grading a bug report. The four setups below run on the same tasks. The cascade rows are scored only on tasks their rule was not built from. The arithmetic is simple. A Jev answer costs $0.0000196. A GLM-5.3 answer costs $0.0008091, and the cascade buys one for 8.2% of tasks. That adds $0.0000663, for $0.0000860 a task, or $86 a million. The wider cascade re-asks 17.5% of tasks. It gains 0.4 points and nearly doubles the bill. Both prices are list rates: Jev and Tev charge $0.042 per million input tokens with output free, and GLM-5.3 lists at $1.40 in and $4.40 out on Cloudflare Workers AI. A token is about three quarters of a word. Which answers get re-asked, and why those three? Three of Jev's answers get re-asked: respond_directly, approve_store_credit and approve_full_refund. Each passed two tests. Jev was right less than 90% of the times it gave that answer, and GLM-5.3 did better on those same tasks. Precision means how often a model is right when it gives a particular answer. It is the number that tells you whether to trust an answer you are holding. | Task family | Jev's answer | Times Jev gave it | Jev's precision | GLM-5.3 right on those tasks | | Agent routing | respond_directly | 3 | 66.7% | 100.0% | | Return requests | approve_store_credit | 8 | 75.0% | 100.0% | | Return requests | approve_full_refund | 25 | 88.0% | 100.0% | Two of the three are return rulings. Jev scored 90% on return requests, its weakest family, while GLM-5.3 got every one right. The second test kept two answers off the list. Jev also gave frontend (bug triage) and remove_harassment (moderation) with under 90% precision. Neither made the list, because GLM-5.3 did no better on those tasks. Re-asking them would add cost and no right answers. The routing rule reads only Jev's answer, never the correct one, so it works on live traffic. A plain lookup table decides. Why the cheaper first model made a worse cascade. Tev made a worse first model, even though it costs about half as much per task. With Tev first, the best cascade got 97.1% while re-asking 39% of tasks, for $326 per million. Jev alone scored 97.2% for $20. The reason is where each model's mistakes land. A cascade by answer works when the mistakes cluster in a few answers. Jev's did. Tev's spread out: it gave 17 different answers with precision under 90%, in seven of the eight task families. Jev gave five. * `require_human_review`, 61%. Tev sends safe agent actions to a person. It errs toward caution. * `deny_outside_window`, 64%. Tev declines returns as too late when a different ruling was right. * `remove_scam`, 60%, and `escalate_self_harm`, 50%, in moderation. * `negative`, 67%, in review sentiment. To catch those, the cascade has to re-ask a large share of Tev's traffic. Each re-asked task pays for both calls, so Tev's lower price is gone long before the accuracy catches up. The first model's price matters less than how neatly its mistakes group. Jev vs Tev has the full head to head. How to build your own cascade. Build it from your own labelled data. Its list of three answers fits its 400 tasks, not yours. You need a few hundred labelled examples and both models' answers on them. * Label a few hundred real examples. Write the correct answer for each. Include the rare answers, since those are where the cheap model slips. Its guide to testing a classifier covers writing test items in pairs that differ by one edit. * Run both models on all of them. Use the prompts you will ship. Record every answer, the cost and the time. * Work out the cheap model's precision for each answer. For each answer it gave, count how often it was right. * Pick the risky list on part of the data and score it on the rest. XY Space Inc. built the list on four fifths of the task pairs and scored the other fifth, five times over, on three shuffles. Keep the setup that is cheapest within the accuracy you need. * Price the re-asked tasks at both calls. The cheap call has already happened when you decide to re-ask. * Re-check as your data changes. New products, new policies and new phrasing move precision. Re-run the labelled set on a schedule. The confidence version follows the same steps with a confidence bar in place of the list. How to use Jev shows where to set that bar. Its inbox uses it: Jev keeps the items it is sure of, and Claude decides the rest. Jev vs GLM-5.2 has those results. Honest limits. Four limits apply to these numbers: * The 98.1% is slightly optimistic. XY Space Inc. tried 20 cascade setups and picked the best one using the same results. The rule itself was built and scored on separate tasks, but the choice of threshold was not. * The set is small. 400 tasks means wide error bars. Some answers appear only a handful of times. respond_directly was Jev's answer three times. * Claude wrote the tasks, and the same kind of model checked the labels, with spot checks. None came from public datasets. * The list is specific to these task families. Yours will differ. That is why step 1 above exists. Latency has a cost too. A re-asked task waits for Jev and then GLM-5.3. At the two medians that adds up to about 3.3 seconds, against 459 milliseconds for Jev alone. The cascade's median stays at 479 milliseconds because only 8.2% of tasks wait. Jev ran through the AI Space gateway and GLM-5.3 through AI Space with reasoning at its defaults, so both times include that extra hop. For when a large model alone is worth the price, see decision model vs LLM. For comparing models by cost per right answer, see the cheapest model per solved task. For what a decision model is, see what is Jev. Faq. What is the difference between an LLM cascade and model routing? The words overlap. A cascade usually runs the cheap model first and escalates after seeing its answer. Routing can also mean choosing a model before any call, from the task itself. Its setup is a cascade: Jev always answers, and its answer decides whether GLM-5.3 is asked too. Does an LLM cascade always save money? No. It saves money when the cheap model is right most of the time and its mistakes cluster in answers you can spot. Its Tev-first cascade re-asked 39% of tasks and cost $326 per million. That was worse than Jev alone on both accuracy and price. Should I route by confidence or by answer? Route by answer when the cheap model's mistakes cluster in a few answers. Route by confidence when it reports a trustworthy confidence and mistakes are spread out. You can test both on the same labelled set. Its inbox routes by confidence; this benchmark routes by answer. How many examples do I need to build a cascade? A few hundred labelled examples is a practical start. XY Space Inc. used 400. Rare answers need enough examples to measure precision, so collect more of those. Score the rule on examples it was not built from. Is the large model's answer always right when it is re-asked? No. On its three risky answers, GLM-5.3 got every re-asked task right. Across all 400 tasks it scored 99.0%. Keep a person on the decisions where a wrong answer is costly.
Equinix, Together AI, Nvidia partner on Inference Exchange. Aims to give enterprises access to secure, low-latency AI inferencing September 03, 2026 Data center operator Equinix is working with Nvidia and Together AI on an Inferencing solution. Described as a distributed AI inference platform, the Inference Exchange uses Nvidia's Enterprise Reference Architecture and infrastructure and Together AI's inference platform, and is delivered through Equinix's global data centers. Equinix has a footprint of more than 280 data centers across 77 metros, 230 cloud on-ramps, and more than 10,500 businesses interconnected on its exchange. Via Equinix Fabric, the Inference Exchange will connect to clouds, networks, and AI providers. "AI is transforming enterprise technology at extraordinary speed, and the infrastructure decisions enterprises make today will define their competitive position for years to come. Equinix is uniquely positioned to deliver what this moment demands based on our nearly three decades building the trusted exchange where the world's enterprises run, connect, and orchestrate their most critical workloads," said Adaire Fox-Martin, CEO and president, Equinix. "Our longtime relationship with Nvidia delivers the accelerated computing foundation at the heart of modern AI, while Together AI's commitment to open ecosystems gives enterprises the flexibility to scale on their terms. Equinix Inference Exchange will enable architectures that are neutral by design, open by default, and engineered for exceptional performance." "Together AI was built on the conviction that open, accessible AI is what will define the industry moving forward, because enterprises shouldn't have to choose between model performance and operational flexibility," added Vipul Ved Prakash, co-founder and CEO of Together AI. "What we are building with Equinix and Nvidia proves that model choice and performance are not trade-offs. They are the foundation of enterprise AI done right." Among the use cases the offering is targeting are metro Edge inference for organizations that require low latency for AI workloads, open model migration for enterprises looking to move from closed proprietary models to open-source alternatives using Together AI's platform, and sovereign AI workloads that need to run in specific locations. The Inference Exchange will be available from Q1 2027. AI cloud company Together AI raised $800m in a Series C funding round in July 2026. The company is a known customer of Rum Group and is planning to deploy hardware at an L&T data center, while Hypertec and 5C are aiming to roll out 2GW of capacity for the company. More in cloud & hyperscale.
Equinix has launched Inference Exchange, a distributed AI inference programme combining NVIDIA Enterprise Reference Architectures with Together AI's inference platform, supporting over 200 open-source models. The solution will be delivered through Equinix's global data centre network. The collaboration addresses the challenge of deploying AI inference closer to users, data and applications across clouds, models and geographies. Equinix operates more than 280 data centres across 77 metropolitan areas, with 230 cloud on-ramps and over 10,500 interconnected businesses. Together AI brings open-model flexibility to the platform. The solution aims to reduce deployment complexity by connecting to inference providers across major metropolitan areas whilst providing access to an extensive ecosystem of clouds, networks and AI providers. Equinix Inference Exchange will be available from the first quarter of 2027.