
Work Here?
LM Studio provides a local-first AI platform that lets users download, install, and run large language models directly on their own computers, avoiding cloud services. It supports multiple LLM frameworks and offers an easy interface for configuration and customization, so users can tailor models for tasks like natural language processing, text generation, and data analysis. By running models locally, LM Studio keeps data on the user’s device, giving them greater privacy and control while removing reliance on external servers. This makes LM Studio different from many competitors that require cloud access or single-framework support, as it emphasizes local execution, privacy, and broad compatibility. The goal is to help developers, researchers, and enthusiasts use AI tools securely and efficiently, integrating local LLMs into their workflows and projects.
Industries
Data & Analytics
Consumer Software
Enterprise Software
AI & Machine Learning
Company Size
11-50
Company Stage
N/A
Total Funding
N/A
Headquarters
New York City, New York
Founded
2023
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Health Insurance
401(k) Retirement Plan
Remote Work Options
Paid Vacation
Flexible Work Hours
Wellness Program
Mental Health Support
Conference Attendance Budget
Stock Options
Company Equity
Professional Development Budget
Tuition Reimbursement
Meal Benefits
Phone/Internet Stipend
Home Office Stipend
Parental Leave
Family Planning Benefits
LM Studio launches Bionic: A new agent for open models. 11 Aug 2026 #ai #opensource #privacy #llm #productivity LM Studio has released Bionic, an agent built for real work and coding tasks that runs on open models - and can operate entirely on your own device. Local AI has quietly reached a point where it can be trusted with actual production work, not just experiments. Here's what Bionic brings to the table. Documents, slides, and PDFs. Bionic creates and edits documents directly within the conversation. Every change is saved automatically, so you can work freely without worrying about losing progress. Coding. Everything you'd expect from a coding agent, Bionic delivers. Just describe what you need, and it writes the code. On-device voice. You can talk to Bionic naturally, with speech recognized in real time. All audio processing happens locally - your voice never leaves your computer. Models download right inside the app. You can pull the latest local LLMs for either simple chat or complex agentic tasks. Under the hood, LM Studio's engine runs on MLX and llama.cpp. For tasks that demand more horsepower, you can spin up open-source frontier models in the cloud, such as GLM and Kimi. Servers are based in the US, with Zero Data Retention enabled by default - meaning no data is stored. Privacy is built into LM Studio's core design. Bionic is now available for Mac and Windows. You can try it at lmstudio.ai. As local models continue to close the gap with cloud-based frontier systems, tools like Bionic point to a future where powerful, private AI agents live directly on your hardware - no data pipeline to a third party required. For teams weighing privacy, cost, and control against raw capability, this release is worth a serious look. For AI Agents Read with AI. Short prompt for a summary, takeaways, and applying this to your task. Audio Version Alpha version: audio generated via local TTS, errors possible. Want to discuss your own task? Tell AImpress Ltd. about the workflow you want to improve. AImpress Ltd. will help you identify the practical next step.
Run Muse Glimmer locally. Aug 10, 2026 · LM Studio partnered with Meta to bring launch day support for Muse Glimmer in LM Studio Bionic! Muse Glimmer is a new 30B open-source model from Meta that excels in agentic tasks, and can fit right on your laptop. The model is available starting today for downloading and running locally in LM Studio Bionic. Muse Glimmer is available under the Apache 2.0 license. Run Muse Glimmer locally with Bionic. * Download and install LM Studio Bionic * Go to Settings | Explore to find Muse Glimmer from the catalog * Select the model and click Download Once the download finishes, Muse Glimmer will appear as a local model option for any existing or new session. Muse Glimmer for local agentic work. Muse Glimmer is a 30B-parameter open-source model optimized for local agent workflows. It is capable of multi-step task execution for open ended agentic tasks, can call tools, and supports image input. The model is a great fit for working with the Bionic agent locally. Bionic's friendly interface makes it suitable for both technical and non-technical users, while Muse Glimmer is optimized for running locally on both Mac and PC machines. Pairing the two together means you can run agentic workloads on your own hardware 24/7, without you having to worry about data leakage or token costs. BionicBench v0.1. To understand how well local models handle end-to-end agentic work, LM Studio is developing BionicBench: an internal evaluation of real-world workflows in Bionic. Its first iteration includes 18 tasks spanning coding and following repository instructions, handling text files and attachments, understanding images and screenshots, editing Word documents, PowerPoint decks, and Excel spreadsheets, and generating PDFs. The cases test whether an agent can follow detailed instructions, create or modify the requested artifact, and preserve formatting and file structure where required. Using the latest graded outcome for each task under the same local test setup, Muse Glimmer completed 83.3% of tasks. This was ahead of Gemma 4 31B and Qwen 3.6 27B, both at 77.7%, demonstrating strong breadth for a model that runs on your own hardware. BionicBench v0.1 results for Muse Glimmer Here are some ways to use Muse Glimmer with Bionic: 1. Deep research agent, right on your Mac or PC. Research a topic with Bionic and preview the PDF summary right in the app Ask Bionic to research a topic and leverage its native web search functionality to gather current information from the web. Point Bionic to a local folder or drag in files to give Bionic context, then ask it to summarize documents or produce detailed research reports. Bionic's agentic document understanding allows you to work seamlessly across large numbers of files and a variety of sources. 2. Use Muse Glimmer to convert an outline into a slide deck. Use Muse Glimmer in Bionic as a thinking partner and ask it to help you shape thoughts into a clear outline. Then, ask Bionic to turn that outline into a polished deck right in the same session. Iterate and refine your work by asking the agent for changes, and easily rollback changes to previous versions with automatic checkpointing. 3. Run multi-step coding workflows. Create a new project in Bionic and toggle on Allow Coding to use Muse Glimmer with a fully local coding agent. Toggle on Allow Coding when creating a project to let Bionic run commands in a local directory Bionic is great at multi-step coding workflows, such as updating release notes on recent git commits, refactoring across files, or fixing failing tests. None of your data goes to third party providers and you pay no per token inference cost. 4. Automate everything with the Bionic local API and Muse Glimmer. The local API conforms to OpenAI's Chat Completions. This means you can point a wide variety of tools and frameworks to it seamlessly. From the terminal. lms server start From the app. Under Settings | Local Model API, toggle on the Local API server to point your projects at Bionic and Muse Glimmer. You can use this with automation scripts, or wire it into batch jobs and CI to have a capable agent running entirely on your hardware.
LM Studio Bionic now supports the Kimi K3 model by Moonshot AI. Jul 28th, 2026 4:36 AM EEST News LM Studio just added support for Moonshot AI and its flagship Kimi K3 model inside the new Bionic app. Bionic is a desktop agent designed for complex coding and document workflows. By bringing Kimi K3 into the mix, users get access to a massive 2.8 trillion parameter model that can handle huge amounts of information at once. The update gives developers and researchers a powerful new option for long-horizon tasks. Video Muted Kimi K3 brings new features to the local desktop agent. The integration into Bionic gives users several major upgrades for their daily workflows. Moonshot AI built this specific model to handle large amounts of data without losing track of the original instructions. People using Apple computers can run the Bionic app to take advantage of these new capabilities right now. Don't miss the best of The Mac Observer. Set The Mac Observer, Inc. as a preferred source and its Apple reporting ranks higher in your Google Search results and Discover feed - one tap, no account changes. Or get it by email Here are the main features included with the Kimi K3 integration: * A 1 million token context window for reading entire code repositories or large document folders at the same time. * Native vision support that allows the agent to see and understand dropped images or screenshots. * Adjustable reasoning levels that let you choose exactly how much effort the model spends on complex problems. The cloud model protects user privacy with strict data rules. Even though Bionic is known for running models locally on your macOS hardware, it offers Kimi K3 as a cloud option for heavier workloads. Since this model is so large, running it in the cloud makes it much easier to use on standard desktop machines without slowing them down. LM Studio set up its inference servers in the United States to handle the processing. The company enabled Zero Data Retention by default for this new integration. This means it will not save your files or use any of your private information to train future versions of the software. You can safely drop in your private code or research documents without worrying about privacy leaks. The addition of Kimi K3 makes Bionic a much more capable tool for developers and writers who need help organizing large projects. Users can download the latest version of the app today to start testing the open model for themselves.
LM Studio shipped Bionic, an agent built for open models. The cloud tier riding shotgun is the tell. July 17, 2026 · News Tl;dr. On July 16, LM Studio released Bionic, a standalone desktop agent it bills as "the AI agent made for open models, built to get things done." It inspects and edits codebases with inline diffs, works over PDFs, slide decks, and spreadsheets in a sandbox with automatic checkpoints, takes voice input through Voxtral running entirely on-device, and has web search built in. You can point it at models on your own GPU, at your own hardware remotely via LM Link, or at a brand-new hosted tier called LM Studio Secure Cloud that serves frontier open models under negotiated zero-data-retention terms. Preview builds are live for Windows x64 and Apple Silicon Macs. The agent is the headline. The cloud tier is the strategy. An agent, not another chat window. LM Studio has spent years as one of the most widely used desktop apps for downloading and running open-weight models locally. Bionic is a different product: a separate app where the model does things instead of just saying things. The launch post organizes those things into a few buckets. * Code: inspect a local codebase, explain it, make edits with inline diffs, and run agentic code search across it. * Documents: work over PDFs, presentations, and spreadsheets inside a sandboxed environment, with automatic checkpoints so a bad agent step can be rolled back. * Voice: what LM Studio calls "state-of-the-art local voice transcription," powered by Mistral's Voxtral and running fully on-device. * Web search: native, for research tasks that need sources fresher than the model's weights. The obvious reference points are Claude Code and the wave of open harnesses like OpenCode. Bionic's pitch is that none of those were designed around open models as the default brain, and that the team shipping the most popular local runtime is best placed to build the harness that treats a quantized model on your own GPU as a first-class citizen rather than a fallback. Three lanes for your tokens. The architecture question that matters for this crowd is where inference happens, and Bionic gives you three answers. Fully local, with models running on your machine. Through LM Link, LM Studio's remote-access layer, so the agent can drive hardware you own from a device you are holding. Or through LM Studio Secure Cloud, a new hosted service offering what the company calls "the largest frontier open source models," with GLM 5.2 and Kimi K2.7 Code the names dropped in the launch post. On data handling, the launch post commits to "Zero Data Retention and never training on your data," and says cloud requests are processed transiently and not retained after the request completes. Pressed on the Hacker News thread, LM Studio's founder was more specific: "We negotiated ZDR with our providers. We consider that a condition to make things available to our users." That phrasing confirms Secure Cloud brokers to upstream inference providers rather than racks LM Studio runs itself. The cloud tier is the tell. If this move feels familiar, it should. Ollama, the other household name in local inference, raised a $65M Series B earlier this month that was widely read as funding for its cloud business. Now LM Studio ships an agent whose flagship models happen to be the ones too big for the audience's hardware, next to a hosted tier with an account system and billing. The local-runtime companies have all reached the same conclusion: the runtime is the funnel, the cloud is the product. The Hacker News reception split accordingly. Praise for the harness itself, with one commenter calling it "one of the better agent harnesses I've seen for inspecting reasoning chains." Skepticism about the trajectory, with another noting that every VC-backed local LLM startup eventually launches a cloud offering because that is the only path to venture-scale returns. And the perennial complaint, sharpened here by the branding: neither LM Studio nor Bionic is open source. It is an agent for open models in an app you cannot open. Why a local harness is genuinely hard. It would be easy to dismiss Bionic as a UI over an OpenAI-compatible API, and some commenters did. But agentic workloads are uniquely punishing on local hardware, and LM Studio has been laying pipe for this for weeks. A June update to its MLX engine added KV-cache checkpointing for long-context agentic workflows, which matters because an agent loop keeps returning to the same giant context over and over. The KV cache is the model's working memory of everything already in the context. Checkpointing it is like leaving a bookmark in an 800-page novel: without it, every agent turn starts the book again from page one, and on a home GPU that re-read is most of your wall-clock time. Chat apps can shrug this off. Agents that take fifty turns against the same repo cannot. What to watch before you commit. Some straight-faced caveats. This is an initial preview, and it shows: no Linux build at launch, which for a self-hosting audience is a strange first omission. Secure Cloud pricing was not published at launch; the flow just asks you to create an account and set up billing. The app is closed source, so the privacy story ultimately rests on the company's word and its provider contracts rather than anything you can audit. And the quiet tension in the product remains: the models the launch post celebrates, GLM 5.2 and the Kimi K2.7 line, are far too large for most home rigs. Run Bionic fully local and you are in small-model territory, where agentic reliability is still hit-or-miss. Run the models it actually advertises and you are a cloud customer. That is not a scandal, but it is the business model wearing a local-first jacket. Key takeaways. * LM Studio launched Bionic on July 16: a standalone agent app for open models covering code edits with inline diffs, sandboxed document work with automatic checkpoints, on-device Voxtral voice input, and built-in web search. * It runs three ways: fully local, against your own hardware via LM Link, or on the new LM Studio Secure Cloud hosting frontier open models like GLM 5.2 and Kimi K2.7 Code. * The company commits to zero data retention on cloud requests, negotiated with upstream providers, and says it never trains on your data. * Preview builds cover Windows x64 and Apple Silicon Macs only; no Linux, no published cloud pricing, and neither app is open source. * Like Ollama's cloud-flavored Series B, Bionic signals where local-AI companies see the money: the runtime is the funnel, the hosted tier is the product. AI LM Studio local AI agents open models GLM 5.2 Kimi self-hosted AI Bacon Weekly Liked this? Get smarter about AI, weekly. The model launches that matter, the local-LLM tricks worth stealing, and the tools actually shipping. Hand-curated into one email a week, cut down to signal. No hype, ever. Frontier models Local LLMs Builder tools Zero hype Free and curated. One tap to unsubscribe, always. Error: Domain verification failed (missing Origin header).
Improving LM Studio's MLX Engine for agentic workflows. Jun 5, 2026 · LM Studio recently released mlx-engine v1.8.5 in LM Studio. This update dramatically improves performance for repeated, long-context agentic workflows by checkpointing your KV cache. It also adds continuous batching for VLM requests. This work is open source; you can view the PR here. In this post, I'll explain the cache-reuse problem this solves, why current open-source LLM models make rewinding harder, and how the new disk-backed cache works. Its benchmarks show up to 80% lower extra RAM usage, up to 2x more throughput, and up to 3.5x faster processing for image requests. Adrien uses the new mlx-engine for a local review of a URL-shortener app with codex -oss. What is mlx-engine? MLX Engine (mlx-engine) is an MIT-licensed inference engine optimized for Apple silicon. It was created and maintained by LM Studio. It uses Apple's MLX machine learning library, and builds on projects such as mlx-lm and mlx-vlm. MLX Engine is LM Studio's backend for all MLX inferencing. Current model architectures, and the shortcomings of mlx-engine. Two of the most popular open-source models right now are Qwen 3.5 (and 3.6) and Gemma 4. As part of each model's architecture, they use some nifty tricks to reduce the size of the KV cache at large context lengths. Qwen 3.5 uses a hybrid architecture and Gemma 4 uses a sliding window architecture. These attention strategies reduce memory usage at large context lengths, but they make the KV cache not arbitrarily rewindable. Let's walk through how Gemma 4 handles inference. This example is focused on Gemma 4 E2B; it interleaves "local" attention layers (sliding window of 512 tokens), and "global" attention layers. Gemma 4 interleaves local and global attention layers. Rewinding after a reasoning-heavy agent turn can leave parts of the local KV cache missing. Step 1: Prompt prefill. Compute the KV cache for the system prompt and user message. Step 2: Decode. Build up the KV cache while computing the assistant reasoning content and assistant message. Step 3: Rewind. Trim the KV cache to step (1) and append the assistant message without the prior reasoning content. So, a key problem that inferencing engines have to solve is avoiding re-computation when rewinding the KV cache to prepare a follow-up response. How LM Studio improved prompt caching in mlx-engine. LM Studio devised a solution for KV cache rewinding for these agentic use cases. By saving and restoring prompt cache to disk, the KV cache for follow-up requests does not need to be recomputed. Saving the KV cache to disk. Copying and storing these KV caches at 256-token boundaries lets LM Studio restore exact cached prefixes when the corresponding KV cache tensors are still present. If part of the prompt was edited, never computed, or evicted from the disk cache, mlx-engine falls back to recomputing that suffix. 256 tokens is small enough to avoid wasting much work on recomputation, while large enough to keep the disk cache efficient. mlx-engine saves KV cache records at fixed 256-token boundaries, then restores the longest available cached prefix for follow-up requests. First, at every boundary of 256 tokens (sequence len % 256 == 0), stream a copy of the local attention layers' KV cache to a disk-writer backend. While the model is processing the prompt or generating new tokens, a background disk-writing process is running. At every 256 token boundary, the system copies the KV cache corresponding to the most recent 256 tokens and sends it to the disk writer which then persists that block to disk. Since Apple silicon has a unified-memory architecture, LM Studio commit the local attention KV cache to disk and evict it from memory. This ensures that mlx-engine's memory usage footprint scales with active sequences, rather than all previously seen sequences. Restoring the KV cache from disk. First, calculate a key for each block of 256 tokens. Then, determine which global and local KV cache blocks need to be retrieved. Using the list of keys and cache types for the prompt, load as much KV cache as LM Studio can from the disk. For prompt sections that never had their KV cache computed (or had their KV cache evicted from disk), schedule those sections for prompt prefill. The disk cache is an LRU store, so whenever LM Studio save to or load from its disk store, the store evicts the least recently used KV cache tensors. This ensures that its disk store optimizes for the usage pattern. If the engine is sent short prompts using the same system prompt, the system prompt's local attention KV cache will not be evicted, but the KV cache of stale conversations will be evicted. And, if the engine is only receiving requests for one ever-growing conversation, the earlier local attention KV cache will get evicted in order to make room for the longer global attention KV cache. Disk cache design. LM Studio designed the disk cache to clean itself up after the model is unloaded. In other words, the cache is temporary and will not leave persistent files. The disk cache is one scratch file, not a folder full of independent cache files. LM Studio pack many cache records into that one file. Each KV cache entry is a serialized safetensors blob, and the engine keeps an in-memory table saying: "entry X starts at byte offset Y and is Z bytes long." When KV cache entries are evicted, their byte ranges are returned to a free list and reused by later records; if free space reaches the end of the file, the file is shrunk. LM Studio make the disk cache temporary by using the operating system's temporary-file mechanism in /tmp, and by treating all lookup metadata as model-lifetime state only. On model unload, the cache store clears its in-memory index and closes the scratch file. If the model process exits, the OS closes the file handle and releases the storage. And, continuous batching. LM Studio also added continuous batching to its vision model runner. Plenty of ink has already been spilled on the implementation and benefits of continuous batching; Hugging Face has a great explainer. Continuous batching allows users to use the same model for concurrent request processing. Along with the KV cache improvements described earlier, mlx-engine can now be used for serious agentic workloads. Benchmarks. To make the performance improvements more concrete, LM Studio ran a few end-to-end LM Studio API benchmarks on an M3 Max MacBook Pro with 36 GB of RAM, using lmstudio-community/Qwen3.6-27B-MLX-4bit. These benchmarks focus on the workloads that this update is intended to improve: parallel chat, long-prompt processing, and repeated high-resolution image prompts. Benchmark: parallel chat throughput. Setup: The model was loaded with parallel=4, then four short chat requests were sent concurrently through the LM Studio API. Each response was allowed to stop naturally. Parallel chat throughput mlx-engine Output tokens End-to-end output tok/s Total tokens End-to-end total tok/s 2.2x faster Result: for this four-way parallel chat workload, mlx-engine v1.8.5 completed the run about 2.2x faster end-to-end, with nearly identical output token counts. Benchmark: memory under parallel long prompts. Setup: The model was loaded with parallel=4, then four large prompts were sent concurrently through the LM Studio API. RAM usage was measured after the model loaded and again after the run completed. Memory under parallel long prompts mlx-engine Input tokens Output tokens Total tok/s RAM after load RAM after run Extra RAM after run 82% less extra RAM Result: for this parallel long-prompt workload, mlx-engine v1.8.5 used about 82% less extra RAM after the run, while maintaining similar wall-clock time and slightly higher total token throughput. This is the expected benefit of moving inactive prompt-cache records out of unified memory. The active sequences still need to stay resident, but stale cache records no longer have to keep accumulating in RAM. Benchmark: repeated high-resolution image prompt. Setup: The same image prompt was sent twice, generating one token per request. This isolates the cost of processing the image-expanded prompt and restoring the prompt cache. mlx-engine Cached prompt tokens Uncached prompt tokens
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Consumer Software
Enterprise Software
AI & Machine Learning
Company Size
11-50
Company Stage
N/A
Total Funding
N/A
Headquarters
New York City, New York
Founded
2023
Find jobs on Simplify and start your career today