
Work Here?
Ollama provides software to run large language models locally on Windows and other platforms. It enables deployment, management, and GPU-accelerated inference for models like Llama 2 and Code Llama, and offers compatibility with the OpenAI Chat Completions API to connect local models with familiar tools. It differentiates itself by focusing on local, private deployment with a model library and licensing/subscription access, rather than cloud-only solutions. The goal is to let users run advanced AI models on their own hardware, reducing reliance on cloud services.
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
51-200
Company Stage
Series B
Total Funding
$65.1M
Headquarters
Palo Alto, California
Founded
2021
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$65.1M
Above
Industry Average
Funded Over
2 Rounds
Industry standards
AI developer digest - Sep 24, 2026. Fabio Pacifici Full-stack developer & educator Local models got more useful this week. Ollama shipped two quality-of-life changes developers have been asking for, Perplexity brought its agentic "Computer" fully on-device to Windows, and there's a paper that quietly reframes how Fabio Pacifici should think about what an LLM is. Let's go. 1. Ollama now streams tool calls. Ollama added streaming responses with tool calling. The old behaviour forced you to wait for the full generation, then parse the JSON to figure out whether the model wanted to call a tool or just return text. The new parser reads each model's template directly to understand tool-call structure as it streams, so your chat applications can render content and fire tool calls in real time instead of stalling at the end of the response. Why it matters: this is the difference between an API that's usable for agent loops and one that feels like a batch job. Streaming tool calls make local agent orchestration feel native - and the parser handles models that weren't even explicitly trained on tool tokens, which is a nice hedge for the wide zoo of open models on Ollama. 2. Turn thinking on and off. Ollama also added a think parameter (and -hidethinking / /set nothink in the CLI) so you can toggle a model's reasoning pass per application. Want the 8B model to mull over multiple viewpoints before answering? Enable thinking. Need sub-200ms latency for a UI autocomplete? Disable it. For scripting, -hidethinking strips the thinking trace and returns only the answer. Why it matters: as reasoning models proliferate, having explicit control over when the "think" pass runs - and hiding it from the output stream when you don't need it - is a real workflow lever. It's a small API detail with big UX impact for anyone building chat or agent interfaces. 3. Perplexity's local agent lands on Windows RTX. Perplexity shipped "Portable Computer" in its Windows app for GeForce RTX and RTX PRO systems (24GB+ VRAM). It's a local version of the Computer agent that plans and executes multistep tasks using on-device models (e.g. Qwen 3.8 27B, post-trained and RTX-optimised), keeps sensitive data on the machine, and only asks before sending anything to the cloud for heavier reasoning. Connectors for Outlook, Drive, Gmail, Slack and GitHub round it out. Why it matters: this is the on-device agentic trend maturing - real multistep agents running fully local, no per-query credits, private by default. For developers it signals where local inference and agent orchestration are converging: the PC as a first-class agent runtime rather than just a cloud client. 4. A probabilistic lens on LLMs (arXiv). "The Probabilistic Structure of Large Language Models" gives a self-contained account of LLMs as probability measures over token sequences - training as maximum-likelihood estimation, generation as sequential simulation of that stochastic process - and ties the asymmetry of KL-divergence to familiar failure modes like hallucination and the gap between statistical plausibility and truth. Why it matters: most working material treats models either as magic or as black boxes. A clean probabilistic framing is genuinely useful for reasoning about why models deviate from the training distribution, and it connects nicely to diffusion models for engineers who think in terms of score functions. Worth the read. What to watch. Local tool-calling and local agents are both getting good enough to build real products on. The interesting question over the next few weeks: which app patterns actually benefit from staying on-device, versus the ones that are better served by cloud reasoning? Building something local? The tooling has never been closer to turnkey. Sources. * Ollama - "Streaming responses with tool calling" - https://ollama.com/blog/streaming-tool - May 28, 2025 * Ollama - "Thinking" - https://ollama.com/blog/thinking - May 30, 2025 * NVIDIA - "Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX" - https://blogs.nvidia.com/blog/local-ai-perplexity-windows-pcs/ - Sep 2026 * arXiv - "The Probabilistic Structure of Large Language Models" - https://arxiv.org/abs/2609.25134 - Sep 2026
Ollama v0.34.0 brings its models to ChatGPT Desktop. 2026-09-09 Ollama has released v0.34.0, adding a way to use Ollama models directly in ChatGPT Desktop. The release also updates several parts of its OpenAI-compatible interface. The announcement is documented in the v0.34.0 release notes. Setup is available through the Ollama app on MacOS. The update also improves structured output performance on Apple Silicon, adds support for client tool search and response compaction, and fixes image handling through compacted responses. For people making things, the main change is workflow flexibility. You can keep using ChatGPT Desktop as the place where you write prompts, review responses, and work through a project, while selecting Ollama models for supported tasks. That reduces the need to move text between separate applications or build a custom interface before testing an open model. The release is especially relevant for creators who already use structured output. Structured output is useful when a model needs to return predictable fields instead of free-form prose. A content pipeline might ask for a title, summary, tags, and scene list in a fixed format. A design tool might request layout values or a list of assets. Better performance on Apple Silicon should make those interactions more responsive on compatible Mac hardware, although the announcement does not provide benchmark results. The OpenAI-compatible updates are useful for developers building around a common client interface. Tool search can help a client find or select tools during an interaction. Response compaction is intended to reduce the amount of conversation context that must remain active as a session grows. Together, these additions matter when a project involves repeated model calls, external actions, or long-running sessions. They may also make it easier to test an Ollama-backed workflow without rewriting every integration around a different API shape. Image support is another practical detail. The release notes say images now work correctly through compacted responses. That matters for applications that combine visual inputs with longer conversations. For example, a maker could keep an image review session going while the system compresses earlier context, rather than treating image handling as a separate path that may break when the response is compacted. On Mina Labs, Ollama v0.34.0 is available for 8 per image. That gives makers a direct way to try the release from the platform rather than setting up a local environment first. The listing is the useful starting point if you want to evaluate the model behavior, test a prompt flow, or compare the output with an existing project. The price token refers to one image; the release itself includes broader text, tool, and client-integration changes. Minaxlab would use this release first for structured creative briefs. A prompt could request a fixed object containing a project concept, target audience, visual direction, shot list, and production notes. The resulting fields could then feed a second step in a workflow. The performance improvements on Apple Silicon would be useful when iterating over many versions of the same brief locally. Minaxlab would also test it for image-aware review sessions and tool-driven production tasks. An image could be submitted for analysis, followed by requests for alternate compositions, captions, or asset metadata. Tool search and response compaction could support a longer session where the model needs to find an available action and retain the important decisions from earlier turns. These are practical tests for the release because they exercise the specific areas Ollama has changed: structured output, compatible clients, compacted responses, and images.
Understanding ollama: the complete guide to running Self-Hosted Local AI & open-source LLMs securely. WebVorta Tech Team 30 Aug 2026 Artificial Intelligence (AI) has fundamentally transformed enterprise workflows, powering automated analytics, customer engagement, and software engineering. However, relying entirely on proprietary cloud AI APIs like OpenAI ChatGPT or Anthropic Claude presents two formidable obstacles for growing enterprises: exponentially escalating API token costs and stringent data privacy & compliance risks. In response, forward-thinking organizations are embracing Private AI and Self-Hosted Local LLMs (Large Language Models). At the vanguard of this movement is the world's most popular open-source LLM runtime: Ollama. What is ollama? Ollama is a lightweight, extensible open-source framework designed to let developers and enterprises bundle, run, customize, and orchestrate Large Language Models locally on personal computers, dedicated GPU workstations, or private cloud servers. If Docker revolutionized software development by standardizing application containerization, Ollama is the "Docker for AI Models". It eliminates the complexity of Python virtual environments, CUDA driver compilation, and manual model quantization. With a single terminal command: ollama run llama3.3 Ollama downloads the weights, optimizes memory allocation across GPU VRAM and CPU RAM, and spins up an interactive AI interface with a built-in REST API in minutes. The ollama phenomenon: Surpassing 100,000 GitHub stars, Ollama has become the de facto foundation for enterprise internal AI assistants, private document RAG pipelines, and automated agent workflows worldwide. Why is ollama the leading choice for enterprise Private AI? * 100% Data Sovereignty & Confidentiality: Since all inference executes locally within your private firewall, proprietary intellectual property, medical records, financial spreadsheets, and internal chats never touch external cloud servers. * Zero Token & Subscription Fees: Run millions of queries, document extractions, and agentic workflows without worrying about unpredictable per-token monthly billing. * Offline & Low-Latency Execution: Operates completely detached from internet connectivity, providing ultra-responsive sub-millisecond network latency over local area networks (LAN). * Standardized OpenAI-Compatible REST API: Ollama exposes a native REST API on port 11434 that adheres to the /v1/chat/completions format, integrating out of the box with Open WebUI, n8n, Dify, LangChain, and custom enterprise web applications. * Customizable Modelfiles: Package custom system instructions, temperature thresholds, and domain-specific knowledge into version-controlled Modelfile definitions. Top open-source LLM architectures supported by ollama. Ollama supports the world's most capable open-source models, delivering performance competitive with proprietary closed APIs: * Llama 3.3 & Llama 3.1 (Meta AI): The gold standard for general intelligence, complex multi-step reasoning, 128K context window comprehension, and multilingual tasks. * DeepSeek-R1 & DeepSeek-V3: State-of-the-art reasoning models exhibiting exceptional proficiency in mathematics, logical chain-of-thought analysis, and automated coding. * Qwen 2.5 (Alibaba Cloud): Top-tier benchmark leader for structured JSON data extraction, technical document comprehension, and concise instruction adherence. * Mistral & Mixtral (Mistral AI): Renowned for blazing-fast inference speeds and innovative Mixture of Experts (MoE) efficiency. * Nomic Embed & BGE: High-density embedding models optimized for enterprise document vector search and Retrieval-Augmented Generation (RAG). Hardware sizing guide for running ollama. Through modern 4-bit and 8-bit GGUF quantization, open-source models run smoothly across consumer and enterprise hardware configurations: | Model Class | Representative Models | Minimum Specifications | Optimal Recommended Setup | | Compact (1B - 3B) | Llama 3.2 3B, Qwen 2.5 1.5B, Phi-3 Mini | 8 GB System RAM (CPU only) | 16 GB RAM / Apple M1/M2/M3 Silicon | | Standard (7B - 8B) | Llama 3.1 8B, Mistral 7B, Qwen 2.5 7B | 16 GB System RAM (CPU) | NVIDIA 8GB VRAM (RTX 3060/4060) or Apple Silicon | | Mid-Range (14B - 32B) | Qwen 2.5 14B/32B, DeepSeek-R1 14B/32B | 32 GB System RAM | NVIDIA 16GB - 24GB VRAM (RTX 4090 / RTX 3090 / A4000) | | Enterprise Tier (70B) | Llama 3.3 70B, Qwen 2.5 72B, DeepSeek-R1 70B | 64 GB System RAM | Multi-GPU (2x RTX 4090 / 2x RTX 3090 / A100 80GB) | High-Value enterprise use cases for local ollama deployments. * Confidential Document Intelligence (Private RAG): Index thousands of proprietary company reports, NDAs, clinical trials, or compliance handbooks for instant, secure employee semantic search. * Automated Agent Workflows with n8n: Power background n8n automation nodes for invoice parsing, lead qualification, and email drafting without variable token costs. * Private Developer AI Copilot: Connect local IDE plugins to your private Ollama instance, keeping proprietary source code strictly on-premise. Ready to deploy Private AI & local LLM infrastructure for your enterprise? WebVorta delivers turnkey Private AI solutions: from enterprise GPU server configuration, secure Ollama deployment, custom document RAG integration, to bespoke web application interfaces. Pertanyaan sering diajukan (FAQ). Is Ollama completely free for commercial enterprise use? Yes. Ollama is licensed under the permissive MIT open-source license, allowing unrestricted commercial, corporate, and private deployment without software licensing fees. Can Ollama run effectively on servers without dedicated GPUs? How easily can Webvorta integrate Ollama with its existing custom software?
Cut AI Agent costs by 80%: the n8n + Ollama local stack. What are You Looking For? August 24, 2026 Local AI agents with n8n + Ollama: cut cloud costs 80% and keep data in-house. Local n8n and Ollama stack reduces AI costs. By Andres SEO Expert. Key takeaways. * Self-hosted AI agents eliminate per-token cloud API charges. * n8n orchestrates and Ollama serves models locally inside your perimeter. * Hybrid routing cuts costs 60-80% while preserving data sovereignty. Table of contents. Local agents cross the production threshold. Autonomous AI agents that run entirely on an organization's own infrastructure have moved from architectural theory to production practice in 2026. n8nlab.io published a technical implementation guide that pairs n8n's orchestration engine with Ollama's local model serving to create a closed environment that keeps every request inside the server perimeter. Open-weight models including Llama 3.1, Mistral, Qwen, and DeepSeek execute locally under this design. The immediate effect is the elimination of metered cloud API charges for workloads routed through the stack. The architecture answers two operational pressures. Under the implementation documented by n8nlab.io, high-volume classification and extraction tasks stop accumulating per-token fees. Regulated or proprietary information no longer passes through a third-party model provider. The closed-loop orchestration stack, layer by layer. In this deployment model, n8n acts as the control plane while Ollama operates as the local inference layer. The data flow moves through five stages: n8n catches inbound payloads, structures the system prompt, calls the local model, evaluates the returned response, and updates downstream systems. The execution path is deliberately linear, but the routing logic around it is where the architecture earns its production credentials. * Trigger: A webhook, schedule, or application event starts the workflow. * Orchestration: n8n structures the inbound data and prepares the exact system prompt. * Model serving: n8n calls the Ollama API, either through its native integration or an OpenAI-compatible shim. * Token generation: Ollama processes the prompt locally against the loaded open-weight model. * Response routing: n8n applies conditional logic and writes the output back to internal systems such as Slack, a database, or a CRM. Two paths to the Ollama API. The native Ollama Chat Model node is the recommended path for new builds. It maps directly to Ollama's API structure and supports AI Agent and Basic LLM Chain configurations without translation layers. For teams migrating existing OpenAI-based workflows, the second path matters more. Ollama exposes an OpenAI-compatible endpoint at the /v1 suffix, allowing the standard OpenAI Chat Model node to treat the local server as a drop-in replacement. Both paths require precise configuration. The base URL must point to the correct Ollama port, the model string must match the exact pulled tag, and the temperature should sit low enough to keep structured business tasks deterministic. * Base URL: Typically localhost:11434 for same-host deployments, or a server IP when n8n runs in a separate container. * Model: Must match the exact pulled tag, such as llama3.1. * Temperature: Set at 0.1 to 0.2 for deterministic output in structured automation tasks. Matching model tier to task reality. Hardware provisioning is the silent failure point of self-hosted AI. The sizing matrix maps model classes to RAM requirements, warning that oversized models on undersized machines trigger memory swapping, latency spikes, and deployment failure. * 7B to 8B models: 8 to 16GB of RAM, suited for classification, routing, and summarization. * 14B models: 16 to 32GB of RAM for advanced extraction and structured formatting. * 70B-plus models: 64GB or more of RAM, or dedicated GPU, for multi-step agentic reasoning. The capability trade-off is explicit. Cloud frontier models still hold an edge in complex logic, advanced tool calling, and strict adherence to nested JSON schemas. Local models win on bounded, repeatable work: categorical classification, unstructured data extraction, summarization, and processing highly sensitive material. Hybrid routing eliminates the either-or mentality. The most operationally valuable pattern is the hybrid routing architecture. A Switch node immediately downstream of the trigger evaluates the payload along two dimensions: data sensitivity and task complexity. Payloads flagged as sensitive or classified as simple route to the local Ollama branch. Complex multi-step analysis routes to the cloud tier, where Claude or GPT handles the reasoning load. This is not a migration strategy; it is a standing architecture. High-volume basic tasks avoid per-token charges, sensitive data stays inside the perimeter, and frontier compute is reserved for the work that genuinely needs it. Economic and sovereignty pressures rewrite the automation playbook. The vendor-reported estimate puts typical monthly AI API cost reduction at 60 to 80 percent for organizations routing 10,000 or more basic operations per day to local infrastructure. That range has not been independently audited at production scale, but it reflects the structural economics of moving token generation from metered cloud APIs to fixed-cost local hardware. The deeper shift is in cost architecture. Once classification and extraction workloads leave per-token pricing, the marginal cost of high-volume automations collapses toward infrastructure depreciation rather than usage metering. Data sovereignty carries equal weight. Compliance frameworks including HIPAA, SOC 2, and GDPR impose hard boundaries on where regulated data can transit. Local execution removes the third-party model provider from the sensitive path entirely. The framework is unusually candid about where local models fall short. It treats cloud frontier models as necessary complements for reasoning-heavy loops, not as a universal substitute. That combined posture changes how automation teams evaluate infrastructure. The relevant question is no longer whether a workflow can be self-hosted, but which tier each payload belongs in. Infrastructure control becomes the default posture. Self-hosted AI is no longer an experimental corner of automation. It is the default baseline for teams that handle regulated data or high-volume structured automations. For teams building n8n and Ollama automation pipelines that need to scale without losing infrastructure control, programmatic SEO AI automation is how Andres SEO Expert approaches production-grade orchestration - contact Andres SEO Expert. Frequently asked questions. How do n8n and Ollama work together for local AI automation? What are the hardware requirements for running local open-weight models in n8n workflows? How can I migrate existing OpenAI-based n8n workflows to Ollama? What is hybrid routing in an n8n and Ollama architecture? How much can self-hosted AI reduce monthly AI API costs? Why does local execution matter for data sovereignty and compliance? When should local models be preferred over cloud frontier models? August 24, 2026
Ollama added Laguna to Apple GPUs. The macOS chat path still needs work. Poolside says Laguna XS 2.1 is compact enough for a Mac with 36 GB of RAM. Ollama's current Laguna XS 2.1 model page also warns that chat through ollama run or /api/chat may return empty output on macOS/Metal. The first claim concerns whether a model can fit on a machine; the second concerns whether a particular interface path returns usable work. That distinction changes today's decision for teams considering a local coding model. Ollama v0.32.4 provides a documented Apple-GPU route worth evaluating on one exact configuration, while the publisher's warning argues against standardizing the affected chat route. After reading, you should be able to separate the model, runtime, hardware backend, quantization, chat template, API endpoint, and observed response before deciding whether to evaluate now or wait. In its Laguna XS 2.1 announcement, Poolside describes a mixture-of-experts coding model with 33 billion total parameters and 3 billion activated for each token. A mixture-of-experts model routes each token through part of the network instead of using every parameter for every token. Poolside also says the model is compact enough for a Mac with 36 GB of RAM and supports a 262,144-token context window. Those specifications identify the model and its intended hardware range. They do not report memory use, speed, output quality, or chat reliability for Ollama on a particular Mac. The 3-billion-active figure also describes computation inside the model; it is not a 3-billion-parameter memory claim. Quantization creates another boundary. It stores model values at lower precision to reduce the resources needed to run them, with tradeoffs that depend on the format and implementation. Ollama's model page lists several Laguna variants, including q4_K_M, q8_0, bf16, and mlx-bf16, so the model name alone does not identify the build a team will load or how much memory it will require. Ollama's stable v0.32.4 release, published July 25, says it added Laguna support on Apple GPUs through the MLX engine. MLX is the execution layer Ollama uses here to run the model on Apple hardware. This release claim establishes a supported implementation path; it does not establish identical behavior across model variants, Macs, or interfaces. The merged Laguna MLX implementation adds model-loading and model-creation paths for Laguna XS 2, XS 2.1, and S 2.1. It handles dense and routed expert layers, mixed quantization, routing, fused projections, and focused tests while staying on maintained MLX operations. Those details explain what Ollama added below the interface layer. A separate MLX memory-residency commit configures Metal residency after the model weights materialize. It caps resident memory at the smaller of active model memory and the device's recommended working set, then warns and continues with pageable memory if setup fails. That behavior concerns how loaded weights use memory; it is neither a general speed guarantee nor evidence that a chat response will contain output. Ollama's model page says reasoning and tool calling use the built-in laguna template. In a chat route, the application sends role-based messages, the template converts that history into the model's expected input form, the runtime executes the model, and the API returns the result. Runtime support can therefore coexist with an interface failure that leaves the caller without usable content. The same page says chat works as expected on Linux/CUDA but may return empty output on macOS/Metal. It does not identify affected macOS versions, Apple chips, quantizations, or the frequency of the failure. The page also does not establish a root cause, so the warning cannot support broader claims about every Apple-GPU run or a particular faulty component. For Mac users, Ollama points to a Linux/CUDA host or /api/generate with "raw": true as alternatives. Raw generation is a different request path: it does not prove that role-based chat, the built-in template, multi-turn history, or tool behavior works on the warned route. It may help a team inspect model output while keeping the chat compatibility question open. A useful evaluation starts with an exact candidate rather than "Laguna on a Mac." Keep the model tag or digest and quantization fixed. Pin the Ollama version, record the Mac model, Apple chip, memory, and macOS version, and identify whether the request uses MLX on Apple hardware. Then choose the interface being evaluated: ollama run or /api/chat with the Laguna template, or /api/generate with raw output. For chat, require non-empty content that answers a representative coding request, then repeat the same request on fresh and already-loaded sessions. If the intended application uses multiple turns or tools, test those behaviors directly: history must remain coherent, a tool request must be parseable, and the model must use the returned tool result to continue. A single completion from a different endpoint cannot establish that path. Capture the complete API response and Ollama logs for each run, including any error and completion fields the endpoint returns. Treat an empty result as a failed run even when the request appears to finish. This keeps endpoint behavior visible instead of reducing the outcome to whether the model loaded. Test raw generation separately if it is useful to the workload. A usable raw result paired with an empty chat result shows only that the two paths behaved differently on that configuration. Its guide to production-shaped model comparison explains why the exact model, runtime, hardware, quantization, template, and interface path must stay attached to the result. A bounded evaluation makes sense when the team wants to measure Laguna's coding quality, memory use, and latency on one available Mac and can tolerate changing the interface or stopping when chat returns empty output. Keep the result attached to the exact quantization, runtime, hardware backend, and endpoint used. That work can show whether Laguna deserves another review without turning a successful sample into a shared standard. Wait to standardize when the intended workflow depends on macOS chat, preserved multi-turn history, tool calls, or unattended automation. Revisit the decision after the public warning changes, then rerun the same workload on the required route. A changed page or runtime version is a reason to test again; repeatable non-empty output and correct task behavior provide the evidence for the local decision. Compatibility is only the immediate issue. Once the route behaves consistently, teams still have to own updates, monitoring, recovery, security, and cost; its article on the open-model production ownership gap covers that broader deployment decision without folding it into this release. BaristaLabs can help through AI consulting or a focused local model compatibility review. The work compares one coding workload across the exact model, quantization, runtime, hardware backend, and interface path so the team can decide whether the route should remain in evaluation. Use Ollama v0.32.4 and Laguna XS 2.1 for a bounded Apple-GPU evaluation if that evidence is useful now. Wait to standardize the warned macOS chat route until the exact path repeatedly returns non-empty, correct results for the workload the team intends to run. Local model compatibility review Decide whether the route should remain in evaluation BaristaLabs can help reproduce one coding workload across the exact model, quantization, runtime, hardware backend, and interface path, then define the evidence needed for standardization. Best fit for teams deciding whether a local coding model route is evaluation-only or dependable enough to standardize. Turn this idea into a pilot Which workflow should go first? Use the readiness check to compare impact, effort, risk, owner, and next step before booking a call. * 3-5 minutes * Deterministic score * No sensitive data July 27, 2026
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
51-200
Company Stage
Series B
Total Funding
$65.1M
Headquarters
Palo Alto, California
Founded
2021
Find jobs on Simplify and start your career today