Weaviate

Weaviate

Open-source vector database for semantic search

Overview

Weaviate provides an open-source vector database for AI-powered apps. It stores data objects together with their vector embeddings, enabling fast semantic and similarity searches across text, images, and audio. The system is cloud-native and modular, letting developers plug in models from providers like OpenAI, Cohere, and Hugging Face, and it supports hybrid vector/keyword search plus RAG features. Weaviate offers both a self-hosted option and managed cloud services (Weaviate Cloud and Enterprise Cloud) to scale with demand.

About Weaviate

Simplify's Rating
Why Weaviate is rated
B-
Rated B on Competitive Edge
Rated B on Growth Potential
Rated C on Differentiation

Industries

Data & Analytics

Enterprise Software

AI & Machine Learning

Company Size

51-200

Company Stage

Series B

Total Funding

$67.7M

Headquarters

Amsterdam, Netherlands

Founded

2019

Get referred to Weaviate

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Ricoh invested March 13, 2026, validating enterprise AI and Japan expansion.
  • Weaviate claims 1.6 million downloads, signaling strong community pull into 2026.
  • Free tiers, Query Agent, and Agent Skills lower adoption friction for new developers.

What critics are saying

  • CVE-2026-11500 exposed authorization bypass in 1.37.7; patching pressure hit June 2026.
  • Pinecone, Milvus, and Postgres extensions commoditize vector search and squeeze cloud pricing.
  • If open-source adoption outpaces monetization, Weaviate Cloud becomes a feature, not a business.

What makes Weaviate unique

  • Weaviate 1.38 shipped MCP Server and HFresh disk vectors on June 25, 2026.
  • Engram adds managed agent memory, extending Weaviate beyond plain vector retrieval.
  • Open-source core plus Weaviate Cloud and free tiers creates developer-to-paid upgrade path.

Help us improve and share your feedback! Did you find this helpful?

Funding

Total Funding

$67.7M

Above

Industry Average

Funded Over

3 Rounds

Series B funding is typically for startups that have proven their business model and need more funding to expand rapidlyβ€”often by entering new markets or adding more products. Investors are usually venture capital firms that specialize in later-stage investments.
Series B Funding Comparison
Above Average

Industry standards

$35M
$45M
Linktree
$50M
Weaviate
$65M
Substack
$100M
ClickUp

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Company Match

Remote Work Options

Home Office Stipend

Company Equity

Professional Development Budget

Growth & Insights and Company News

Headcount

6 month growth

↓ -6%

1 year growth

↓ -6%

2 year growth

↑ 5%
Apify
Sep 16th, 2026
What is Pinecone and why use it with your LLMs?

What is Pinecone and why use it with your LLMs? Pinecone is one of the best-known purpose-built vector databases, and it now positions itself more broadly as an AI knowledge platform. Here's what it does and when it earns its place in your stack. Sep 16, 2026 by What is the Pinecone vector database? In simple terms, Pinecone is a fully managed vector database. These days, Pinecone describes itself more broadly as an AI knowledge platform, with the vector database as the foundation alongside its Nexus and Marketplace products. By representing data as vectors, Pinecone can quickly search for similar data points in a database. That makes it a fit for retrieval-augmented generation (RAG) and agent memory, which is what most teams use it for today, as well as semantic search, similarity search across images and audio, recommendation systems, record matching, and anomaly detection. What are vector databases? Vector databases are designed to handle the unique structure of vector embeddings, which are dense arrays of numbers that represent meaning in text, images, audio, or video. They're used in machine learning to capture the meaning of words and map their semantic meaning. Vector databases index these representations so they can quickly compare them and retrieve the most similar results. That makes them useful for natural language processing, recommendation systems, semantic search, multimodal retrieval, and other AI-driven applications. Pinecone use cases. * RAG and question answering: retrieve relevant passages from a knowledge base before an LLM generates an answer * Semantic and hybrid search: find relevant content by meaning, keywords, or a combination of both * Recommendation systems: retrieve products, media, users, or other items that are similar to a query or existing item * Multimodal retrieval: search images and other content using vector embeddings * Matching and anomaly detection: identify similar records, duplicates, unusual items, or suspicious patterns Pinecone launched its vector database as a public beta in January 2021, straight into the generative AI boom, and became the best-known name in vector search. The category has since crowded. Qdrant, Weaviate, Milvus, and Chroma all compete for the same workloads, general-purpose engines like Elasticsearch and OpenSearch added vector search, and Postgres with pgvector now handles a large share of smaller deployments. In the beginning, most Pinecone use cases were centered around semantic search. Today, they have a broad customer base, from hobbyists interested in vector databases and embeddings to ML engineers, data scientists, and systems and production engineers who want to build chatbots, large language models, and generative AI models integration. It was obvious to me that the world of machine learning and databases were on a head-on collision path where machine learning was representing data as these new objects called vectors that no database was really able to handle. - Edo Liberty, founder of Pinecone Why use Pinecone with large language models? Perhaps the biggest use case for the Pinecone vector database is natural language processing (NLP) software, a category featured on Spotsaas. You can use Pinecone to build NLP systems that can understand the meaning of words and suggest similar text based on semantic similarity. That's why Pinecone is so useful for large language models. You can use Pinecone to extend LLMs with long-term memory. You begin with a general-purpose model, like GPT-4, but add your own data in the vector database. This process is essential when considering how to build your own LLM model, as it allows you to fine-tune and customize prompt responses by querying relevant documents from your database to update the context. You can also integrate Pinecone with LangChain, which combines multiple LLMs together. This is the main reason vector databases are all the rage these days. And while there are some excellent open-source alternatives, such as Weaviate, Milvus, and Chroma, which are also big players, Pinecone remains the leader in this field. Pinecone key features. * Fully managed: no infrastructure to run, and indexing happens automatically * Dense, sparse, and full-text indexes: semantic, keyword, and hybrid search in one database * Built-in embedding and reranking: Pinecone Inference generates embeddings and reranks results, so you don't need a separate provider * Namespaces: partition one index per tenant, user, or document set * Scales without re-architecting: from a free index up to dedicated read nodes, with backups, object-storage import, and a 99.95% uptime SLA on Enterprise * Runs where you do: AWS, Azure, and GCP, plus bring-your-own-cloud for teams that need the data in their own account How much does Pinecone cost? Pinecone has four plans, as of September 2026: * Starter: free, up to 2 GB of storage, one project, AWS Apify-east-1 only * Builder: $20 a month flat, for solo developers and small teams, with your choice of cloud and region * Standard: $50 a month minimum usage, then pay as you go, with a three-week trial that includes $300 in credits * Enterprise: $500 a month minimum usage, adding bring-your-own-cloud, private endpoints, audit logs, and a 99.95% uptime SLA On Standard, usage is billed at about $0.33 per GB of storage per month, $16 to $18 per 1 million read units, and $4 to $4.50 per 1 million write units, depending on cloud and region. Embedding and reranking through Pinecone Inference are billed separately. Pinecone is also available through major cloud marketplaces. Check Pinecone's pricing page before you budget, since its plans and pricing have changed more than once. Conclusion. If you're a developer working with generative AI (that's probably most of you now), learning how to use Pinecone and similar vector databases will certainly be worth your time. And if you need a web scraping tool to collect data for your vector databases, you might want to consider Website Content Crawler while you're at it. Get better data for AI Website Content Crawler was specifically designed to extract data for feeding, fine-tuning, or training large language models (LLMs) such as GPT-4, ChatGPT, or LLaMA

Weaviate
Jul 2nd, 2026
Native MCP + HFresh disk-based vectors are now GA.

Native MCP + HFresh disk-based vectors are now GA. Hello Weaviate Community! Weaviate is excited to share the Weaviate 1.38 release, bringing the HFresh disk-based vector index and built-in MCP Server to general availability. This update also rebuilds cluster-wide async replication and introduces previews for the Boost API and Nested Object Filtering. Latest AI & tech insights. Explore its recent Weaviate content: Read. * | Weaviate 1.38 Release: The full changelog and engineering context behind everything in this release. Read the blog * | Import & Vectorize Data with Weaviate at Scale: This guide covers the production patterns that keep imports fast and safe, like server-side batching, error handling, data type decisions, and ingesting media and PDFs without standing up an OCR pipeline. Read the blog Watch. Product highlights. * HFresh Vector Index (GA): Disk-based index that keeps memory low, especially for continuously changing streaming workloads, with built-in RQ-1 quantization. * MCP Server (GA): Native LLM/agent access to inspect schemas, search, and write data - no glue code required. * Async Replication, Everywhere: Re-architected to run cluster-wide from a single scheduler, rather than being configured and run separately. * Boost API (Preview) & Nested Filtering (Preview): Fine-tune rankings with query-time boosting and filter nested object properties using dotted paths. * πŸ†“ Weaviate Cloud free tier: A fully managed vector database that is free forever - no credit card required and no expiration. Comes with 100,000 objects, Query Agent, and Weaviate Embeddings built in. * | Engram is GA: Give your AI agents persistent, personalized memory that carries across sessions. Engram runs its own managed pipelines, so it works with your existing stack; no separate database required. Try it now * | New Support Portal: Weaviate launched support.weaviate.io, a single home for support across all its products. Open tickets through guided forms, track their status, and see your full interaction history all in one place! Ready to start building? Jump right in and spin up your free cluster with Weaviate Cloud. Or check its GitHub and star Weaviate while you're there. Company updates. Weaviate is excited to welcome Harneet! Join its team - Weaviate is hiring across various teams! Check out its career page for exciting opportunities in product, research, growth, and more. Hungry for more? Have a question or want to connect? Join its Weaviate Forum to engage in community conversations. See you in two weeks,

AISO Tools
Jul 1st, 2026
Weaviate review 2026: pricing, features, pros & cons.

Weaviate review 2026: pricing, features, pros & cons. Weaviate is the open-source vector database built around native hybrid search - combining vector similarity with keyword search in a single query. Here's an honest look at the new 2026 cloud pricing, self-hosting tradeoffs, and how it compares to Pinecone. Quick verdict. Overall Rating 100K objects $45/mo min Best for: Teams building RAG applications that need hybrid (vector + keyword) search and want the option to self-host without vendor lock-in. Less ideal for teams wanting the absolute simplest managed setup - Pinecone still wins there. What is Weaviate? Weaviate is an open-source vector database designed for AI applications, most commonly retrieval-augmented generation (RAG) pipelines. Its defining feature is native hybrid search: a single query can combine dense vector similarity search with sparse BM25 keyword search, weighted by an alpha parameter, so teams don't need to bolt a separate keyword search system onto their vector store. Because the core database is open source, Weaviate can be self-hosted via Docker or Kubernetes at no licensing cost, or run as a fully managed service through Weaviate Cloud. It integrates as a pluggable module with OpenAI, Cohere, HuggingFace, and other embedding and reranking providers, and ships with actively maintained Python, JavaScript/TypeScript, Go, and Java client libraries. In October 2025, Weaviate restructured its cloud pricing model entirely, replacing the old Serverless tier with a new Free / Flex / Premium structure billed transparently on vector dimensions, object storage, and backup storage. Weaviate pros & cons. Pros. * - Open source at its core: Weaviate can be self-hosted for free with the full feature set, giving teams a real exit from vendor lock-in that fully managed competitors like Pinecone don't offer * - Hybrid search built in: Weaviate natively combines vector similarity search with BM25 keyword search in a single query, weighted by an alpha parameter - useful for RAG systems where pure semantic search misses exact keyword or acronym matches * - Generous free cloud tier: The Free plan includes 100,000 objects, 1GB memory, and 10GB disk on an always-free cluster per user, enough to build and test a real RAG prototype before committing to a paid tier * - Transparent, dimension-based billing on paid tiers: Flex pricing is calculated from vector dimensions, object storage, and backup storage rather than opaque per-query pricing, making cost more predictable to forecast as usage scales * - Modular vectorizer and reranker modules: Weaviate integrates directly with OpenAI, Cohere, HuggingFace, and other embedding/reranking providers as pluggable modules, so you don't need to manage embedding generation in a separate service * - GraphQL and REST APIs with strong client libraries: Python, JavaScript/TypeScript, Go, and Java clients are actively maintained, and the schema/class model maps cleanly onto most application data models Cons. * - 2025 pricing restructure raised the entry point: Weaviate Cloud's paid tier starting price moved from $25/month under the old Serverless plan to $45/month minimum under the new Flex tier - a real cost increase for small production workloads migrating off the free tier * - Premium tier is a steep jump: after Flex ($45/month minimum), the next tier up (Premium, with 99.95% uptime and dedicated deployment options) starts at $400/month on a prepaid contract - there's a wide gap between hobby-scale and enterprise-scale pricing * - Self-hosting requires real operational expertise: running Weaviate yourself in Docker or Kubernetes for production means you own sharding, backups, and upgrades - the free-to-self-host framing understates the DevOps investment needed at scale * - Class/schema model has a learning curve: developers coming from a simple key-value or pure-vector mental model need to learn Weaviate's schema, cross-references, and module configuration before getting hybrid search and filtering working well * - Smaller managed-service footprint than Pinecone: Pinecone's fully managed simplicity and broader out-of-the-box framework integrations still win for teams that want zero infrastructure decisions and the fastest path from prototype to production Weaviate pricing 2026. Free. * - 100,000 objects * - 1GB memory / 10GB disk * - 1 collection * - 1 always-free cluster per user * - Best-effort availability Prototyping and small RAG projects before committing to a paid tier Flex. From $45/mo * - Unlimited objects * - Up to 1,000 collections * - Shared cluster with replication * - RBAC security * - 99.5% uptime SLA * - Pay-as-you-go dimension-based billing Production workloads that need real scale without a long-term contract Premium. From $400/mo * - Unlimited objects and collections * - Shared or dedicated deployment * - AWS, GCP & Azure coverage * - 99.95% uptime SLA * - Phone/Slack support + Technical Account Team Enterprise workloads needing dedicated infrastructure and faster incident response A Dedicated Enterprise tier is also available from $400/month prepaid with 99.95% uptime, 1-hour Severity 1 response, and HIPAA compliance on AWS. Weaviate can also be self-hosted for free. Weaviate vs Pinecone. | Feature | Weaviate | Pinecone | | Open source / self-hostable | | Fully open source | | Managed only, closed source | | Hybrid search (vector + keyword) | | Native BM25 + vector | | Sparse-dense hybrid supported | | Free tier | | 100K objects, always free | | Serverless free tier | | Entry paid price | $45/mo (Flex) | Usage-based, no fixed minimum on Serverless | | Enterprise tier | $400/mo+ (Premium/Dedicated) | Custom Enterprise pricing | | Setup complexity | | Schema/class model, more config | | Minimal config, fastest to production | | Vendor lock-in risk | | Low - can self-host anytime | | Managed-only, no self-host option | Who should use Weaviate? RAG apps needing hybrid search. Native vector + BM25 keyword search in one query makes Weaviate a strong fit for RAG systems where pure semantic search misses exact terms, product codes, or acronyms. Teams avoiding vendor lock-in. Because the core database is open source, teams can self-host on their own infrastructure at any point, giving an exit path that fully managed-only competitors don't provide. Cost-conscious startups building on the free tier. The always-free 100,000-object tier is real enough to build and test a production-shaped RAG prototype before committing to Flex or Premium. Enterprises needing dedicated deployment. Premium and Dedicated Enterprise tiers offer HIPAA compliance, choice of cloud provider, and 1-hour incident response for regulated or mission-critical workloads. Frequently asked questions. Is Weaviate worth it in 2026? For teams that want hybrid search (vector + keyword) built in and value the option to self-host without vendor lock-in, Weaviate is a strong pick - the always-free 100K-object tier is genuinely usable for prototyping. Teams that want the absolute fastest path to production with zero infrastructure decisions may still prefer Pinecone's simpler managed model, especially since Weaviate's October 2025 pricing restructure raised the paid entry point from $25/month to a $45/month minimum. How is Weaviate priced? Weaviate Cloud has three tiers: Free ($0, 100,000 objects, 1 always-free cluster), Flex (from $45/month minimum, pay-as-you-go based on vector dimensions, object storage, and backup storage, 99.5% SLA), and Premium (from $400/month prepaid, 99.95% SLA, shared or dedicated deployment). Weaviate can also be fully self-hosted for free via Docker or Kubernetes if you're willing to manage the infrastructure yourself. What is hybrid search in Weaviate? Hybrid search combines dense vector similarity search with sparse keyword search (BM25) in a single query, weighted by an alpha parameter you control. This matters for RAG systems because pure semantic search sometimes misses exact keyword matches, acronyms, or product codes that keyword search catches - hybrid search gets the best of both without running two separate systems. How does Weaviate compare to Pinecone? Weaviate is open source and self-hostable, giving teams an exit from vendor lock-in that Pinecone (managed-only, closed source) doesn't offer. Weaviate also has native hybrid search built into its core query model. Pinecone counters with a simpler, faster path to production and a broader out-of-the-box integration ecosystem across RAG frameworks. Cost-conscious teams comfortable with more setup often pick Weaviate; teams that want zero infrastructure decisions often pick Pinecone. Can I self-host Weaviate for free? Yes - Weaviate's core database is open source and can be self-hosted via Docker or Kubernetes at no licensing cost. The tradeoff is that you take on sharding, backup, and upgrade operations yourself, which requires real DevOps capacity at production scale. Weaviate Cloud's managed tiers exist specifically to remove that operational burden for a monthly fee. What changed in Weaviate's 2026 pricing? In October 2025, Weaviate restructured its cloud pricing entirely: old tier names were retired, the starting price for paid usage moved from $25/month (the old Serverless tier) to a $45/month minimum under the new Flex tier, and billing switched to a transparent model based on vector dimensions, object storage, and backup storage rather than opaque per-query pricing. Compare vector databases. See how Weaviate stacks up against Pinecone and other AI infrastructure tools.

Weaviate
Jun 18th, 2026
Free Tiers in Weaviate Cloud, new Playground demos and a Ricoh investment.

Free Tiers in Weaviate Cloud, new Playground demos and a Ricoh investment. Hello Weaviate Community! The best developer platforms don't just provide infrastructure - they make it easy to learn, experiment, and build. That's why this week Weaviate is introducing free tiers across Weaviate Cloud, launching the new Weaviate Playground with hands-on AI demos, and celebrating a strategic investment from Ricoh as Weaviate continue expanding the AI-native platform for developers and enterprises alike. Latest AI & tech insights. Explore its recent Weaviate content: Read. * | Building AI should start with building - not billing: Weaviate has always been free to self-host. Now, Weaviate is extending that philosophy across the entire Weaviate Cloud platform with a free tier for the managed database, Query Agent, and Engram. Read the blog * | Ricoh invests in Weaviate: A new strategic investment from Ricoh through the RICOH Innovation Fund marks another exciting milestone as Weaviate continue building the AI-native database powering the next generation of AI applications. Read the announcement * | Memory changes everything for AI agents: Learn how Engram gives agents persistent memory and context, enabling them to remember past interactions, learn over time, and power more capable production-ready applications. Read the announcement * | Import & Vectorize Data at Scale: Moving from a prototype to millions of objects introduces new challenges. Learn production best practices for high-volume imports, server-side batching, retries, multimodal ingestion, and building resilient pipelines with Weaviate. Read the guide Watch. * | Knowledge Engineering with Dr. Bradley Allen - Weaviate Podcast #139: Dr. Allen explores five decades of AI history, from expert systems and knowledge graphs to today's LLMs, explaining why knowledge engineering remains essential for building trustworthy enterprise AI. Watch the full podcast * | Can Weaviate Win Europe's Biggest Hackathon?: Follow the Weaviate team as Weaviate take on one of Europe's largest AI hackathons, building an agentic content platform in just 26 hours while exploring the future of Europe's startup and builder ecosystem. Watch the journey * | Building a Production-Ready Legal RAG App... in One Prompt: See how Weaviate built a production-ready legal assistant using Weaviate Query Agent, multi-vector retrieval, and coding agents - transforming weeks of development into a single guided prompt. Watch the tutorial Product highlights. * πŸ†“ Weaviate Cloud Free Tiers: Weaviate Cloud now includes free tiers (with no credit card required) for both Database and Engram. Launch a fully managed cluster, experiment with Query Agent and persistent memory, and upgrade whenever you're ready. Start free * | Introducing the Weaviate Playground: Explore a growing collection of interactive demos built with Weaviate - from Engram and Query Agent to multimodal search and AI-powered recommendations. It's the easiest way to experience what's possible before building your own applications. Explore the Playground * | Engram General Availability: Its managed memory service for AI agents is now production-ready, enabling persistent context across sessions so agents can remember, learn, and improve over time. Try it now * | Weaviate Core v1.36.18: The latest patch release delivers stability improvements, bug fixes, RBAC enhancements, and infrastructure updates to keep production deployments running smoothly. Read the release notes. Ready to start building? Jump right in and spin up your free cluster with Weaviate Cloud. Or check its GitHub and star Weaviate while you're there. Hungry for more? Have a question or want to connect? Join its Weaviate Forum to engage in community conversations. See you in two weeks,

Action Intelligence
Jun 16th, 2026
Ricoh invests in AI database developer Weaviate through Innovation Fund.

Ricoh invests in AI database developer Weaviate through Innovation Fund. On June 16, Ricoh Company, Ltd. announced that it has invested in Weaviate, a Netherlands-based developer of an AI-native vector database for unstructured data, through the Ricoh Innovation Fund. Ricoh said the investment was made on March 13 and is intended to support development of new solutions that combine Ricoh's data-capture technologies with Weaviate's context-aware To see the rest of this post, please view membership options or log in to your account below.

Recently Posted Jobs

Sign up to get curated job recommendations

Weaviate is Hiring for 2 Jobs on Simplify!

Find jobs on Simplify and start your career today

Don't see your dream role? Check out thousands of other roles on Simplify. Browse all jobs β†’