
Work Here?
VAST Data provides a data platform that stores and manages very large data sets for AI and big-data workloads, unifying structured and unstructured data in a global namespace and scaling to petabytes. It uses a scale-out, multi-tenant, zero-trust architecture that combines storage and data management in one stack for fast, secure access and processing. It differentiates itself with a new scale-out design, a single global namespace, strong security, and partnerships with companies like NVIDIA to reduce cost and complexity compared to traditional cloud storage. Its goal is to help customers store, access, and analyze massive data sets for AI and other compute-heavy tasks at a lower cost.
Industries
Data & Analytics
Enterprise Software
Cybersecurity
AI & Machine Learning
Company Size
1,001-5,000
Company Stage
Growth Equity (Non-Venture Capital)
Total Funding
$881M
Headquarters
New York City, New York
Founded
2016
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$881M
Above
Industry Average
Funded Over
7 Rounds
Flexible Work Hours
Building scalable, multi-tenant AI Factories and Gigafactories with OpenNebula and VAST Data. AI Factories require much more than GPU clusters. As organizations invest in AI infrastructure, they need a platform capable of orchestrating compute, networking, storage, virtualization, Kubernetes, and AI services while providing the performance, scalability, and multi-tenant isolation required by modern AI workloads. OpenNebula Systems and VAST Data address this challenge by combining OpenNebula's end-to-end AI infrastructure management platform with VAST's AI-native data platform. Together, they enable organizations to build and operate AI Factories, AI Gigafactories, and next-generation Neoclouds that support AI training, fine-tuning, inference, scientific computing, and other data-intensive workloads at scale. OpenNebula provides a unified control plane for managing heterogeneous AI infrastructure, while VAST delivers the scalable data foundation required by large-scale AI deployments. The integration extends multi-tenancy across the infrastructure and data layers, enabling tenant-aware provisioning and isolation of compute and storage resources within a consistent operational model. This complementary approach simplifies infrastructure operations while giving organizations the flexibility to design environments that meet their performance, scalability, and operational requirements. The joint solution is designed for national AI initiatives, sovereign AI Factories and Gigafactories, enterprise AI platforms, research infrastructures, and AI Neocloud providers. It supports deployments built on NVIDIA accelerated computing platforms across on-premises, hybrid, and sovereign cloud environments. By combining unified AI infrastructure orchestration with AI-native data services and integrated multi-tenancy across compute and storage, OpenNebula and VAST provide an integrated software foundation that helps organizations accelerate the deployment of production-ready AI infrastructure while reducing operational complexity. Discover how OpenNebula Systems and VAST Data can help you design and operate scalable AI Factories, AI Gigafactories, and AI Neoclouds. Contact its team to discuss your infrastructure requirements and deployment strategy. Alberto P. Martí. Director of Strategic Alliances at OpenNebula Systems Aug 10, 2026.
Alphabet's CapitalG, Nvidia in talks to fund Vast Data at up to $30 billion valuation, sources say. Reuters · 06 Aug 2026 Alphabet's NASDAQ:GOOG growth-stage venture arm CapitalG and Nvidia NASDAQ:NVDA are in talks to invest in artificial intelligence infrastructure provider Vast Data in a new funding round that could value the startup as high as $30 billion, two sources said. The startup is raising several billion dollars from tech giants, private equity and venture capital investors, which could make it one of the most valuable AI startups, the two sources with knowledge of the matter said, as companies building the backbone for the AI boom come into sharper focus. CapitalG and existing backer Nvidia are in discussions to participate in the round, which could close in the next few weeks, according to the sources, who requested anonymity to speak on private matters. New York-headquartered Vast Data develops storage technology specifically designed for large AI data centers, enabling efficient data movement between graphics processors (GPUs) made by the likes of Nvidia. Its clients include companies such as Elon Musk's xAI and CoreWeave NASDAQ:CRWV, and its value in the AI supply chain makes it an attractive acquisition target, bankers and analysts said. Nvidia declined to comment, while Vast Data and CapitalG did not respond to requests for comment. TechCrunch earlier reported Vast Data's fundraising efforts, but the valuation of up to $30 billion and the expected involvement of CapitalG and Nvidia have not been reported previously. Vast Data CEO Renen Halak has said the company is free cash flow positive. The company earned $200 million in annual recurring revenue (ARR) by January 2025, with a strong backlog of orders and projections to grow ARR to $600 million next year, according to a separate source familiar with its financials. The company has raised roughly $380 million to date, and its last funding round in 2023 valued it at $9.1 billion. Vast Data has said it would consider an initial public offering at the right time. While no listing is imminent, according to another source familiar with the matter, investors and bankers view the data infrastructure firm as a likely IPO candidate. Vast Data last year hired Amy Shapero, its first chief financial officer, who was previously in the same role at publicly listed e-commerce giant Shopify NASDAQ:SHOP, in a move that could signal preparations for an IPO. Mergers and acquisitions activity has also been heating up in the sector, and Nvidia has been acquiring companies that add complementary software and hardware products beyond its flagship GPUs. In 2020, it bought networking chip and cable maker Mellanox, which has helped Nvidia build integrated systems featuring its latest Blackwell chips. It has also acquired software companies such as Run:ai, which helps engineers optimize data center AI hardware. Vast Data's storage architecture is based on a system of flash storage devices and other off-the-shelf hardware, combined with its specialized software for data access and movement. The company says adopting its technology can reduce the cost of building and running large AI models. Several companies, such as Weka and DDN, are pursuing similar efforts, but analysts and industry executives say Vast Data's technology is more mature than that of its rivals. Shares of Nvidia slumped 2.3% and Alphabet stock sunk 1.4% during regular trading on Friday.
Optimizing AI infrastructure performance. The collaboration introduces several technical integrations to improve AI infrastructure efficiency. Notably, VAST selected 6th Gen AMD EPYC processors to power its 6th-generation CBox and 3rd-generation EBox platforms. These processors support PCIe Gen-6, which doubles the input-output bandwidth generationally to improve file and object storage performance. "AI is entering an operational phase where infrastructure efficiency matters as much as model performance." John Mao, Vice President of Global Technology Alliances at VAST Data Furthermore, early testing of the AMD Instinct MI355X GPU showed a 9X speedup in time-to-first-token. The system also achieved 9.7X more token throughput when using VAST for key-value cache offloading. These results are critical for delivering low-latency and energy-efficient deployments. Reference architectures and ecosystem growth. In addition, VAST, AMD, and DriveNets developed a reference architecture featuring AMD Helios rack-scale AI infrastructure. This architecture provides guidance for model training, inference, and reinforcement learning workloads. Meanwhile, software innovators like TensorMesh and EmbeddedLLM are collaborating to accelerate production-ready deployments. "Our expanded collaboration with VAST combines AMD EPYC CPUs and Instinct GPUs with the software foundation customers need to accelerate inference." Derek Dicker, Corporate Vice President at AMD The integration also includes automated key-value cache lifecycle management. This feature uses native data lifecycle policies to automatically delete cached data containing sensitive information. As a result, enterprises can maintain regulatory compliance without manual operational overhead. The system manages these tasks through its core database and event streaming capabilities. Industry adoption and future outlook. Several AI cloud providers, including Core42, Crusoe, and Vultr, are deploying these joint technologies. These companies require architectures that combine accelerated computing with context management. Ultimately, the unified platform helps organizations transition from pilot projects to production environments.
VAST Data, AMD expand collaboration to accelerate enterprise AI infrastructure. VAST Data and AMD have expanded their collaboration to help enterprises build AI infrastructure for large-scale inference and agentic AI. VAST Data and AMD have expanded their collaboration to help AI cloud providers and enterprises build high-performance AI infrastructure designed to support training, inference and agentic AI workloads at scale. The collaboration combines the VAST AI Operating System with 6th Gen AMD EPYC processors and AMD Instinct GPUs, bringing together compute, data management and software to address the growing demands of production AI environments. The companies said the joint platform is intended to improve infrastructure efficiency as organisations deploy increasingly complex inference, retrieval-augmented generation (RAG) and agentic AI applications. The expanded collaboration includes new AI infrastructure reference architectures developed alongside DriveNets, support for AMD's latest EPYC processors in VAST's next-generation platforms, and software integrations with TensorMesh and EmbeddedLLM. It also introduces KV cache optimisation capabilities, automated cache lifecycle management for enterprise compliance, and networking integration using the AMD Pensando Pollara 400 AI NIC. According to VAST Data, early testing using an AMD Instinct MI355X GPU demonstrated a 9X improvement in time-to-first-token and 9.7X higher token throughput through KV cache offloading for high-concurrency agentic AI workloads. "AI is entering an operational phase where infrastructure efficiency matters as much as model performance," said John Mao, Vice President, Global Technology Alliances at VAST Data. "The industry is discovering that inference is fundamentally a data problem. Success depends on how effectively organizations can bring data, compute, memory and intelligence together as a single system." Derek Dicker, Corporate Vice President, Enterprise Business Group at AMD, added: "Our expanded collaboration with VAST combines AMD EPYC CPUs and Instinct GPUs with the software foundation customers need to accelerate inference, improve infrastructure efficiency and deploy AI at scale." According to VAST Data, the collaboration reflects growing demand for AI infrastructure that supports large-scale inference and agentic AI alongside traditional model training.
Cloudera and VAST Data take aim at GPU starvation with joint AI factory stack. by Harold Fritts on July 15, 2026 Cloudera and VAST Data have entered into a partnership to build a unified AI factory architecture for enterprises running continuous AI training, inference, and analytics workloads. The joint offering combines Cloudera's containerized data services with the VAST AI Operating System, targeting a problem that has become increasingly common in enterprise AI deployments: expensive GPU clusters sitting idle while they wait on data. That idle-GPU problem, often called GPU starvation, tends to show up when organizations bolt AI workloads onto data architectures that were never built for the continuous, high-throughput demands of modern AI pipelines. Data preparation, training, inference, and analytics all compete for the same infrastructure, and when the data layer can't keep pace, the accelerators stall. Cloudera and VAST are positioning their combined stack as a fix for that bottleneck, aiming to keep GPUs fed with low-latency data throughout the full AI lifecycle. How the architecture is built. On the Cloudera side, the company brings its lakehouse architecture, which packages data engineering, streaming, analytics, machine learning, and AI services into portable containers that can run across hybrid and multi-cloud environments. That gives customers a consistent operational model regardless of where the workload actually executes. VAST contributes its Disaggregated Shared Everything (DASE) architecture as the underlying data infrastructure, scaling to exabyte levels while integrating vector database services with NVIDIA cuVS for GPU-accelerated vector indexing and search. The VAST AI OS is built on the NVIDIA AI Data Platform reference design and is designed to take latent enterprise data and turn it into an AI-ready state that can be consumed directly by training and inference pipelines. Cloudera then layers its data engineering, governance, and AI services on top of that foundation. The combined stack is meant to cover the full path from raw data ingestion through model deployment, with consistent operations across data centers, private cloud, and public cloud. Cloudera and VAST say the architecture eliminates GPU starvation through high-bandwidth, low-latency data pipelines, which should, in turn, improve sustained GPU utilization and compute efficiency. The platform is also built to handle structured, unstructured, and multimodal datasets at scale, with enterprise governance and compliance controls intended for private and sovereign AI deployments. NVIDIA integration and inference. The partnership leans heavily on NVIDIA's stack. Alongside the AI Data Platform reference design underpinning VAST's AI OS, the companies are integrating NVIDIA AI Enterprise software into the joint architecture. Cloudera's AI Inference Service uses NVIDIA NIM microservices to enable organizations to deploy and scale models directly on their own data, including NVIDIA's newer Nemotron open models. On the data engineering side, customers can accelerate Apache Spark workloads using NVIDIA cuDF, which integrates transparently with Cloudera Data Engineering. That allows Spark jobs to tap into VAST's high-throughput data services with GPU-accelerated processing, which should help on data-heavy stages of AI pipelines rather than just training and inference. For regulated industries in particular, Cloudera and VAST are framing this as a "silicon-to-application" solution, covering everything from the underlying NVIDIA hardware and software up through production AI applications, deployable on-premises or in the cloud depending on data residency and compliance requirements. Availability. The companies say the partnership combines 60 exabytes of customer-managed data across their installed bases, giving both vendors a large pool of existing enterprise customers to target as demand for private AI infrastructure grows. The joint Cloudera-VAST AI factory solution is available now through both companies' enterprise sales teams and partner networks. Cloudera and VAST plan to expand the portfolio through 2026 with additional reference architectures, validated deployment patterns, and industry-specific solutions. Engage with StorageReview Harold Fritts. I have been in the tech industry since IBM created Selectric. My background, though, is writing. So I decided to get out of the pre-sales biz and return to my roots, doing a bit of writing but still being involved in technology.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
Cybersecurity
AI & Machine Learning
Company Size
1,001-5,000
Company Stage
Growth Equity (Non-Venture Capital)
Total Funding
$881M
Headquarters
New York City, New York
Founded
2016
Find jobs on Simplify and start your career today