
Work Here?
Capy builds an AI-native software engineering platform that runs autonomous AI agents in secure, dedicated virtual environments. Its core product, Scrapybara, offers instant, secure virtual desktops via an API, enabling AI agents to perform long-horizon tasks on large codebases without needing external browsers or ad-hoc APIs. The system executes on its own virtual machine asynchronously, reducing deployment time from hours to seconds and providing a scalable, secure environment for agents to browse, process, and modify code and tasks. Capy is differentiating itself by delivering a purpose-built infrastructure for AI agents to operate inside a controlled computing environment, rather than relying on general APIs or unstable browser contexts. The company’s goal is to provide infrastructure-as-a-service that enables companies to deploy autonomous AI agents to handle complex operational tasks across software development, data processing, and workflow automation.
Industries
Data & Analytics
Enterprise Software
Cybersecurity
AI & Machine Learning
Company Size
1-10
Company Stage
Seed
Total Funding
$130K
Headquarters
San Francisco, California
Founded
2024
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$130k
Below
Industry Average
Funded Over
1 Rounds
Industry standards
Company Equity
Health Insurance
Dental Insurance
Vision Insurance
Unlimited Paid Time Off
Capy v2: a cloud coding agent with 32 vCPU sandboxes and sub-second VM boots, bring your own Codex or Grok subscription. Capy launched v2, a cloud coding agent that spawns VMs with up to 32 vCPU and 128GB RAM at sub-second boot, runs agents in parallel, and lets you bring your own Codex or Grok subscription. It also claims to beat Claude Code, Codex, Devin and Cursor on DeepSWE while being 50% cheaper, with no score published for itself or for any of the four. The infrastructure claims are specific and testable, the cost claim follows from the billing model rather than from efficiency, and the benchmark claim currently has nothing attached to it. Links & resources. | Resource | Link | | The announcement | @justinsunyt, August 12 | | The benchmark | DeepSWE, via datacurve-ai/pier | | Comparison point | Grok 4.6's DeepSWE row | Capy launched v2 on August 12, a cloud coding agent that spawns VMs with up to 32 vCPU and 128GB RAM at sub-second boot, lets you bring your own Codex or Grok subscription, and claims to "beat Claude Code, Codex, Devin and Cursor on DeepSWE while being 50% cheaper." The infrastructure claim is the specific one and the benchmark claim is the loud one. They need different treatment. The benchmark claim has no published run. Four named competitors, one benchmark, one percentage implied and none given. The post carries no score for Capy v2, no scores for the four it says it beat, no harness description, no date, no link to a leaderboard entry. DeepSWE is a real benchmark and it is not obscure. xAI put it in Grok 4.6's launch table this week at v1.1, where GPT-5.6 Sol scores 73%, Fable 5 scores 70% and Grok 4.6 scores 65.9%. Those are model scores rather than harness scores, so they are not the right comparison, but they establish that the benchmark has published numbers other people can check against. Without a number, "we beat Claude Code, Codex, Devin and Cursor" is a claim about four products by name with nothing attached. That is the kind of statement that gets a correction rather than a citation, and it costs nothing to fix: one table. The infrastructure claim is the interesting one. Strip the benchmark line and what is left is a real product with real specifics: * up to 32 vCPU and 128GB RAM per agent, which is a workstation per task rather than a container * sub-one-second VM boots, which is the number that decides whether spawning a fresh environment per attempt is practical * parallel agents, so you can fan out across approaches * bring your own Codex or Grok subscription, so the model cost sits on a plan you already pay for That last one is the commercially clever part and it is what makes "50% cheaper" plausible without touching a benchmark. If you are already paying for a Codex subscription, a harness that spends your existing quota rather than reselling you tokens is cheaper by construction. The saving comes from the billing model, not from the agent being more efficient. Sub-second boots at that memory size is the claim I would test first, because it is falsifiable in an afternoon and it is what the rest depends on. An agent that can throw away a broken environment and start clean in under a second can afford strategies that a thirty-second boot forbids. The comment-farming line. Worth naming because it explains the engagement. 198,000 views and 1,654 likes on an account with 8,016 followers is not organic reach, it is a credit giveaway. That does not make the product worse and it does make the numbers under the post meaningless as a signal. What to do with this. If you run coding agents in the cloud and boot latency or memory ceilings are your constraint, this is worth a trial, and the BYO-subscription model means the trial is cheap. Ignore the DeepSWE line until there is a table. Four competitors named without a score is not a result, and the fastest way for Capy to convert this post into credibility is to publish the run: harness version, model, sample, date, and the four baselines measured the same way. And test the boot time yourself. It is the number the whole architecture rests on and the only one in the post you can verify without their cooperation.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
Cybersecurity
AI & Machine Learning
Company Size
1-10
Company Stage
Seed
Total Funding
$130K
Headquarters
San Francisco, California
Founded
2024
Find jobs on Simplify and start your career today