Pika

Pika

AI-powered video creation and editing tools

Growth Engineer

Full-TimeUpdated on 9/29/2026
No salary listed
Mid
Remote in USA
Remote

About the job

Requirements
  • At least 4 years of software engineering experience, ideally on product or growth teams.
  • A proven record of building and optimizing user-facing features, preferably in consumer or software-as-a-service products.
  • Experience with experimentation frameworks, A/B testing, and data analysis.
  • Strong problem-solving and analytical skills, with the ability to translate insights into scalable solutions.
  • Proficiency with modern web or backend technologies such as Python, Go, Node.js, or TypeScript.
  • Familiarity with growth tools and analytics platforms such as Amplitude, Segment, or Google Analytics.
  • Ability to collaborate with cross-functional teams.
Responsibilities
  • Design, develop, and launch growth features and experiments across the Pika platform.
  • Implement scalable A/B tests and personalization strategies to optimize user acquisition and retention.
  • Analyze data to discover new growth levers and track the effectiveness of growth strategies.
  • Work closely with product, marketing, and data teams to prioritize and execute high-impact growth initiatives.
  • Rapidly prototype and ship new product surfaces, onboarding flows, and engagement loops.
  • Optimize performance, reliability, and user experience for features targeting key growth metrics.
  • Contribute to product virality, referral, and sharing capabilities to expand reach.
  • Advocate for a data-driven, experimentation-focused engineering culture.
Desired Qualifications
  • Experience in generative artificial intelligence, creative tools, or social platforms.
  • Background in high-growth startups or rapid experimentation roles.
  • Open-source, hackathon, or growth-hacking experience.

About the company

Pika.art builds an online platform that turns ideas into videos using AI. It offers text-to-video, image-to-video, and video-to-video transformations so users can create, edit, and extend videos with simple text commands. The product works through AI-powered video generation and editing tools accessible in a browser-like interface, enabling users to modify scenes, extend runtimes, and resize canvases without deep technical skills. Pika differentiates itself by targeting a broad range of creators—from casual meme makers to professional filmmakers—with a user-friendly, text-driven workflow and a subscription-based model that likely includes premium features for advanced tools. The company’s goal is to democratize video creation by making powerful AI-assisted video production accessible and customizable to everyday users and professionals alike.

Company Size

51-200

Company Stage

Series B

Total Funding

$135M

Headquarters

Palo Alto, California

Founded

2023

Get referred to Pika

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • August 2026 API Club cut prices up to 88% below aggregators, attracting developers.
  • Pika SFX and Pika Audio create a fuller workflow, reducing tool-switching friction.
  • Fortune cited 16.4 million users in 2025, proving broad consumer demand and retention.

What critics are saying

  • Runway, Higgsfield, and OpenAI Sora outspend Pika; feature parity compresses margins by 2027.
  • Pika's $470 million valuation is unchanged since June 2024; funding drought constrains hiring.
  • If Adobe or Apple ships native generative video, Pika loses distribution and relevance.

What makes Pika unique

  • September 2025 Adobe Firefly Boards integration put Pika inside Adobe's creator workflow.
  • Pika API Club bundles 100+ models through one API, launched August 2026.
  • Pika SFX delivers sub-second audio generation, adding native sound to video creation.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

↑ 2%

1 year growth

↑ 2%

2 year growth

↑ 5%
Semafor
Sep 17th, 2026
Pika launches faster AI video tool to compete with $5.4B Higgsfield at half the price

Video AI company Pika has launched a programme to accelerate AI video creation by reducing the need for detailed prompts. The new product automatically selects appropriate AI models for different video elements and allows users to input basic details to generate an editable, shot-by-shot list. Pika is also releasing a suite of apps, including tools for relighting and swapping characters or outfits. The company, valued at $470 million and backed by Spark Capital and Lightspeed Venture Partners, faces competition from rivals like Higgsfield AI, valued at $5.4 billion, and Runway, which already offer integrated creator platforms. Pika says its product will be available at half the price of competitors, including Seedance.

Aagey Se Right
Aug 21st, 2026
AI is closing the gap between idea and execution!

AI is closing the gap between idea and execution! This week at a glance. * Developer AI: Google launches Gemini 3.7 Flash with stronger coding and automation performance at half the introductory price of its predecessor. * AI Video: Higgsfield releases Cinema Studio 4 with eight new production controls for camera movement, lighting, color grading, and character emotion. * AI Music: Pika Labs launches four AI audio models for speech, music, sound effects, and synchronized soundtracks at up to 20x lower prices than rivals. * Creative AI: FLORA launches Fashion Studio, taking apparel brands from hand- drawn sketches to Shopify-ready campaign assets in one workspace. * AI Video: Topaz Labs releases Hyperion 2.5, bringing AI-generated video into professional HDR workflows for broadcast and streaming. Google launches Gemini 3.7 Flash with stronger coding performance and half the previous price. Higgsfield launches Cinema Studio 4 with eight new production controls for solo creators. Topaz Labs releases Hyperion 2.5, letting AI generated video enter professional HDR workflows. Pika Labs launches four AI audio models at prices up to 20 times lower than rivals. FLORA launches Fashion Studio, taking apparel brands from hand drawn sketch to Shopify campaign in one workspace.

Communeify
Aug 19th, 2026
AI daily|openai pauses RL training for safety; GLM-5.3 released; Claude adds Gmail & Drive connectors.

AI daily|openai pauses RL training for safety; GLM-5.3 released; Claude adds Gmail & Drive connectors. August 19, 2026 Updated Aug 19 Model releases & updates. GLM-5.3 - Zhipu AI (Z.ai). * TL;DR: Zhipu AI officially released GLM-5.3, securing a top score on the Artificial Analysis index through advanced post-training improvements. * Key Highlights: * Scores 60 on the Artificial Analysis Intelligence Index, matching Kimi K3 as one of the highest-performing open base models available. * Significant performance jump driven purely by post-training: Terminal-Bench 3.0 surged from 4.6 to 28.3, and DeepSWE v1.1 rose from 46.2 to 66.9 compared to GLM-5.2. * Retains identical pricing as GLM-5.2 with lower output token generation per task, accessible immediately via Z.ai API, OpenRouter, and Vercel AI Gateway. * Specs: Mixture-of-Experts (753B Total / 40B Active) / Open Weights (releasing within a week) / 1M Context Window * Links: | Z.ai Documentation Pika Soundtrack - Pika labs. * TL;DR: Pika introduced Pika Soundtrack, a latent diffusion transformer model engineered for precise video-to-audio synchronization. * Key Highlights: * Cross-attends latent video tokens with semantic audio prompt tokens to generate motion-aware sound effects, ambient audio, and score tracks. * Ranked #1 for semantic alignment (0.2457 ImageBind) and audiovisual synchronization (0.5537 DeSync) across a 67-chunk benchmark test. * Released via the Pika API Club at up to 50% lower cost compared to competing generative audio models. * Specs: Proprietary Video-to-Audio Model / Latent Diffusion Transformer / Multi-Modal Alignment * Links: Product releases & updates. Direct Gmail & Google Drive connectors - Anthropic (Claude). * What's New: Anthropic integrated direct Google Workspace connectors into Claude. Users can now prompt Claude to search Google Drive files or draft and send emails directly inside Gmail. Built-in permission gates allow users to specify when human confirmation is required before any email is dispatched. * Who It's For: Knowledge workers, managers, and enterprise users automating daily email and document workflows. * Try It: Origin native git hosting - Cursor. * What's New: Cursor launched Origin, a native code-hosting platform built on database-grade storage infrastructure. Designed for high uptime and performance, Origin features bi-directional GitHub synchronization, PR management, and inline repository editing directly within the Cursor ecosystem. * Who It's For: Software engineers and enterprise teams seeking high-reliability git infrastructure for AI-driven development. * Try It: TensorRT Model Connect - NVIDIA. * What's New: NVIDIA released the public preview of TensorRT Model Connect (TRTMC), an open-source tool that converts Hugging Face or local model checkpoints directly into native C++ TensorRT inference engines in two commands, completely eliminating intermediate ONNX conversions. * Who It's For: AI infrastructure engineers, C++ developers, and high-performance inference builders. * Try It: | GitHub Repository ChatGPT for Teens - OpenAI. * What's New: OpenAI introduced ChatGPT for Teens for users aged 13-17. The feature automatically turns on enhanced content safety guardrails, Study Hours scheduling, parental controls, and a step-by-step "Study Mode" designed to guide students through homework problems rather than providing direct answers. * Who It's For: Teenagers, parents, and educators looking for safe, educational AI assistance. * Try It: OpenAI Announcement AI SDK "code Mode" Sandbox execution - Vercel. * What's New: Vercel added "Code Mode" to the AI SDK. Instead of relying solely on multi-turn JSON tool calls, models generate JavaScript/TypeScript code that executes within an isolated QuickJS sandbox to programmatically orchestrate and run tool sequences in a single turn. * Who It's For: Full-stack developers building complex, multi-tool AI agents and autonomous web applications. * Try It: fx lightweight agent harness - Vercel Labs. * What's New: Vercel Labs open-sourced fx, a minimalist coding agent harness written natively in Zig. Featuring a 6.3MB binary, single-digit megabyte memory usage, and 10-microsecond cold-start execution, fx is optimized for high-speed terminal development, WebAssembly embedding, and lightweight agent evaluation. * Who It's For: Developers seeking fast, open-source terminal coding agents and embeddable harnesses. * Try It: Industry news. OpenAI pauses frontier RL training to strengthen AI safety protocols. * What Happened: OpenAI announced a two-week pause on reinforcement learning (RL) training for its next-generation deployment models - including the upcoming Astra model - to harden internal research environments and expand red-teaming. Additionally, OpenAI is dedicating up to 20% of its research inference compute toward multi-stage chain-of-thought monitoring to detect concerning model behaviors prior to public release. * Why It Matters: This marks a major voluntary pause by a leading frontier lab triggered by internal capability thresholds. As AI capabilities in cybersecurity advance, safety verification and alignment monitoring are increasingly determining the pace of frontier model deployments. * Source: OpenAI Announcement | OpenAI, NVIDIA, and SB Energy Partner on 8GW Data Center Project. * What Happened: OpenAI, NVIDIA, and SoftBank's SB Energy finalized an agreement to construct the PORTS-Pike data center campus in Ohio. The project targets 10GW of clean energy generation and up to 8GW of dedicated AI factory compute capacity, with NVIDIA providing up to $105 billion in financial backstop guarantees for the long-term lease. * Why It Matters: The deal illustrates the huge capital requirements and novel financial structures needed for next-generation compute scaling, with hardware providers directly underwriting data center infrastructure to secure capacity for frontier AI labs. * Source: Pew survey: over 50% of young adults in the US concerned about AI job impact. * What Happened: A new Pew Research survey revealed that 55% of US adults under 30 express more concern than excitement regarding artificial intelligence. Furthermore, 73% of young respondents expect AI to reduce overall jobs over the next two decades, up from 61% in 2024. * Why It Matters: Public sentiment among younger demographics is increasingly shifting toward economic displacement concerns, which could impact future US regulatory policy, labor legislation, and consumer adoption rates. * Source: Vercel launches $1 million hacker challenge for sandbox security. * What Happened: Vercel announced a 1,000,000publicbugbountyprogramforVercelSandbox.SecurityresearchersandAIdevelopersareinvitedtousefrontierAImodelstoattemptescapingtheFirecrackermicroVMorbypassinghost−sidenetworkisolation,withbountiesreachingupto1,000,000publicbugbountyprogramforVercelSandbox.SecurityresearchersandAIdevelopersareinvitedtousefrontierAImodelstoattemptescapingtheFirecrackermicroVMorbypassinghost−sidenetworkisolation,withbountiesreachingupto50,000 per verified report. * Why It Matters: As autonomous coding agents gain permissions to execute untrusted code in production, verifying microVM boundaries against AI-driven exploits is becoming critical enterprise infrastructure security work. * Source: Vercel Blog | Research papers. Autonomous protein design via frontier LLMs - Anthropic Research. * Motivation: Computational protein design traditionally relies on human domain experts to manually configure complex multi-step pipelines, creating a major bottleneck in drug discovery and biological research. * Key Innovation: Anthropic evaluated Claude's ability to autonomously run computational protein design workflows without human intervention. Using expert-written protocols, Claude designed novel protein binders for 15 biological targets, which were subsequently synthesized and laboratory-tested by Adaptyv Bio and Twist Bioscience. * Results: Claude achieved a 22.6%-35.1% binding success rate (354 out of 1,320 designs bound successfully across 14 targets) - roughly doubling the traditional human expert benchmark of 10%-15%. * Paper: Anthropic Research | Other Highlights. Mojo programming language officially open sourced - Modular. * Overview: Modular officially open-sourced the Mojo compiler and toolchain under the Apache 2.0 license (with LLVM exceptions) following its 1.0 release. Designed to simplify high-performance GPU programming with Python-inspired syntax, Mojo's core codebase is now publicly available on GitHub. * Link: Modular Blog Miles v0.1 open-source RL framework - RadixArk. * Overview: RadixArk released Miles v0.1, an open-source reinforcement learning framework for LLMs and vision-language models. Developed by 72 contributors, Miles focuses on RL debugging and hardware utilization, passing 85 end-to-end GPU CI tests across models like DeepSeek V4, Kimi K3, and Qwen 3.8. * Link: Featured Partners

AIKONS
Aug 18th, 2026
Pika Audio: full-stack video + Sound for AI spots.

Pika Audio: full-stack video + Sound for AI spots. AIKONS · August 18, 2026 Property of AIKONS The silent film era of AI video is over. For too long, the most compelling AI-generated visuals have emerged into the world without a voice, a score, or a soundscape. Adding audio has been a fractured, multi-step process involving separate tools, disjointed workflows, and creative compromises. This has been the primary obstacle between a promising clip and a finished AI spot. Pika's recent launch of four frontier audio models changes the landscape. By integrating sound generation directly into its video platform, Pika has created the first true full-stack solution for AI video creation. This isn't just an update; it's a fundamental shift in the AI filmmaking workflow. Pika's new audio models: why this matters for AI spots. Sound is not an accessory in advertising. It drives emotion, directs attention, and defines brand identity. The inability of AI video models to handle audio natively has been a critical limitation for creators and advertisers. Pika Audio directly addresses this gap. By turning AI video generation into a comprehensive workflow - from silent prompt to fully realized audiovisual asset - it collapses the production timeline. This allows for a level of creative velocity and cohesion previously impossible. Now, a concept for an AI spot can be developed, iterated, and finished with a synchronized soundtrack, voiceover, and sound effects, all within a single ecosystem. What Pika Audio adds to the AI video workflow. The traditional AI video post-production process is a patchwork. A creator might generate video in one platform, find stock music on another, use a separate text-to-speech tool for narration, and then manually assemble everything in a video editor. Each step introduces friction and potential for creative disconnect. Pika's integrated suite streamlines this process into a single, fluid motion. The workflow compresses from days or hours to minutes. * Before: Generate video clips. Export. Import into an editor. Search for music. Generate voiceover. Find SFX. Manually sync every element. * After: Generate a video clip. Use Pika Soundtrack to create a synchronized audio layer. Add custom voiceover with Pika Speech. Refine with Pika SFX. Export a finished asset. This efficiency is more than a convenience. It enables rapid A/B testing of different audio treatments, faster turnaround on creative, and a more intuitive, holistic approach to storytelling. The four models: soundtrack, music, SFX, and speech. The power of Pika Audio lies in its four specialized foundation models, each targeting a critical component of sound design. These tools are available exclusively through the Pika API Club, reinforcing the focus on an integrated developer and creator ecosystem. Pika Soundtrack: the end of silent AI video. Pika Soundtrack is the cornerstone of the announcement. This video-to-audio AI analyzes a video clip and generates a complete, synchronized soundscape. It produces music, ambient noise, and motion-aware sound effects that align with the on-screen action. This model alone solves the "silent video" problem that has plagued the space, turning raw visual output into a textured, immersive experience. Pika Music: composing on command. Pika Music functions as an in-house composer. It can generate complete songs from simple text prompts, supplied lyrics, or even reference audio tracks. For advertisers and filmmakers, this means the ability to create bespoke, royalty-free scores that perfectly match the mood and pacing of their visuals, bypassing the limitations and costs of stock music libraries. Pika SFX: precision sound design. Every impactful video relies on subtle audio cues. Pika SFX is a generative audio model designed to create clean, high-fidelity sound effects from text prompts. Need the sound of a futuristic car door closing or a specific brand of soda being opened? This tool provides the precision required for professional ad production and post-production workflows. Pika Speech: giving AI a voice. Narration and dialogue bring a story to life. Pika Speech is an advanced text-to-speech AI that creates expressive, realistic human speech. It supports a range of preset voices and, crucially, offers voice cloning AI capabilities. This allows brands to create a consistent, ownable voice for their AI spots or clone an actor's voice for localized campaigns without requiring new recording sessions. Why full-stack video + Sound is a competitive advantage. In a rapidly crowding field of AI tools, building a moat is essential. Pika's strategy is clear: the advantage isn't just in the quality of the video or audio, but in the seamless integration of the two. By offering a full-stack platform, Pika controls the entire creative journey. This vertical integration creates a powerful feedback loop where improvements in video generation can inform audio generation, and vice versa. For creators, it eliminates the "context switching" that fragments creative energy. For Pika, it builds deep user loyalty by making the platform the most efficient place to get from idea to finished product. What advertisers can do with Pika Audio today. The practical applications for advertisers and marketing teams are immediate and significant. The integration of high-quality generative audio models unlocks new possibilities for creating dynamic, cost-effective AI spots. * Rapid Prototyping: Develop complete audiovisual concepts for social ads in minutes, not days. * Performance Creative at Scale: Generate dozens of variations of a single ad, each with a unique soundtrack or voiceover, to A/B test what resonates with audiences. * High-Fidelity Product Demos: Enhance product visuals with custom SFX and clear, professional narration. * Consistent Brand Storytelling: Use voice cloning AI to establish a single, recognizable brand voice across all video content. * Concept Visualization: Show stakeholders a nearly-finished spot, complete with sound, to get buy-in before committing significant production resources. Pricing, access, and API availability. Pika has positioned its audio models not only as a creative advantage but also as a significant cost advantage. The models are available through the Pika API Club, signaling a focus on developers and professional creators building custom workflows. The company has made direct comparisons on pricing. Public statements have claimed that its audio generation can be significantly cheaper than competing, specialized audio models - in one case, up to 20 times less expensive for comparable outputs. This pricing strategy is designed to make Pika the default choice for studios and agencies looking to produce AI-driven creative at scale. Best use cases for AI filmmakers and brands. The implications extend beyond short-form advertising. The new toolset empowers both independent creators and established brands. For AI filmmakers, Pika Audio provides the tools to build richer narratives. An entire short film can be scored, voiced, and sound-designed within the same environment where the visuals were born. This allows for a director-level of control over the complete sensory experience. For brands, it opens the door to creating a consistent sonic identity across global markets. A campaign's narration can be cloned and adapted to different languages while retaining the original tone and cadence. Product demos can be accompanied by precise, satisfying sound effects that enhance the perceived quality of the item. What this means for the AI creative stack. Pika's move signals a maturation of the AI creative market. The era of fragmented, single-purpose tools is giving way to powerful, integrated platforms. Value is no longer just in generating an image, a video clip, or an audio file. The true value lies in orchestrating the entire creative workflow. This consolidation is a natural and necessary step. It reduces friction for creators and enables a more ambitious and holistic vision for what AI-native storytelling can be. The focus is shifting from the novelty of generation to the craft of production. The tools are becoming more powerful. The workflows are collapsing. The question is what stories you will tell with them. If you're ready to explore what's possible at the intersection of narrative and AI, let's build the future of storytelling together. Start a project with AIKONS. Pika audio models Pika Soundtrack Pika Music Pika SFX Pika Speech AI video sound design generative audio models AI spots video-to-audio AI voice cloning AI text to speech AI AI filmmaking workflow

Pika
Aug 18th, 2026
Now hear this: Pika SFX.

Now hear this: Pika SFX. Production-ready sound from language, generated faster than real-time August 18, 2026 - By Pika Team Today, Pika Labs Inc. is introducing Pika SFX, a new model that turns a written description into a focused, ready-to-use sound effect in real-time. Describe a single event - glass shattering or a cartoon whoosh - or direct a complete sound sequence with timing, environment, texture, and mood. Pika SFX generates an audio clip that follows the request while avoiding unrelated speech, music, and background noise unless they are explicitly part of the prompt. Listen to the model. A single event. glass shattering Material, space, and perspective. A heavy metal door slamming shut in a large empty warehouse, followed by a long metallic echo. A sequence of events. A cork popping out of a champagne bottle, followed by fizzy liquid being poured into a glass, the bubbles softly crackling as they settle. Layered sound design. An old malfunctioning machine: gears grinding irregularly, steam hissing from a valve, electrical sparks crackling, a loose bolt rattling, ending with a heavy clunk and a power-down whir. Designed fantasy sound. A magical spell being cast: shimmering sparkles building up, an ethereal choir-like swell, then a powerful whoosh release with glittering chimes trailing off. The technical challenge. A useful sound prompt can be as short as "a metal door slamming in a warehouse." Or it can describe a sequence of events: "A cork pops out of a champagne bottle, followed by fizzy liquid pouring into a glass, with bubbles softly crackling as they settle." Prompt fidelity, however, is only the first of three constraints. The second is audio quality: sounds should have crisp, accurate transients, plausible acoustics and reverberation, and no unwanted content beyond what the prompt specifies. The third is efficiency: how can Pika Labs Inc. generate high-quality audio quickly and cost-effectively, without sacrificing fidelity or controllability? Pika SFX is designed to tackle all three challenges at once, combining prompt fidelity, audio quality, and efficient generation in a single system. Methods. Architecture overview. Pika SFX translates a natural-language description into a complete 44.1 kHz stereo audio waveform. A transformer-based text encoder first maps the prompt into conditioning embeddings that capture the requested sound sources, physical actions, acoustic environment, timing, texture, and creative direction. The generative core is a text-conditioned diffusion transformer (DiT) operating in a compressed acoustic latent space. R ather than generating waveform samples directly, the transformer progressively constructs a compact representation of the sound. This reduces the sequence length and computational cost while preserving both fine acoustic details - such as transients, texture, and reverberation - and broader temporal structure. The model can therefore represent anything from a single impact to evolving ambience and layered multi-event sequences. A semantic-acoustic autoencoder then decodes the generated latents into the final stereo waveform. Pika Labs Inc. designed the system around three goals: * Prompt fidelity: preserve the objects, actions, environment, and sequence described by the user. * Production quality: generate clean, balanced audio with minimal artifacts or unrelated content. * Efficient generation: preserve useful sound quality within a distilled, few-step inference path suitable for interactive applications. This architecture provides the foundation; the inference stack determines how quickly it can be served. Training for efficient generation. Pika Labs Inc. train Pika SFX in three stages. First, the model learns to transform random noise into structured audio using a flow-matching objective. During training, it observes intermediate points between noise and real audio and learns the direction - or velocity - that moves each point toward a clean audio representation. This teaches the diffusion transformer to model diverse sounds, temporal structure, and text-audio relationships. Next, Pika Labs Inc. apply teacher-student distillation. A high-quality teacher follows a longer, multi-step generation path, while the student learns to predict the teacher's final result from an intermediate noisy state. This effectively shortens and straightens the generation trajectory, allowing the model to produce useful audio with far fewer inference steps. Finally, Pika Labs Inc. refine the distilled student model through post-training on high-quality curated data with human feedback. Together, these stages allow Pika Labs Inc. to generate high-fidelity content efficiently. Real-time interactive inference. Pika Labs Inc. treat inference as an end-to-end systems problem, optimizing the model, hardware, and serving stack for continuous generation. At a high level, its approach combines two ideas: * Few-step generation. Pika Labs Inc. distill the model to generate high-quality audio in fewer inference steps, substantially reducing sequential computation. * Deployment-aware optimization. Pika Labs Inc. optimize for the specific combination of model architecture, hardware, output duration, and serving configuration instead of treating each component in isolation. Together, these choices turn inference from a cold batch job into a warm, continuous generation loop. Evaluation. Evaluation setup. The benchmark uses 50 prompts and 454 seconds of requested audio per model, with requested durations from 2 to 20 seconds. It covers short and detailed prompts across impacts, Foley, animals, nature, vehicles, human sounds, ambience, spatial motion, and layered multi-event sequences. Generation performance. Across the same 50 prompts of requested audio per model, Pika SFX completed its locally hosted end-to-end path in 0.847 seconds on average. For context, the fastest hosted APIs in this benchmark averaged 2.47 seconds for ElevenLabs v2 and 2.64 seconds for Stable Audio 3 SFX. Figure: Generation latency versus requested audio duration across text-to-SFX models. Pika SFX achieves sub-second end-to-end generation, including CPU processing and file saving. Lines show means; shaded bands show +/-1 standard deviation. Lower is faster. Audio quality. Evaluating generated sound is difficult. A clip can match a prompt in embedding space and still be noisy, poorly balanced, or awkward to place in an edit. Pika Labs Inc. therefore evaluated both prompt-related metrics and the practical qualities of the resulting audio. What Pika Labs Inc. measure. The evaluation separates prompt alignment from practical audio quality: * CLAPScore: similarity between the generated audio and prompt in a shared embedding space. * AQAScore: whether the prompted events are present, as assessed by a multimodal model. * Content usefulness: whether the generated content is practically usable, using the corresponding Meta Audiobox Aesthetics axis. * Production quality: whether the audio is clean and ready for creative work, using the corresponding Meta Audiobox Aesthetics axis. Results. Across this benchmark, Pika SFX achieved the highest content usefulness and production quality scores among the systems evaluated. Pika SFX leads the measurements closest to using the result in a creative workflow - whether the content is useful and whether the audio sounds production-ready. Applications and future work. Pika SFX is designed as a building block. It can generate an isolated effect for an edit, create a sequence from a detailed direction, supply ambience for a scene, or bring sound generation directly into an application through the API. This release is one part of its broader work on generative audio. Pika SFX begins with language. Pika Video-to-SFX begins with the picture itself, generating effects and ambience aligned with visible events and motion across a video. Availability. Pika SFX is available now through the Pika API Club. Developers can generate downloadable MP3 sound effects from text prompts, with configurable durations from 1 to 20 seconds and optional controls for negative prompts, seed, inference steps, and guidance scale. Pika Labs Inc. is excited to hear what you create.