Full-Time
Develops AI image-generation models and licenses
€130k - €340k/yr
Freiburg im Breisgau, Germany
Hybrid
Hybrid arrangement: 2+ days on-site weekly in Freiburg or SF; remote with a monthly in-person week option.
See people who can refer or advise you
Black Forest Labs builds AI-powered image generation tools. Its flagship model, FLUX.1, delivers strong prompt adherence, high visual quality, and diverse outputs, and is offered through partnerships and licensing for commercial use, as well as customized enterprise solutions. They serve a broad audience from individual developers to large enterprises. For non-commercial use, models like 1 dev and 1 schnell are available on platforms such as HuggingFace and under Apache 2.0, providing options for local development and personal use. The company differentiates itself through a combination of high-quality image generation, flexible access (enterprise licensing, partnerships, and non-commercial options), and a clear focus on practical deployment. The goal is to advance AI-based image generation and make reliable, high-quality models accessible to a wide range of users and applications.
Company Size
51-200
Company Stage
Series B
Total Funding
$431M
Headquarters
Freiburg, Germany
Founded
2024
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Health Insurance
Dental Insurance
Vision Insurance
Paid Vacation
Paid Sick Leave
Paid Holidays
Flexible Work Hours
Hybrid Work Options
Remote Work Options
401(k) Retirement Plan
401(k) Company Match
Wellness Program
Gym Membership
Mental Health Support
Conference Attendance Budget
Professional Development Budget
Training Programs
Tuition Reimbursement
Stock Options
Company Equity
Phone/Internet Stipend
Home Office Stipend
Family Planning Benefits
Fertility Treatment Support
Adoption Assistance
Parental Leave
Relocation Assistance
Meal Benefits
Wellness Program
Employee Discounts
Company Social Events
PixPix integrates FLUX Upscale for native 4K AI Video enhancement. PixPix integrates Black Forest Labs' FLUX Upscale, enabling AI-generated and existing videos to be enhanced at resolutions up to native 4K. PixPix brings AI image generation, video generation, editing and enhancement tools together in one creative workspace. PixPix adds Black Forest Labs' FLUX Upscale, bringing native 4K AI video enhancement into its image and video creative workflow. FLUX Upscale gives creators a practical way to bring AI-generated video closer to production-ready 4K quality within the PixPix workflow." - Will Chen, Marketing Director at PixPix DOVER, DE, UNITED STATES, August 21, 2026 /EINPresswire.com/ - PixPix, an AI creative platform for image and video production, has integrated FLUX Upscale, the newly released AI video upscaling technology from Black Forest Labs. The integration gives creators a streamlined way to enhance existing and AI-generated videos at resolutions up to native 4K. Black Forest Labs introduced FLUX Upscale on August 20, 2026 as a standalone FLUX Tool and API endpoint. Designed for video inputs starting from 480p, the model regenerates footage at higher resolutions rather than relying only on conventional pixel enlargement. This approach is especially relevant to AI-generated video, where simply increasing resolution may leave existing visual defects untouched. FLUX Upscale is designed to improve common issues such as blurred facial details and grid-like artifacts that can appear in complex textures including water and grass. By adding FLUX Upscale to PixPix, creators can access the new video enhancement capability as part of a broader image and video production environment, without having to build and manage a separate API workflow. FLUX Upscale Offers Precise and Creative Modes FLUX Upscale provides two processing modes for different enhancement needs. Precise mode uses four processing steps and focuses on maintaining identity, composition and reference details more consistently. It is suited to footage involving people, products, branded assets and other subjects where preserving the original appearance is important. Creative mode uses eight processing steps and allows stronger detail reconstruction. It can be useful for landscapes, environmental textures and AI-generated scenes that require more extensive visual repair. Because Creative mode introduces more newly generated detail, it may also produce greater changes to some reference elements. The model supports 1.5x, 2x and 3x upscale factors. Depending on the source resolution and selected factor, videos can be regenerated at progressively higher resolutions, with supported workflows reaching up to native 4K. A More Accessible 4K AI Video Workflow As AI-generated video becomes more widely used in advertising, e-commerce, social media and commercial content production, creators increasingly need practical ways to prepare generated footage for higher-resolution delivery. PixPix combines AI image generation, video generation, editing and other creative tools in one workspace. With FLUX Upscale integrated into the platform, users can move from content generation to video enhancement within the same broader creative workflow. Potential use cases include improving AI-generated advertising footage, enhancing product videos, preparing social media creative for larger displays, and refining generated scenes before they enter a wider editing or campaign process. The workflow may also be useful for creators producing short-form video, product demonstrations and AI-generated campaign assets, where higher-resolution output provides greater flexibility for editing, cropping and distribution across different screen sizes and publishing platforms. FLUX Upscale is not limited to videos created with FLUX models. Black Forest Labs states that the standalone tool can process supported video inputs from 480p and above, making the technology applicable to a wider range of existing and AI-generated footage. The integration expands PixPix's video capabilities as generative and post-production models continue to reshape visual content workflows. It also provides creators with another option for improving the quality of selected AI-generated clips before they are used in campaigns, storefronts, social content or other production environments. Creators can explore PixPix and its AI image and video tools at https://www.pixpix.com/. About PixPix PixPix is an AI creative workspace for image, video and commercial content production. The platform brings AI generation, editing and production tools into one environment, helping brands, sellers, marketers and creators turn ideas and source assets into finished visual content more efficiently. Legal Disclaimer: EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.
FLUX 3 Video Part 1 drops: what the August release actually delivers. Black Forest Labs released FLUX 3 Video Part 1 on August 4, 2026. Here is what the generation capabilities include, how it compares to Veo 3.1 and Runway, and what is missing from the roadmap. #flux 3 #ai video #black forest labs #video generation Some links are partner links: if you subscribe through them, we may earn a commission, at no extra cost to you. The crowd verdicts stay independent. FLUX 3 Video Part 1 launched on August 4, 2026. Black Forest Labs, the team behind FLUX.2 (71% crowd approval in the image arena), shipped generation capabilities for text-to-video, image-to-video and video continuation, with native audio and clips up to 20 seconds. The headline numbers FLUX 3 Video generates HD clips up to 20 seconds with native audio. Early testers preferred it over Runway Gen-4.5 in 77% of head-to-heads and over Luma Ray 3.2 in 93%. No pricing yet. What Part 1 generation actually includes. The Part 1 release covers five input modes: * Text-to-video: describe a scene in plain language; the model handles motion, scene logic and audio together. * Image-to-video and keyframes: start from a still image, specify an end frame, or set multiple keyframes for controlled transitions. * Video continuation: extend an existing 4-second clip with coherent audio and visuals. * Multiple shots: generate varied scenes and camera angles within a single prompt. Audio is baked into the video generation, not added as a post-process. The model produces dialogue in English, Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi and Punjabi with lip-synced output. Technical specs at launch. | Spec | Value | | Max duration | 20 seconds | | Resolution at launch | 720p (1080p rolling out) | | Native audio | Yes (dialogue, SFX, ambient) | | Input modes | text, image, video, audio | | Reference images | up to 10 | Black Forest Labs also added a Draft Mode for fast iteration at lower cost before committing to a full render. How it stacks up against competitors. The AI video field in August 2026 has clear leaders. Here is how FLUX 3 compares to what is already shipping: | Model | From | Resolution | Max duration | Native audio | | FLUX 3 Video | TBA | 720p (1080p soon) | 20s | Yes | | Google Veo 3.1 | $20/mo | 1080p | 8s (148s extended) | Yes (48kHz spatial) | | Runway Gen-4.5 | $12/mo | 1080p | 10s | Yes | | Seedance 2.0 | $15/mo | 2K | 15s (30s on 2.5) | Yes | | Kling 3.0 | $7/mo | 4K | 15s | Yes | In early blind tests, FLUX 3 beat Runway Gen-4.5 77% of the time and Luma Ray 3.2 93% of the time. It ties Seedance 2.0 in image-to-video and "surpasses other alternatives," per the official announcement. The 20-second clip ceiling matters. Kling AI caps at 15 seconds, Veo 3.1 at 8 seconds per generation. For single-generation length, FLUX 3 is already competitive, though Scene Extension on Veo chains clips up to 148 seconds. The open-weight champion: 32B FLUX.2 wins 66.6% of head-to-heads vs open rivals, API from $0.014 per image $3/mo 71%(3) Partner link. The crowd verdicts stay independent. What is missing from Part 1. The August release is generation only. The FLUX 3 roadmap promised four pillars: image, video, audio and action (robotics). Part 1 ships video. What is not here yet: * FLUX 3 Image: editing and generation, "coming weeks" per the blog post. * FLUX 3 Dev (open weights): planned for later in 2026, no date locked. * Pricing: the BFL pricing page lists FLUX.2 and FLUX Tools, but FLUX 3 has no row yet. Access today is through the BFL API and select partners. The consumer web app and open weights are downstream. The pricing gap. Black Forest Labs has not published FLUX 3 pricing. For context, here is what FLUX.2 costs: * FLUX.2 [klein]: $0.014/image (4B) to $0.015/image (9B) * FLUX.2 [pro]: $0.03/image * FLUX.2 [max]: $0.07/image (first megapixel) About $3 covers 100 [pro] images. Video will cost more per second than images cost per frame, but the BFL team has historically undercut competitors on price. The 71% crowd approval for FLUX.2 in the AI image leaderboard suggests early trust. The open-weight champion: 32B FLUX.2 wins 66.6% of head-to-heads vs open rivals, API from $0.014 per image $3/mo 71%(3) Partner link. The crowd verdicts stay independent. The verdict. * Part 1 is real: 20-second clips with native audio, 5 input modes, multilingual lip sync. * Performance is strong: 77% win rate vs Runway Gen-4.5 in blind tests, 93% vs Luma Ray 3.2. * Gaps remain: no pricing, 720p at launch, no open weights until later in 2026. * For FLUX.2 users: the image model (71% recommend, $3/100 images) is the entry point. FLUX 3 Video builds on that foundation. The August release positions Black Forest Labs as a serious contender against Veo, Runway and Seedance. Pricing and open weights will determine whether that translates to adoption. Check the AI video leaderboard for live crowd scores as the rollout continues. Keep exploring Every claim above is backed by the arena's live data: crowd votes, verified pricing, honest pros & cons.
Black Forest Labs opens FLUX 3 Video API for 20-second Full HD clips. Key points * FLUX 3 Video is available through Black Forest Labs' dashboard and API, generating clips up to 20 seconds in HD or Full HD.[[1]] * Text-to-video and image-to-video cost $0.17 per second in HD or $0.29 per second in Full HD; synchronized audio is included.[[2]] * BFL's ranking above Seedance 2.0 and Gemini Omni Flash is a vendor benchmark, and its earlier head-to-head tests showed only a 52% preference for FLUX 3 against each rival.[[3]] Black Forest Labs has made FLUX 3 Video available through its dashboard, API and partner platforms, moving its first video-generation model beyond the early-access program announced on July 23. The system produces HD or Full HD clips lasting up to 20 seconds and can generate dialogue, sound effects and ambience alongside the frames.[[1]] The release supports text-to-video, image-to-video, keyframe transitions and video continuation. BFL also promotes multi-scene generation, typography embedded in moving scenes and multilingual dialogue with synchronized lip movement. Its public product and API pages, however, describe the speech support as multilingual without publishing a tested language list, so the reported claim of more than 14 languages remains vendor-supplied.[[1]] [[4]] Longer clips, familiar competitive claims. FLUX 3 costs $0.17 per second for standard HD text-to-video or image-to-video and $0.29 per second for Full HD. That puts a 20-second generation at $3.40 or $5.80, respectively. Draft-mode HD costs $0.06 per second, and audio carries no additional charge.[[2]] BFL's Elo tables rank FLUX 3 first for text-to-video and image-to-video, ahead of Seedance 2.0 and Google's Gemini Omni Flash.[[4]] Those results are not independent benchmarks. In BFL's preliminary early-access comparison of 10-second, 720p clips, raters preferred FLUX 3 over Seedance and Gemini Omni in 52% of comparisons, close to an even split.[[3]] The feature list also overlaps with those rivals. ByteDance says Seedance 2.0 generates multi-shot audio-video clips up to 15 seconds and accepts text, images, video and audio references.[[5]] Google's preview Gemini Omni Flash creates three- to 10-second 720p videos with audio, readable text and conversational editing.[[6]] FLUX 3's clearest documented distinction is its longer single-generation limit, rather than an independently established quality lead. Companies mentioned. At the intersection of AI, tech, and markets. The stories that matter, in one email. Free - unsubscribe anytime.
Black Forest Labs releases FLUX 3: A multimodal flow model for image, video, audio and robot action prediction. July 26, 2026 Black Forest Labs (BFL) has released FLUX 3, a multimodal foundation model that learns from images, videos and audio inside a single architecture. It is also the first FLUX model to ship video, audio and action prediction from one set of weights. The Black Forest Labs (BFL) research team argues that no single modality gives a complete description of the world. Images capture spatial structure at one instant. Video restores time and exposes physical dynamics. Audio reveals causal relationships between mechanical events and sound. Each is treated as a lossy projection of the same underlying reality. Training on all of them at once means the modalities constrain each other. The sound has to match the impact. The motion has to obey the mass. The research team calls FLUX 3 its first model built entirely on that principle. The method underneath: Self-Flow. FLUX 3 builds on Self-Flow, BFL's method for aligning multimodal generation and understanding in one architecture. Self-Flow combines the flow matching objective with a self-supervised feature reconstruction objective. The reference implementation on GitHub is Apache-2.0 and uses SiT-XL/2 with per-token timestep conditioning. It trains with a 25% per-token mask ratio and self-distillation from an EMA teacher at layer 20 to a student at layer 8. That released checkpoint is an ImageNet 256x256 research model, not FLUX 3. BFL states that it 'significantly scaled up compute and data resources' on the same approach to train FLUX 3 across video, images and audio simultaneously. Self-Flow itself was introduced in March 2026, so it is not new to this launch. What is new is the scale. What FLUX 3 Video does. FLUX 3 Video generates clips up to 20 seconds long in a single generation, with native audio. The supported modes cover text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video for controlled transitions, and generative video-audio continuation from input video and audio. BFL also lists multilingual dialogue, agentic chaining of clips into multi-shot sequences, and strong typography generation with animated designs. The BFL team reports particular strength in human facial expressions and in associating sounds with physical events. Performance. BFL team published preliminary human preference results. The setup was 10-second text-to-video clips at 720p with audio. FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%. Against Grok Imagine Video the figure is up to 69%, then Kling v3 Pro at 60%, Happy Horse v1 at 59% and Happy Horse 1.1 at 57%. Against Seedance 2.0 and Gemini Omni Flash the result is 52%, close to a coin flip. Interactive explorer. bfl@ flux-3 :~/real-world-models Early Access flux3 $ flux3 -spec One backbone, four output types. * Model class - Multimodal flow matching foundation model * Joint training modalities - image + video + audio, one architecture * Extended output - Action prediction (robot state) * Max video length - 20 seconds in a single generation * Audio - Native, generated jointly with video * Video modes - text-to-video, image-to-video, video-to-video, keyframe-to-video, video+audio continuation * Other capabilities - multilingual dialogue, agentic multi-shot chaining, typography and animated design * Announced - 23 July 2026, Black Forest Labs * Pricing - not published BFL describes each modality as a lossy projection of one underlying reality. Training on all of them at once lets their mutual constraints supervise each other. flux3 $ enter
FLUX 3: one AI model for Video, audio, and factory robots. On July 23, 2026, Black Forest Labs (BFL) launched FLUX 3, and the model is unlike anything the company has released before. Where earlier FLUX models stopped at still images, FLUX 3 is trained simultaneously on images, video, and audio inside a single architecture. The same backbone that generates a 20-second clip with synchronized audio is also driving factory robots at Audi. For enterprise teams building creative workflows or physical AI systems, that convergence is the story worth understanding. One architecture, four product lines. BFL frames FLUX 3 around a thesis it calls "visual intelligence": models that can perceive, predict, and act across physical and digital environments. The claim is that a model must learn a representation of the world, not just its still-frame snapshots, to generate convincing video or reliable robot actions. FLUX 3 is the first public result of that research direction. The model ships in four product lines: | Product | Capability | Status (July 25, 2026) | | FLUX 3 Video | 20-second clips with synchronized audio, text-to-video and image-to-video | Gated early access | | FLUX 3 Image | Advanced image synthesis and editing | Coming in weeks | | FLUX 3 Action | Action prediction for robotics, starting with FLUX-mimic | Gated early access | | FLUX 3 Dev | Open-weight multimodal backbone (video, audio, image, action) | Later in 2026 | Early access for Video and Action is available by application at bfl.ai. FLUX 3 Image and open-weight Dev access follow on rolling timelines that BFL has not pinned to specific dates. What FLUX 3 Video actually does. FLUX 3 Video generates clips up to 20 seconds long, with audio produced alongside the visual content and matched to what happens on screen: dialogue, sound effects, ambient noise. In early human-preference evaluations, FLUX 3 Video was preferred over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93%, according to BFL's launch materials. The model also handles multilingual dialogue and supports agentic chaining of individual clips into longer multi-shot sequences while maintaining character and visual reference consistency. Enterprise creative teams running campaigns at scale should note the workflow implication: a single generation call can cover the video, its audio track, and the keyframe images from which both can be edited, without handing off between separate vendor APIs. The company's existing distribution makes that consolidation credible. Earlier FLUX models power generative features in Adobe Photoshop, Picsart, and Nous Research's Hermes Agent. More than 500 million downloads of FLUX models have been recorded to date, per the company's press release. FLUX 3 Video is already in early testing with Canva, Burda, Magnific (formerly Freepik), Krea, and Picsart. Physical AI: from content to factory floor. The part of FLUX 3 that signals a longer strategic bet is FLUX-mimic. BFL partnered with Swiss robotics firm mimic robotics to build a video-action model on top of the FLUX 3 backbone. The model is being tested and deployed in Audi production facilities for tasks that have resisted conventional automation: fitting flexible door seals, kitting parts into structured trays, inserting electronic control units into tight-fitting fixtures, and handling soft, flexible materials. The efficiency claim is the headline: FLUX-mimic can be fine-tuned for a specific manipulation task with as little as 30 minutes of robot data. Prior approaches required 30 or more hours. BFL attributes this to the FLUX 3 backbone already encoding how the physical world behaves, so the robot only needs to learn how a specific task maps onto knowledge it already holds. The full system reacts in approximately 101 milliseconds, in the range of human visual reaction time. "We have seen these robots solve complex soft-body manipulation work that would have been simply impossible with conventional robotics," Christoph Schneider of Audi Production Lab said in BFL's announcement. "This can have a major impact in assisting our employees, increasing efficiency, and expanding flexible automation across production and logistics operations." For enterprise buyers outside manufacturing, the Audi deployment is a proof point worth tracking: it is the first public case of a foundation model trained on consumer video content being transferred, with minimal task-specific data, into a production industrial setting. Why the unified architecture matters for enterprise buyers. Most enterprise creative stacks today involve four to six distinct vendors: one for image generation, one for video, one for audio, one for 3D rendering or product visualization, and so on. Each vendor requires its own integration, contract, and fine-tuning dataset. FLUX 3's unified architecture is a direct challenge to that structure. VentureBeat's coverage quotes BFL's pitch to enterprise software companies: "A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models." For GTM teams and agencies running content at scale, the practical benefit is fewer round-trips and less context loss when an asset moves from concept to finished format. For AI-native platforms embedding generative features, a single fine-tune on FLUX 3 that updates image, video, and audio quality simultaneously is meaningfully cheaper than maintaining separate model versions. This connects to a pattern Enera covered in the broader enterprise AI price war: as foundation model providers consolidate capabilities into fewer endpoints, the cost-per-workflow calculation shifts in ways that favor unified vendors over specialist ones. What is still missing. FLUX 3 launched without several things enterprise buyers typically need before committing: * No public pricing. BFL has not announced API costs, enterprise tiers, or SLAs. * No full benchmark methodology. Published benchmark results are expected alongside broader availability. * FLUX 3 Dev details are incomplete. License terms, parameter counts, quantization options, and hardware requirements for the open-weight version have not been disclosed. * FLUX 3 Image is not live yet. This is likely the most immediately useful tier for the creative teams already using FLUX 1 and FLUX 2. None of these gaps disqualify FLUX 3 as a story worth following. But they do mean that enterprise procurement conversations are premature until BFL releases pricing and the FLUX 3 Image tier goes live. What enterprise teams should watch. If you are running creative operations, product marketing, or physical AI initiatives, three things are worth tracking in the near term: * FLUX 3 Image launch. This is the earliest point at which most enterprise creative teams can evaluate the model in practice, not just in press materials. * FLUX 3 Dev license terms. The license will determine whether enterprises can run the model locally for sensitive workflows, which is how earlier FLUX Dev releases drove wide adoption. As Enera noted in its look at enterprise AI token efficiency, local deployment is increasingly a compliance and cost requirement, not just a preference. * Independent benchmark reproduction. BFL's own evaluations show strong results. Independent testing of FLUX-mimic's sample efficiency claims and FLUX 3 Video's head-to-head numbers will determine whether those numbers hold in realistic deployment conditions. BFL is valued at $3.25 billion and has raised more than $450 million from investors including a16z, NVIDIA, Salesforce Ventures, Adobe Ventures, Figma Ventures, Canva, and Deutsche Telekom's T.Capital. The investor list mirrors the distribution network: the platforms that invest in BFL also embed FLUX models into their products. That alignment makes FLUX 3 more likely to reach enterprise creative stacks quickly than a comparable model from a lab without those ties. If you want to evaluate how a unified visual AI foundation fits your current creative or physical AI stack, let's talk.