images and video model

FLUX — an image and video generator with sound

FLUX models by Black Forest Labs in GPTunneL: photorealism, readable text in frame and video generation with sound. Pay per generation, no subscription.

FLUX 3 — video with sound on the first take

The multimodal generation of the family: one model draws the picture, shoots the clip and scores it — in a single render, no editing pass.

Up to 20 secondsSound in one render720p and 1080pCharacter dialogueDrafts cost a third
A reel on FLUX 3 multimodality — image, video and sound from one model

What FLUX 3 can do

One model instead of a set of separate ones. Pick a mode and the clip on stage changes.

Sound in frame

The clips were generated in FLUX 3 — samples by Black Forest Labs.

What the FLUX family is good at

Traits Black Forest Labs keeps from one generation to the next — and with FLUX 3 they work in video too.

Photorealism

This is what the family is picked for above all: skin, fabric, reflections and light look photographed rather than drawn.

Readable text in frame

FLUX was one of the first to write without turning letters into mush — first on stills, now in motion, in titles and on signage.

Video with sound

FLUX 3 shoots the clip and scores it in one render: character speech, the noise of the scene and music arrive with the picture.

Edits by description

The Kontext models change a finished image on a text request: swap the background, redress the character, remove what you don't need.

A frame as the starting point

The family works from text and from a finished image alike — a shot can be brought to life and a character moved into a new scene.

Open weights

Part of the family is published and can run on your own hardware — a rarity among generators of this level.

Which model to pick

Three tiers for stills and a separate model for video.

devFast draftsproThe default pickultra / maxMaximum qualityFLUX 3Video with sound
OutputStill imageStill imageStill imageVideo with sound
Best forTrying ideas, draft framesFinished images for a site or social mediaPrint, large formats, complex scenesClips up to 20 seconds, scenes with dialogue
DetailMediumHighMaximum720p or 1080p
SpeedHighestHighMediumDepends on clip length
PriceLowestModerateAbove averagePer second of video

The tiers differ in detail and price per image; FLUX 3 is billed per second of video. Exact per-model prices are on the Pricing

Every FLUX model

Release dates follow Black Forest Labs' official announcements.

July 2026

FLUX 3

The multimodal generation: one model for images, video and sound. A clip of up to 20 seconds is generated together with its soundtrack — character lines, foley and music.

November 2025

FLUX.2 [pro]

The current generation: noticeably better at holding a complex prompt together and more accurate at writing text on the image.

July 2025

FLUX.1 [krea]FLUX.1 [srpo]

A version made together with Krea AI: livelier and more varied in aesthetics, with less of the «AI gloss».

May 2025

FLUX.1 Kontext [max]FLUX.1 Kontext [pro]

The generation that can not only draw from scratch but also edit a finished image on a text request.

October — November 2024

FLUX 1.1 [ultra]FLUX 1.1 [pro]

A faster generation plus an ultra version with a higher-resolution frame.

August 2024

FLUX.1 [dev]

The first family from the company founded by the authors of Stable Diffusion. Today's bar for photorealism starts here.

How billing works

There is no subscription: you top up one balance and spend it on any model on the platform.

Pay per generation

Images are charged per frame, video by clip length and resolution. Not using it costs nothing: no limits, no monthly fee.

One balance for every model

FLUX, ChatGPT, Claude, image and video generation — all out of the same wallet. No separate subscription per service.

Top up the way you prefer

International cards, Apple Pay, Google Pay or crypto. The minimum top-up is $5.

Sample video

The clips were generated in FLUX 3 from a text description. Hover a tile for a muted preview; a click opens the clip with sound.

A dinosaur on a night streetcity · vfx
Dancers made of splashing paintgraphics · motion
A wolf in a snowy forestwildlife
A harpist in an ice cavelight · music
Street food, fire in the wokreportage
Droplets forming a logotitles · graphics
A street during an uprisinghistorical scene
An anime rooftop shotanimation

Sample images

Every frame was generated in FLUX from an ordinary text description. The more detail you give, the closer the result gets to what you had in mind — there is no need to hand-pick tags.

How to get started with FLUX

Sign in to GPTunneL
One account for every model on the platform.
A giant fish-shaped pastry towering over a night city street, passers-by filming it on their phones
FLUX.2 [pro]
A frame generated in FLUX from a text description

Sign in whichever way suits you.

The model switches right inside the input — you can change it between frames.

The finished frame arrives at full resolution.

Try FLUX in GPTunneL

Signing up takes a minute, and $5 on the balance is enough to see whether the model fits your task.

Frequently asked questions

No. Requests go through GPTunneL's infrastructure, so the models open like any ordinary website — no VPN and no proxy.

FLUX is a family of image generators from the German company Black Forest Labs, founded by the authors of Stable Diffusion. The model draws a picture from a text description and edits a finished one on request — swap the background, remove an object, redress the character.

The family is picked for two things. First, photorealism: skin, fabric, reflections and light look photographed rather than drawn. Second, text on the image: FLUX was one of the first to write letters without turning them into mush, which is why it is used for signage, packaging and ad layouts. Some models are published with open weights and can run on your own hardware — a rarity among generators of this level.

With the FLUX 3 generation the family moved beyond the still frame. It is a single multimodal model for images, video and sound: it shoots a clip of up to 20 seconds from a description or from a finished frame and scores it right away — character speech, the noise of the scene and music. There is no need to generate the video and the audio separately and then sync them.

FLUX models are available in GPTunneL — worldwide, without a VPN, with one balance and billing per generated frame and per second of video.