FLUX 3 Image is live through a Black Forest Labs-owned Replicate model and partner endpoints, with 4K output and editing from up to 10 references. BFL's native documentation still centers on FLUX 3 Video, while image providers use different schemas, limits, and billing.
Is the FLUX 3 Image API actually available?
The FLUX 3 Image API is available, but “official” needs a precise definition. The strongest evidence is the live black-forest-labs/flux-3-image listing owned by Black Forest Labs on Replicate. It accepts generation requests, switches into edit mode when an image is supplied, and exposes 4k as a resolution value.
| Surface checked on October 2, 2026 | What is live | What it proves |
|---|---|---|
| BFL on Replicate | black-forest-labs/flux-3-image | BFL-owned model; text generation, editing, 4K, and up to 10 references |
| fal partner endpoint | blackforestlabs/flux-3/edit-image | Commercial edit endpoint, 1-10 references, queued API, resolution-based billing |
| Layer API docs | bfl-flux-3-image | 1K/2K/4K generation and editing through an asynchronous workspace API |
| BFL native API docs | FLUX 3 Video documented | No equivalent native FLUX 3 Image route was listed when checked |
flux3api.com and community wrappers | Separate third-party services | A matching name does not establish BFL ownership or current FLUX 3 Image access |
The BFL help article for FLUX 3 describes only the video model; the separate BFL-owned Replicate listing and partner endpoints verify image editing.
Before the endpoint appeared, Reddit user u/rerri predicted an API-first release:
“I wouldn't be surprised if the API-only Flux 3 Image launches first.” — u/rerri on r/StableDiffusion
The rollout fits that prediction, but API access does not mean open weights are available.
What 4K and multi-reference editing mean in practice
FLUX 3 Image exposes a 4k output option and accepts as many as 10 reference images. Neither field guarantees that a result will preserve every identity, product detail, or small line of text; the provider pages publish controls and examples, not independent quality scores.
The BFL-owned Replicate README lists 768sq, 1k, 1.5k, 2k, and 4k. Reference files may be JPEG, PNG, GIF, or WebP, must be at least 256 by 256 pixels, and may be no larger than 16 megapixels. With aspect_ratio: auto, the first reference controls the edit's aspect ratio.
The fal edit schema is similar but not identical. It accepts 1-10 URLs or data URIs, limits each input to 4 megapixels, supports 512sq through 4k, and warns that 4K may take several minutes. Reference order is semantic: “image 1” means the first item in image_urls.
| Control | Replicate | fal | Production consequence |
|---|---|---|---|
| Maximum references | 10 | 10 | Number inputs explicitly in the prompt |
| Maximum input size | 16 MP | 4 MP per image | Validate before routing to a provider |
| Output choices | 768sq, 1K, 1.5K, 2K, 4K | 512sq, 768sq, 1K, 2K, 4K | Do not share one unvalidated enum across providers |
| Automatic ratio | First reference guides ratio | First reference guides ratio | Put the framing reference first |
| Output formats | WebP, JPG, PNG | JPEG, PNG | Normalize downstream file handling |
| 4K latency disclosure | No measured latency published | May take several minutes | Keep 4K out of interactive preview paths |
For multi-reference edits, assign one role to each input: base composition, subject identity, product, or style. fal's own guidance recommends one edit per request. A prompt such as “Use image 1 as the base; replace only the bottle with the product from image 2; preserve camera angle, hands, lighting, and background” is easier to inspect than a request that also changes wardrobe, typography, and location.
A practical queued API workflow
A production FLUX 3 Image API workflow should treat generation as an asynchronous job. The application uploads stable input URLs, submits a narrow request, stores the provider's request ID, polls with backoff, and copies the completed output to its own storage.
The following example uses fal's documented endpoint identifier and request fields. It is an integration template, not a claim that the request below was executed during this review.
import os
import time
import requests
ENDPOINT = "https://queue.fal.run/blackforestlabs/flux-3/edit-image"
headers = {
"Authorization": f"Key {os.environ['FAL_KEY']}",
"Content-Type": "application/json",
}
payload = {
"prompt": (
"Use image 1 as the base. Replace only its package with the product "
"from image 2. Preserve the hands, camera angle, shadows, and background."
),
"image_urls": [
"https://cdn.example.com/base.jpg",
"https://cdn.example.com/product.png",
],
"resolution": "1k",
"aspect_ratio": "auto",
"output_format": "png",
"safety_tolerance": 2,
}
submitted = requests.post(ENDPOINT, headers=headers, json=payload, timeout=30)
submitted.raise_for_status()
job = submitted.json()
status_url = job["status_url"]
response_url = job["response_url"]
while True:
status = requests.get(status_url, headers=headers, timeout=30)
status.raise_for_status()
state = status.json().get("status")
if state == "COMPLETED":
break
if state in {"FAILED", "CANCELLED"}:
raise RuntimeError(status.text)
time.sleep(2)
result = requests.get(response_url, headers=headers, timeout=30)
result.raise_for_status()
print(result.json())
The fal queue documentation linked from the model page also exposes sync_mode, but queued execution is the safer default for 4K because a render can outlive a normal HTTP request timeout. Layer makes this asynchronous contract explicit: submission returns HTTP 202, an inference_id, and a recommended polling interval. Layer also supports idempotency keys replayable for 24 hours, which helps prevent duplicate charges after network retries.
Before enabling traffic:
- Reject images below 256 pixels per side and enforce the selected provider's megapixel ceiling.
- Preserve array order and generate prompts that refer to
image 1,image 2, and so on. - Use a unique idempotency key where the provider supports it; otherwise persist the request before retrying.
- Cap polling time and surface a pending state instead of holding an application request open.
- Copy finished files to controlled storage because hosted result URLs may not match the application's retention policy.
- Log model ID, provider, resolution, reference count, quoted cost, elapsed time, and moderation outcome for every job.
The honest cost and quality trade-off
The defensible cost comparison is limited because providers did not publish a complete per-resolution table on the pages checked. fal advertised a promotional 1K price of $0.024 per image, rising to $0.048 after the promotion; it also stated that reference count does not change the charge. Exact 2K and 4K prices were not printed on that model page, so a 4K budget cannot be derived from the 1K figure.
Use a two-stage policy instead of assuming 4K is always the best setting:
| Stage | Resolution | Purpose | Promotion rule |
|---|---|---|---|
| Prompt and reference validation | 1K | Check composition, identity, product shape, and text | Reject or revise before expensive output |
| Final asset | 2K or 4K | Produce the approved deliverable | Promote only when the target channel needs those pixels |
Higher resolution buys pixels, not better edit fidelity: a bad 1K edit becomes a larger failure at 4K. Reserve 4K for approved edits destined for print, billboard layouts, or aggressive crops.
At application startup, send a minimal valid test job or query the provider's pricing surface, capture the quoted charge, and disable 4K if the quote is missing or exceeds the job budget. Layer's initial response may include estimated_price_creative_units; its public model page did not provide a dollar conversion. Replicate's retrieved model page documented inputs but no fixed price. Those are procurement gaps to resolve in the account dashboard before launch, not numbers to guess in code.
Choose the endpoint by operational fit
Provider choice should follow the contract your application needs. Model ownership alone does not make schemas interchangeable.
- Replicate: choose the BFL-owned listing when provenance is the priority and your stack already uses Replicate's prediction workflow. It documents the broadest input ceiling here at 16 MP and includes optional web/image grounding.
- fal: choose the partner edit endpoint when clear image-edit controls, a queue workflow, and a visible 1K price are more useful. Its 4 MP input ceiling requires earlier downscaling.
- Layer: choose it for workspace organization, a formal HTTP
202contract, polling hints, and 24-hour idempotency. Confirm how Creative Units convert to dollars before setting a budget.
Do not identify a provider from the word “FLUX3” in its domain or repository name. Verify the model ID, owner or partner label, current enum values, commercial terms, and a successful low-cost request. The high-ranking Anil-matcha/Flux-3-Dev-API wrapper still labeled its image routes “coming soon” when checked, while the BFL-owned Replicate and fal partner routes were live.
Production release gate
FLUX 3 Image is suitable for controlled API testing, including 4K and up to 10 references. Ship only after the selected endpoint passes the same representative edit set at 1K and final resolution.
| Check | Pass condition |
|---|---|
| Provenance | Exact BFL-owned or verified partner model ID |
| Availability | A real low-cost request completes, not merely a documented route |
| Reference behavior | Input order and role labels survive representative 2-, 5-, and 10-image cases |
| Quality | Identity, product geometry, text, and untouched regions meet defined review thresholds |
| Cost | Provider returns or displays an acceptable price for every enabled resolution |
| Latency | Measured queue and render times fit preview and batch service targets |
| Reliability | Retries cannot create untracked duplicate jobs or charges |
| Storage | Outputs are copied before provider URLs expire or policies change |
The practical recommendation is to launch 1K editing first, log quote and latency data, and enable 2K or 4K only for approved finals. That keeps the new model's strongest documented capabilities available without making an unverified assumption about high-resolution quality or cost.