AIREITER

FLUX 3 Image API: 4K and Multi-Reference Guide

Last Updated: 2026-10-02 00:26:48

FLUX 3 Image is live through a Black Forest Labs-owned Replicate model and partner endpoints, with 4K output and editing from up to 10 references. BFL's native documentation still centers on FLUX 3 Video, while image providers use different schemas, limits, and billing.

Is the FLUX 3 Image API actually available?

The FLUX 3 Image API is available, but “official” needs a precise definition. The strongest evidence is the live black-forest-labs/flux-3-image listing owned by Black Forest Labs on Replicate. It accepts generation requests, switches into edit mode when an image is supplied, and exposes 4k as a resolution value.

Surface checked on October 2, 2026What is liveWhat it proves
BFL on Replicateblack-forest-labs/flux-3-imageBFL-owned model; text generation, editing, 4K, and up to 10 references
fal partner endpointblackforestlabs/flux-3/edit-imageCommercial edit endpoint, 1-10 references, queued API, resolution-based billing
Layer API docsbfl-flux-3-image1K/2K/4K generation and editing through an asynchronous workspace API
BFL native API docsFLUX 3 Video documentedNo equivalent native FLUX 3 Image route was listed when checked
flux3api.com and community wrappersSeparate third-party servicesA matching name does not establish BFL ownership or current FLUX 3 Image access
Black Forest Labs FLUX 3 Image model page on Replicate

The BFL help article for FLUX 3 describes only the video model; the separate BFL-owned Replicate listing and partner endpoints verify image editing.

Before the endpoint appeared, Reddit user u/rerri predicted an API-first release:

“I wouldn't be surprised if the API-only Flux 3 Image launches first.” — u/rerri on r/StableDiffusion

The rollout fits that prediction, but API access does not mean open weights are available.

What 4K and multi-reference editing mean in practice

FLUX 3 Image exposes a 4k output option and accepts as many as 10 reference images. Neither field guarantees that a result will preserve every identity, product detail, or small line of text; the provider pages publish controls and examples, not independent quality scores.

The BFL-owned Replicate README lists 768sq, 1k, 1.5k, 2k, and 4k. Reference files may be JPEG, PNG, GIF, or WebP, must be at least 256 by 256 pixels, and may be no larger than 16 megapixels. With aspect_ratio: auto, the first reference controls the edit's aspect ratio.

The fal edit schema is similar but not identical. It accepts 1-10 URLs or data URIs, limits each input to 4 megapixels, supports 512sq through 4k, and warns that 4K may take several minutes. Reference order is semantic: “image 1” means the first item in image_urls.

ControlReplicatefalProduction consequence
Maximum references1010Number inputs explicitly in the prompt
Maximum input size16 MP4 MP per imageValidate before routing to a provider
Output choices768sq, 1K, 1.5K, 2K, 4K512sq, 768sq, 1K, 2K, 4KDo not share one unvalidated enum across providers
Automatic ratioFirst reference guides ratioFirst reference guides ratioPut the framing reference first
Output formatsWebP, JPG, PNGJPEG, PNGNormalize downstream file handling
4K latency disclosureNo measured latency publishedMay take several minutesKeep 4K out of interactive preview paths

For multi-reference edits, assign one role to each input: base composition, subject identity, product, or style. fal's own guidance recommends one edit per request. A prompt such as “Use image 1 as the base; replace only the bottle with the product from image 2; preserve camera angle, hands, lighting, and background” is easier to inspect than a request that also changes wardrobe, typography, and location.

A practical queued API workflow

A production FLUX 3 Image API workflow should treat generation as an asynchronous job. The application uploads stable input URLs, submits a narrow request, stores the provider's request ID, polls with backoff, and copies the completed output to its own storage.

The following example uses fal's documented endpoint identifier and request fields. It is an integration template, not a claim that the request below was executed during this review.

import os
import time
import requests

ENDPOINT = "https://queue.fal.run/blackforestlabs/flux-3/edit-image"
headers = {
    "Authorization": f"Key {os.environ['FAL_KEY']}",
    "Content-Type": "application/json",
}
payload = {
    "prompt": (
        "Use image 1 as the base. Replace only its package with the product "
        "from image 2. Preserve the hands, camera angle, shadows, and background."
    ),
    "image_urls": [
        "https://cdn.example.com/base.jpg",
        "https://cdn.example.com/product.png",
    ],
    "resolution": "1k",
    "aspect_ratio": "auto",
    "output_format": "png",
    "safety_tolerance": 2,
}

submitted = requests.post(ENDPOINT, headers=headers, json=payload, timeout=30)
submitted.raise_for_status()
job = submitted.json()

status_url = job["status_url"]
response_url = job["response_url"]
while True:
    status = requests.get(status_url, headers=headers, timeout=30)
    status.raise_for_status()
    state = status.json().get("status")
    if state == "COMPLETED":
        break
    if state in {"FAILED", "CANCELLED"}:
        raise RuntimeError(status.text)
    time.sleep(2)

result = requests.get(response_url, headers=headers, timeout=30)
result.raise_for_status()
print(result.json())

The fal queue documentation linked from the model page also exposes sync_mode, but queued execution is the safer default for 4K because a render can outlive a normal HTTP request timeout. Layer makes this asynchronous contract explicit: submission returns HTTP 202, an inference_id, and a recommended polling interval. Layer also supports idempotency keys replayable for 24 hours, which helps prevent duplicate charges after network retries.

Before enabling traffic:

  1. Reject images below 256 pixels per side and enforce the selected provider's megapixel ceiling.
  2. Preserve array order and generate prompts that refer to image 1, image 2, and so on.
  3. Use a unique idempotency key where the provider supports it; otherwise persist the request before retrying.
  4. Cap polling time and surface a pending state instead of holding an application request open.
  5. Copy finished files to controlled storage because hosted result URLs may not match the application's retention policy.
  6. Log model ID, provider, resolution, reference count, quoted cost, elapsed time, and moderation outcome for every job.

The honest cost and quality trade-off

The defensible cost comparison is limited because providers did not publish a complete per-resolution table on the pages checked. fal advertised a promotional 1K price of $0.024 per image, rising to $0.048 after the promotion; it also stated that reference count does not change the charge. Exact 2K and 4K prices were not printed on that model page, so a 4K budget cannot be derived from the 1K figure.

FLUX 3 Image Edit API model page on fal

Use a two-stage policy instead of assuming 4K is always the best setting:

StageResolutionPurposePromotion rule
Prompt and reference validation1KCheck composition, identity, product shape, and textReject or revise before expensive output
Final asset2K or 4KProduce the approved deliverablePromote only when the target channel needs those pixels

Higher resolution buys pixels, not better edit fidelity: a bad 1K edit becomes a larger failure at 4K. Reserve 4K for approved edits destined for print, billboard layouts, or aggressive crops.

At application startup, send a minimal valid test job or query the provider's pricing surface, capture the quoted charge, and disable 4K if the quote is missing or exceeds the job budget. Layer's initial response may include estimated_price_creative_units; its public model page did not provide a dollar conversion. Replicate's retrieved model page documented inputs but no fixed price. Those are procurement gaps to resolve in the account dashboard before launch, not numbers to guess in code.

Choose the endpoint by operational fit

Provider choice should follow the contract your application needs. Model ownership alone does not make schemas interchangeable.

  • Replicate: choose the BFL-owned listing when provenance is the priority and your stack already uses Replicate's prediction workflow. It documents the broadest input ceiling here at 16 MP and includes optional web/image grounding.
  • fal: choose the partner edit endpoint when clear image-edit controls, a queue workflow, and a visible 1K price are more useful. Its 4 MP input ceiling requires earlier downscaling.
  • Layer: choose it for workspace organization, a formal HTTP 202 contract, polling hints, and 24-hour idempotency. Confirm how Creative Units convert to dollars before setting a budget.

Do not identify a provider from the word “FLUX3” in its domain or repository name. Verify the model ID, owner or partner label, current enum values, commercial terms, and a successful low-cost request. The high-ranking Anil-matcha/Flux-3-Dev-API wrapper still labeled its image routes “coming soon” when checked, while the BFL-owned Replicate and fal partner routes were live.

Production release gate

FLUX 3 Image is suitable for controlled API testing, including 4K and up to 10 references. Ship only after the selected endpoint passes the same representative edit set at 1K and final resolution.

CheckPass condition
ProvenanceExact BFL-owned or verified partner model ID
AvailabilityA real low-cost request completes, not merely a documented route
Reference behaviorInput order and role labels survive representative 2-, 5-, and 10-image cases
QualityIdentity, product geometry, text, and untouched regions meet defined review thresholds
CostProvider returns or displays an acceptable price for every enabled resolution
LatencyMeasured queue and render times fit preview and batch service targets
ReliabilityRetries cannot create untracked duplicate jobs or charges
StorageOutputs are copied before provider URLs expire or policies change

The practical recommendation is to launch 1K editing first, log quote and latency data, and enable 2K or 4K only for approved finals. That keeps the new model's strongest documented capabilities available without making an unverified assumption about high-resolution quality or cost.

Related reading