The fal genmedia CLI handles discovery, schema inspection, async request IDs, downloads, and JSON receipts; the missing layer is an orchestrator that bounds concurrency, retries transient failures, resumes completed work, and records in-flight IDs for async jobs.
Install once, then make every run reproducible
For macOS and Linux, the project README documents:
curl https://genmedia.sh/install -fsS | bash
genmedia setup --non-interactive --api-key "$FAL_KEY" --no-auto-update
Windows uses the documented PowerShell installer:
irm https://genmedia.sh/install.ps1 | iex
genmedia setup --non-interactive --api-key "$env:FAL_KEY"
The genmedia CLI README says genmedia supports FAL_KEY, JSON output, non-interactive setup, and an optional background update checker. In CI, inject the key through the runner’s secret store rather than placing it in a script or command log.
Before a batch, pin the endpoint and inspect its live contract:
genmedia models "image to video" --json
genmedia schema bytedance/seedance-2.0/image-to-video --json
genmedia pricing bytedance/seedance-2.0/image-to-video --json
The fal guide’s useful sequence is models → schema → run → status → download. Do not assume that flags from one video endpoint work on another: the workflow skill specifically recommends rechecking the schema after a validation error.
Design the batch as a queue, not a shell loop
The official genmedia materials document async execution, but they do not promise a native “read this CSV and run 500 rows” command or an automatic retry policy. Treat the CLI as the provider-aware runner and your wrapper as the queue controller.
A durable job record needs at least:
| Field | Why it matters |
|---|---|
id | Stable input identity for resume and deduplication |
endpoint | Exact model route used |
prompt | Reproducibility and audit |
status | pending, submitted, complete, failed, or skipped |
request_id | Required by genmedia status |
attempts | Stops runaway retries |
output | Deterministic local path |
error | Makes failed rows actionable |
Use --async for long image-to-video or video jobs. Save the returned request_id immediately, then poll with both the endpoint ID and request ID:
genmedia run bytedance/seedance-2.0/image-to-video \
--image_url "$IMAGE_URL" \
--prompt "Slow product turn on a studio table; no text or logo" \
--duration 4 --resolution 720p --aspect_ratio 16:9 \
--async --json > logs/shot-001-submit.json
REQUEST_ID=$(jq -r '.request_id' logs/shot-001-submit.json)
genmedia status bytedance/seedance-2.0/image-to-video "$REQUEST_ID" \
--download "outputs/{request_id}_{index}.{ext}" --json > logs/shot-001-result.json
The {request_id}_{index}.{ext} pattern reduces the risk of two jobs silently overwriting one another. Keep the JSON beside the media, not in a separate temporary directory.
Retry only failures that can recover
A practical starting policy is:
- Retry network timeouts, connection resets, HTTP 429, and transient 5xx responses.
- Back off exponentially, for example 2, 4, 8, 16, then 32 seconds, with a small random jitter.
- Cap attempts per row, such as four submissions or five status polls over a time window.
- Do not retry 401/403 authentication errors, 422 schema validation errors, safety refusals, or malformed JSON.
- If a submission times out after the request may have reached the provider, inspect the saved request record before creating a second paid request.
This policy belongs to your wrapper; the public genmedia README documents commands and lifecycle operations, not a guaranteed automatic retry behavior. For a 422, read validation_errors, run genmedia schema again, and fix the named field instead of blindly resubmitting.
A copyable Python batch wrapper
The wrapper below is a synchronous image-batch starter: it uses subprocess.run with an argument list, skips completed output paths, limits concurrent work, retries transient process failures, and writes an atomic manifest. For long video jobs, persist the async submission before polling: submit -> write endpoint/request_id -> restart -> poll saved request_id -> download; only submit again when no request ID exists.
#!/usr/bin/env python3
import concurrent.futures as pool
import json, os, random, subprocess, tempfile, threading, time
from pathlib import Path
ENDPOINT = "fal-ai/flux/dev"
OUT = Path("outputs/images")
LOG = Path("outputs/logs")
MAX_WORKERS = 3
MAX_ATTEMPTS = 4
MANIFEST_LOCK = threading.Lock()
JOBS = [
{"id": "shoe-001", "prompt": "Black running shoe, clean studio product photo", "file": "shoe-001.png"},
{"id": "shoe-002", "prompt": "Black running shoe on wet pavement at dawn", "file": "shoe-002.png"},
]
OUT.mkdir(parents=True, exist_ok=True)
LOG.mkdir(parents=True, exist_ok=True)
MANIFEST = Path("outputs/manifest.json")
PREVIOUS = json.loads(MANIFEST.read_text()) if MANIFEST.exists() else {"results": []}
STATE = {r["id"]: r for r in PREVIOUS.get("results", [])}
DONE = {k: r for k, r in STATE.items() if r.get("status") == "complete"}
MAX_REQUESTS = len(JOBS) * MAX_ATTEMPTS
if MAX_REQUESTS > 100:
raise SystemExit(f"request ceiling exceeded: {MAX_REQUESTS}")
def save_result(result):
with MANIFEST_LOCK:
STATE[result["id"]] = result
payload = {"endpoint": ENDPOINT, "max_requests": MAX_REQUESTS,
"results": list(STATE.values())}
fd, tmp = tempfile.mkstemp(dir=MANIFEST.parent, prefix="manifest.", text=True)
with os.fdopen(fd, "w") as f:
json.dump(payload, f, indent=2)
os.replace(tmp, MANIFEST)
TRANSIENT_WORDS = ("429", "500", "502", "503", "504", "timeout", "temporarily", "connection")
def run_one(job):
target = OUT / job["file"]
receipt = LOG / f"{job['id']}.json"
if job["id"] in DONE and target.exists() and target.stat().st_size > 0 and receipt.exists():
return DONE[job["id"]]
cmd = ["genmedia", "run", ENDPOINT, "--prompt", job["prompt"],
"--num_images", "1", "--download", str(target), "--json"]
last_error = ""
for attempt in range(1, MAX_ATTEMPTS + 1):
try:
p = subprocess.run(cmd, text=True, capture_output=True, timeout=900)
raw = p.stdout.strip()
if p.returncode != 0:
last_error = p.stderr[-1000:] or raw[-1000:]
if not any(w in last_error.lower() for w in TRANSIENT_WORDS):
break
if attempt < MAX_ATTEMPTS:
time.sleep((2 ** attempt) + random.random())
continue
try:
data = json.loads(raw) if raw else {}
except json.JSONDecodeError as exc:
return {**job, "status": "failed", "attempts": attempt, "error": f"invalid JSON: {exc}"}
if target.exists():
(LOG / f"{job['id']}.json").write_text(json.dumps(data, indent=2))
return {**job, "status": "complete", "attempts": attempt, "output": str(target)}
last_error = p.stderr[-1000:] or raw[-1000:]
if not any(w in last_error.lower() for w in TRANSIENT_WORDS):
break
except (subprocess.TimeoutExpired, OSError) as exc:
last_error = str(exc)
if attempt < MAX_ATTEMPTS:
time.sleep((2 ** attempt) + random.random())
return {**job, "status": "failed", "attempts": MAX_ATTEMPTS, "error": last_error}
results = []
with pool.ThreadPoolExecutor(max_workers=MAX_WORKERS) as executor:
futures = [executor.submit(run_one, job) for job in JOBS]
for future in pool.as_completed(futures):
result = future.result()
results.append(result)
save_result(result)
print(json.dumps(results, indent=2))
For video, replace ENDPOINT and add the endpoint’s schema-specific flags. For an image-to-video chain, upload the local image once with genmedia upload ./frame.png --json, pass the returned URL to the video job, and keep both records in the manifest. The wrapper does not claim that genmedia itself estimates or enforces a maximum spend; it only prevents duplicate local work and limits retries.
Put a cost gate before paid generation
genmedia pricing <endpoint_id> --json is a lookup, not a reservation or budget cap. Use it before the batch, then calculate a conservative ceiling from the number of rows, outputs per row, resolution/duration settings, and the maximum retry count.
A practical gate is:
| Control | Implementation |
|---|---|
| Hard row cap | Refuse to start if the manifest exceeds the approved count |
| Model tier | Draft with a cheaper/fast endpoint; render finals only after QA |
| Output cap | Keep num_images explicit rather than relying on defaults |
| Retry budget | Count retries separately from first attempts |
| Resume | Skip rows with verified local outputs |
| Cancellation | Use genmedia status ... --cancel for queued work when appropriate |
Record the pricing response with the manifest because model prices and endpoint availability can change. If the provider does not expose a comparable unit, label the result as a request-count ceiling, not an invoice.
fal genmedia CLI vs Replicate CLI vs a custom script
These tools solve different layers of the problem. Replicate’s official CLI exposes commands for running and streaming predictions, inspecting model schemas, uploads, training, and model management. genmedia is more useful when the job begins with discovering fal endpoints and moving media through fal’s queue and download lifecycle.
| Choose | Best fit | Main trade-off |
|---|---|---|
| fal genmedia CLI | fal-native model search, schema lookup, async jobs, downloads, and agent shell use | Provider-specific; your batch policy still lives outside the CLI |
| Replicate CLI | Replicate predictions, streaming, model/schema operations, and training commands | A different provider catalog and lifecycle; do not expect fal endpoint IDs or flags to transfer |
| Custom Python/HTTP script | Multi-provider routing, approval gates, database state, queues, and billing policy | You own authentication, schema changes, polling, downloads, and error handling |
My recommendation is straightforward: use genmedia directly for exploration and a small Python wrapper for a fal-only production batch. Move to a custom provider abstraction only when switching providers is a requirement, not because a wrapper feels more “enterprise.”
Treat outputs as records, then run QA
The public workflow skill recommends a compact manifest containing the goal, node ID, endpoint ID, request ID, input URLs, output URLs, downloaded files, and defect notes. That is more useful than a folder full of files with generated names.
Before accepting the batch, verify:
- Every
completerow has a local file and a JSON receipt. - No
failedrow was silently omitted. - Images have the expected dimensions and nonzero size.
- Videos open and have the expected duration, resolution, and frame rate;
ffprobeis suitable for this check. - Prompts and endpoint IDs are preserved for the assets you keep.
- Re-running the same manifest produces skips rather than duplicate files.
The fal guide emphasizes keeping generated media close to its JSON metadata. That practice also makes a later provider migration possible: you can compare output, request, and cost records instead of trying to reconstruct a run from filenames.
FAQ
Does fal genmedia CLI natively batch hundreds of prompts?
The documented commands provide model execution, async status handling, JSON output, uploads, and downloads. A manifest reader and concurrency controller are still needed for a resumable hundreds-of-prompts batch.
Is genmedia cheaper than Replicate CLI?
The CLI does not determine the provider’s model price. Compare the specific endpoint pricing, output settings, retry count, and transfer behavior for your workload; a tool-level “cheaper” verdict is not meaningful across different model catalogs.
For the shortest path, use genmedia for discovery and execution, add a manifest-driven wrapper for batching, and choose a custom script only when you need multi-provider routing or centralized job state.