Evidence you can inspect.
How Super Arbor evaluated AI renovation imagery—and what this first benchmark can and cannot tell you.
Test date: September 30, 2026 · Published October 1, 2026 · Publisher: Super Arbor
The test
We attempted 17 fal image-editing endpoints on three supplied photographs: a glass conservatory restaurant, a residential sunroom, and a house exterior. The candidate list came from a user-supplied general image-model leaderboard screenshot. That screenshot’s rankings and prices were not used as interior scores or measured costs.
Each endpoint received the same source photograph and frozen prompt for each scene. We kept the first result, with no masks, model-specific prompt rewriting, or paid retries. Seeds were unset and requests ran with concurrency three. Of 51 requests, 45 returned images. Both MAI endpoints failed all three scenes.
How the scores work
One AI assistant visually evaluated four components on a 0–10 scale: structure, camera and layout preservation (35%); design coherence and material quality (25%); instruction adherence (25%); and photographic realism (15%). The weighted result is expressed out of 100. The interior score is the equal mean of the restaurant and sunroom results. The exterior case is reported separately.
Most review sheets used anonymous letters, but the assistant had access to model metadata and initially saw a named output. This was not an independently blinded review. No people voted, no Elo system was used, and no confidence intervals were calculated. Displayed decimal places record the arithmetic, not measurement certainty. Differences of a few points should be treated as practical ties.
Cost and speed
Total measured usage was $6.3678 USD. Each request’s reported billable units were multiplied by its endpoint’s authenticated unit price saved during the run. The six failed MAI requests reported zero billable units. This is usage accounting, not invoice reconciliation; taxes, storage, and application overhead are excluded.
The leaderboard reports average cost and median latency across all three successful requests for a model, including the exterior. Time runs from submission to the first observed completion, includes provider queue time and two-second polling, and excludes shared upload and output-download time. These observations are not service-level guarantees.
Settings and output sizes
We targeted approximately 2K outputs where available: GPT 2.5 used max quality; GPT 2 and GPT 1.5 used high; Ideogram used high; Grok 2 used medium, the highest setting supported by the inspected schema. Exact settings appear on every model page and in the public JSON.
Actual sizes varied: GPT 2.5, GPT 2, Qwen 3, Seedream 5 Pro, and Ideogram returned 2048 × 1536; Nano Banana variants 2400 × 1792; Grok variants 2368 × 1776; Muse 1792 × 1344; Hunyuan 1024 × 768; Seedream Lite 2304 × 1728; and GPT 1.5 1536 × 1024. This is a practical endpoint comparison, not a strictly equal-resolution experiment. Images shown in the gallery use optimized previews; full-size links preserve the original outputs.
What this does not establish
- Two interiors and one exterior do not represent every room, architectural style, or renovation task.
- One output per scene does not measure repeatability or production reliability.
- Conceptual imagery does not prove exact SKU reproduction, color accuracy, dimensions, quantities, or construction feasibility.
- The exterior prompt ambiguously called the small door plaque a third glazed inset. Exact door-pane count was excluded as a scoring discriminator.
- Prices and endpoint availability are dated observations. Recheck before future use.
Coverage and failed runs
MAI Image 2.5 and MAI Image 2.5 Pro returned HTTP 422 with a generic processing error for all scenes. Their cause was not established. They remain unranked rather than receiving zero quality scores.
Seven exact candidate variants were not found in the connected fal catalog at test time:
- MAI-Image-2.6
- MAI-Image-2.6-Flash
- MAI-Image-2.5-Flash
- HunyuanImage 3.5 Preview
- Qwen-Image-2.1
- Luma UNI1 Max
- HiDream-O1-Edit-15
This does not mean those models are unavailable from every provider. Qwen’s official API description identified the tested endpoint as Qwen Image 3 Pro. Hunyuan 3.0 Instruct had a testable endpoint even though the supplied screenshot said no API.
Inspect, download, or cite
Every ranked model page includes its three results, per-image scores, reviewer observations, costs, dimensions, endpoint, and request settings. Public downloads contain the benchmark measurements and prompts; credentials and private execution metadata are excluded.
Suggested citation
Super Arbor. AI Interior Design Benchmark, September 30, 2026 pilot. Published October 1, 2026. 17 endpoints, two interiors and one exterior, single-assistant visual assessment. https://leaderboard.superarbor.com/
Research ownership and corrections
Super Arbor publishes this benchmark alongside its Design Studio. The scores are an assistant’s assessments of the archived outputs, not provider endorsements or independent laboratory certification. To report an error or suggest a model, contact ProMax@Superarbor.com.