WQ-Bench: prompt examples and visual evidence
What these examples prove
WQ-Bench sends varied human prompts through the real WRIGHT planner, QWAY furnishing and independent geometry verification. It records successful, unresolved and failed results; it is not a catalog of hand-built ideal scenes. RIDLEY is not scored by this spatial benchmark.
Current recorded outcomes
| Outcome | Cases |
|---|---|
| PASS | 0 |
| PASS_WITH_WARNINGS | 29 |
| UNRESOLVED | 27 |
| FAIL | 29 |
| NOT_RUN | 0 |
PASS_WITH_WARNINGS retains orientation/size uncertainty. Automatic success is separate from human screenshot approval. An empty portal list does not prove circulation. Missing asset categories, extra rooms and unsupported requested clearances must not be hidden by a successful compiler return.
Rooms, worlds and narrative prompts
| Family | Recorded outcomes |
|---|---|
| room | {"PASS_WITH_WARNINGS":24,"FAIL":12,"UNRESOLVED":12} |
| vague | {"UNRESOLVED":3} |
| narrative | {"PASS_WITH_WARNINGS":5,"UNRESOLVED":1} |
| constrained | {"FAIL":1,"UNRESOLVED":2} |
| conflicting | {"FAIL":3} |
| world | {"FAIL":8,"UNRESOLVED":4} |
| hidden | {"FAIL":5,"UNRESOLVED":5} |
The corpus includes explicit, conversational, narrative, screenplay-like, vague and contradictory prompts. Literature-inspired prompts are original short paraphrases, not quotations or canonical reconstructions. The held-out split is run only after generic changes freeze; it is an execution holdout, not a secret/blind dataset.
WorldPack and user asset replacement
Open the complete desk-custom replacement example. A room is first generated from standard assets. A different public CC0 desk is uploaded into the user library, registered with the alias desk-custom, and selected through the real prompt replace this desk with my desk-custom. The recorded proof compares IDs, measured geometry and independently satisfied relations before and after. This recorded-proposal follow-up is not counted as another live planner-generalization result.
All examples expose selected assets with scope, semantic category, license, source hash and pinned manifest/profile data. Standard, user and project assets use the same library; the benchmark does not create another catalog architecture.
Constraint-heavy and diagnostic examples
Impossible dining request retains the requested small room and explicit warning rather than claiming an unsupported clearance is satisfied. Laboratory demonstrates an absent asset category. Meeting visibility and access exposes unsupported guarantees. These are valuable regression examples, not successful scene recommendations.
Reproduce a result
Every gallery detail exposes its original prompt, QWAY source, typed WRIGHT topology, project envelope, diagnostics, selected assets and pinned replay dependencies. Renderable accepted results include a self-contained WorldPack for the standalone player. Load that package without an authoring API.
node wq-bench/run.mjs --split development
node wq-bench/run.mjs --split regression
node wq-bench/replay.mjs
node wq-bench/check.mjs --require-mappingPlanner mode uses the configured API-only provider and incurs model usage. Recorded-proposal mode is cheaper but is not new language understanding. Exact source replay uses no AI and checks scene identity, transforms, pins, relations and independent geometry against the recorded result.
Documentation fixture policy
Documentation images are selected from actual successful benchmark cases through wq-bench/documentation-map.json. The capture tool reloads those projects in isolated Studio processes, selects the relevant workspace/panel and records actual screenshots. It does not replace furniture, hide collisions or invent successful source for the camera. A failed mapped case blocks validation.
The low-level box/desk compiler examples elsewhere remain small API exercises, not representative Studio screenshots. The furnished-world gallery and mapped player show real public Kenney CC0 assets.
Limits and failure interpretation
85 of 85 corpus prompts have recorded planner outcomes in this build. The gallery keeps non-executed cases explicit when a run is incomplete. Scene appearance uses the currently available low-poly catalog; historical mood words cannot create absent models. Exact anatomical ergonomics, arbitrary visibility and requested clearance magnitudes may be unsupported. Internal repair counts and complete dropped-intent detection remain unknown where the runtime does not expose evidence.