Loop Engineer

OpenAI Build Week · Education track · decision book

Loop EngineerFive original 3D games about making change trustworthy.

One shared lesson. Five completely different fantasies. Each can become a complete desktop, mobile, and WebXR judged demo in four focused implementation hours.

Competition clock: submissions are due Tuesday, July 21, 2026 at 5:00 p.m. Pacific. The recommended bet is Kaiju QA; the recommended backup is The Museum of Almosts.

A friendly baby kaiju in a safety helmet walking through a miniature city test course on a laboratory table.
Recommended: Kaiju QAStage tests · catch regression · ship safely

The lesson is the loop.

Loop Engineering is an emerging practice: define the goal, take a bounded action, observe external evidence, adjust from causality, then pass, stop, roll back, or escalate. Those loops live inside every SDLC phase.

Diagram showing Goal, Act, Observe, Adjust and a pass-stop-escalate gate nested inside seven software development lifecycle phases.
Read the learning-loop diagram as text
  1. Goal: define outcome, proof, constraints, budget, authority, and exit.
  2. Act: make the smallest safe, reversible change.
  3. Observe: inspect tests, traces, users, telemetry, or review.
  4. Adjust: change the plan or implementation from that evidence.
  5. Gate: pass, repeat, roll back, stop, or escalate.

The loop can run inside conceive/plan, specify, design, build/integrate, verify/validate, release/operate, and maintain/retire.

Timeline showing a three-minute lesson: see the failure, earn a first success, read evidence, catch regression, then ship or escalate.
Read the three-minute arc as text
  1. 0:00–0:20: see a deterministic failure and a visible goal.
  2. 0:20–0:45: make one scoped change and earn a first result.
  3. 0:45–1:30: read a partial pass and the actual mismatch.
  4. 1:30–2:20: catch regression and adjust without weakening criteria.
  5. 2:20–3:00: pass every gate, release or escalate, and invite replay.

No concept shortlisted yet.
Your choice is stored only in this browser when local storage is available.

Inside a sky lighthouse, glass prisms on a brass table open a glowing path through a storm for a small airship.
Cinematic rescue · paper score 87

Stormglass

Rotate the prisms of a sky-lighthouse, run repeatable storm tests, and use lightning traces to open one verified route for a courier airship.

Why it is distinct: evidence fills the sky. The player calibrates one optical variable, compares route traces, and decides between a fast fragile path and a longer robust path.

Core verbAlign prisms
EvidenceRoute traces
EmotionUrgency → relief

The three-minute loop

Baseline: pull the lever; crosswind collapses the courier route.

First success: rotate one prism; the nearest cloud ring clears.

New evidence: the fast route overloads the lighthouse.

Adjustment: compare ghost paths and preserve both acceptance criteria.

Release: pass a repeat test and commit the safe corridor.

A small robot tends a circular terrarium with flowers, roots, pollinator, and translucent earlier plant growth around it.
Calm systems mastery · paper score 84

The Tomorrow Garden

Run accelerated seasons in a seedship terrarium, compare the evidence left by each cycle, and cultivate a habitat that can sustain itself after release.

Why it is distinct: the system is grown rather than assembled. Time, dependencies, and stability become visible plant silhouettes instead of charts.

Core verbCultivate
EvidenceSeason overlays
EmotionCare → wonder

The three-minute loop

Baseline: a sprout grows, but the habitat cannot reproduce.

First success: add one sun tile and rerun the season.

Regression: extra water helps one plant and floods another.

Adjustment: compare prior gardens, prune one branch, rebalance one input.

Release: a full unattended season passes and the ark pod seals.

A baby kaiju in a safety helmet walks through a toy city testing course surrounded by robotic laboratory arms and path traces.
Recommended · paper score 92

Kaiju QA

Before a helpful baby kaiju meets a real city, build a tabletop test district, stage edge cases, and prove that its behavior is safe enough to ship.

Why it is distinct: the player is not the monster or its programmer. They are the quality engineer. Test coverage becomes a miniature city and over-broad fixes create funny, visible regression.

Core verbStage tests
EvidenceCoverage paths
EmotionComedy → confidence

The three-minute loop

Happy path: the kaiju carries a stalled car and earns a partial pass.

Edge case: place a fragile tower; helpful behavior knocks it over.

Regression: a broad freeze rule protects the tower but blocks an ambulance.

Adjustment: compare both paths and place a targeted slow-zone boundary.

Release: every scenario passes; the tiny kaiju becomes the city guardian.

Golden trains and translucent rehearsal trains loop above a dark miniature city toward a large sunrise.
Release ritual · paper score 86

Sunrise Express

Rehearse a magical release train across an orbital switchyard, diagnose where dependencies stop, and dispatch a verified sunrise to the city below.

Why it is distinct: the train is the release artifact, not a logistics economy. The player sees integration order, quality gates, rollback, and operation as one civic ritual.

Core verbRoute rehearsal
EvidenceGhost trains
EmotionMomentum → triumph

The three-minute loop

Baseline: one light car arrives, but the full train stops at a dependency.

First success: flip one switch; two cars pass the next gate.

Quality tradeoff: a shortcut is fast but leaves one car unverified.

Adjustment: use the suspended route trace to choose a robust path.

Release: dispatch the train; its wake lights the city and monitor.

An origami bird sculpture in a museum is compared with translucent reference versions through a large inspection lens.
Backup choice · paper score 89

The Museum of Almosts

Restore an unfinished kinetic exhibit by comparing layered prototypes, finding their shared mismatch, and preserving every useful failure in the final verified story.

Why it is distinct: the evidence itself is the toy. An inspection lens, translucent attempts, and one test crank turn diffs, regression protection, and retrospectives into spatial play.

Core verbInspect and compare
EvidencePrototype layers
EmotionMystery → wonder

The three-minute loop

Baseline: turn the crank; the paper bird forms only one wing.

First success: align an expected silhouette and replace one worn part.

New evidence: the wing passes, but flight still fails.

Adjustment: overlay attempts and fix the cause without erasing prior gains.

Release: the bird flies and the exhibit preserves the complete attempt history.

Choose the bet.

Paper scores are a decision aid. The final lock should follow a 30-minute graybox of the riskiest interaction and a thumbnail test of the hero state.

ConceptVerbBest strengthMain riskScoreRecommendation
StormglassAlignCinematic transformationWeather VFX scope87Visual alternative
Tomorrow GardenCultivateGentle systems lessonCan resemble resource balancing84Calm alternative
Kaiju QAStage testsInstant comprehension + regressionAvoid physics92
Sunrise ExpressRouteRelease and monitoring metaphorFamiliar network puzzle86Deployment alternative
Museum of AlmostsInspectOriginality + emotional payoffAbstract objective89Backup
How the paper scores were calculated

Each criterion is scored 1–5. Weighted contribution is (score ÷ 5) × weight; no hidden penalty is applied.

ConceptComprehension /15Verb /15Arc /15Build /20Parity /10XR /10Demo /10Energy /5Total
Stormglass4454455487
Tomorrow Garden4454544384
Kaiju QA5455445492
Sunrise Express4545444386
Museum of Almosts4454555489
Four-hour build schedule prioritizing the complete loop, cross-platform input, feedback, polish, and final validation.
Read the four-hour schedule as text
  1. 0:00–0:30: lock the complete arc, states, failure, and win.
  2. 0:30–1:15: build one semantic verb for pointer, touch, and XR ray.
  3. 1:15–2:10: implement baseline, evidence, adjustment, and rerun.
  4. 2:10–2:50: add shape, motion, and plain-language teaching cues.
  5. 2:50–3:20: polish one hero transformation.
  6. 3:20–4:00: typecheck, build, E2E, XR checks, screenshots, deploy, and README.

Devpost preparation

Design the submission while designing the game.

Recommended track: Education. Position the selected runnable game—not this concept book—as a three-minute practice tool for novice AI-assisted developers.

Suggested tagline: Turn the AI build-test-learn loop into muscle memory.

Working project: one complete desktop/mobile browser path, with WebXR enhancement.
Public demo video: shorter than three minutes, with audio covering the game and Codex/GPT-5.6 use.
Repository: public with relevant licensing, or private and shared with both official judge addresses.
README: setup, testing, asset licenses, old/new work boundary, and concrete human–Codex collaboration.
Codex evidence: include the main feedback session ID where most core functionality was built.
Judge access: keep the project free and available through August 5, 2026 at 5:00 p.m. Pacific.
Four equal criteria: technological implementation, design, potential impact, and quality of the idea.