Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The appearance screen, consolidated — what we measured, what to flip, and what’s left

2026-08-14 · consolidated report, banked numbers only (no new measurements). Written for the three pending promotion decisions. Queue item sim-appearance-consolidated-report; closed-screen rule.

Plain words. Our simulator draws pictures of the robot’s table, and a neural network that has only ever seen real photos can tell the drawings apart from the real thing. Over three days we measured which pixels give the game away, fixed the ones that could be fixed cheaply, and proved two seductive fixes don’t work. The verdict: the fake-looking clutter objects (a gray mug, a gray mouse) were most of the removable problem — pasting real photo crops of them into the scene closes most of the measurable gap. A careful re-paint of the robot arm helps a little. Adding fake surface texture to the arm makes things worse, twice, in two different ways. And everything that remains lives in the rendered arm itself — its shape and how light plays on it — which no re-paint can reach. The owner has three switches to flip; this report says which ones are worth it, in what order, and what each is worth.

Left: every measured arm of the screen on the one shared instrument. Teal = shipped opt-in fixes, amber = refuted texture arms, gray = anchors and diagnostic references. Right: the paired reads the promotion decisions actually rest on.

The instrument (one paragraph)

Every number in this report is the same read: 100 paired slots (20 seeds × 5 draws) of settled reset frames rendered through the production compositor, embedded by the frozen er_60k vision trunk, scored knn5 AUROC vs held-out real rig frames — lower = sim reads more real; 0.5 = indistinguishable; the clean real anchor reads 0.283. Arms are generated by re-running _composite on one hooked production instance with RNG-state restore, so a paired Δknn5 is exactly the visible-pixel effect of the change — same physics, same plate, same sensor noise. CIs are 10k-resample bootstrap over the 100 paired slots. Every leg below was pre-registered in-channel before its read; every read ran all its gates green (in-run v3 anchor dead-center 0.7127, qpos bit-equal, tripwires quiet).

How the screen unfolded

1. Where does the 0.713 live? The top-cam gap decomposition swapped frame regions between sim and real. Swapping the rendered foreground for real dynamic pixels collapses the read to 0.328 — at the clean-repo anchor. The compositing pipeline (plates, feathering, grade, noise) can already reach real-level; only the pixels it pastes are wrong. That number is the screen’s floor: everything the appearance programme could ever recover.

2. Which foreground class? The content split removed each rendered class in turn. Removing the gray clutter stand-ins alone collapses 0.713 → 0.576 (99/100 slots closer) — the unique class past the registered bar, on ~5% of pixels. The arm reads +0.113 worse when removed (every real frame has an arm; absence is itself OOD), so the arm must be fixed in place, not removed.

3. Fix the clutter. The real-crop patch pass mined per-object RGBA crops from the bank episodes and pasted them at the drawn poses through the verified fisheye model, graded by the episode affine. 0.713 → 0.556, 100/100 slots closer, 3× the registered gate — and 0.020 below the removal ceiling: real-looking clutter beats clutter-free. (before/after strip)

4. Which arm pixels? The arm sub-part split: the links carry 88% of the arm signature; the six camera-mount geoms are the per-pixel worst offender; either arm instance alone saturates the encoder, so a fix must treat both.

5. Re-paint the links from measurement. Arm photometrics posed the sim at the recorded joints of 142 real frames and measured the real arm’s pixels under the production fisheye: the real hardware is brighter, cool-cast, and 16–18% glints — the missing term was shine, not paint. Fitted through the composite: 0.713 → 0.698 (CI below zero, 72/100), mechanism arm 0.705 → 0.652. (strip)

6. Fix the white bracket. The mount material split: mechanism decisive (only-mount 0.821 → 0.793, 93/100; presence now beats absence), whole-frame null at ~0.66% of pixels. No solo promotion ask per the frozen rule — but the two-flag material stack reads 0.713 → 0.702, CI entirely below zero: the mount flag rides free if photometrics flips.

7. Texture: refuted, then refuted again. The photometric close left the graded arm locally flat vs real (print-layer contrast 4.7 vs 8.4). Two escalating attempts to add it back both read more fake: composite-stage statistically-matched grain (0.698 → 0.751 — the encoder reads spatial structure, not pooled statistics), then true surface-tracking bands baked into the materials via mjSpec (0.698 → 0.718 — coherence was not the missing ingredient; real print layers are relief, shading structure the classic renderer cannot express). The albedo channel is exhausted; the direction is cold. (the mottling, zoomed)

8. The wrist side is safe. The wrist-view read: the material flags change arm pixels the wrist camera stares at from inches away, but at reset poses it sees ~230 px of graded surface — paired Δ straddles zero (46/100). No regression, no gain; the promotion asks’ wrist-side sanity is measured, not assumed. (The 0.828 rollout-pose wrist gap is a different, still open fact — priced separately.)

9. Do the fixes compose? The full opt-in stack read (all three flags together): 0.5521 — far better than v3 (−2.075e-06, 99/100) but only −0.0040 under clutter-alone, inside the registered ε. The materials’ marginal on top of clutter straddles zero (−5.50e-08, CI [−1.44e-07, +3.37e-08]); the interaction term is +0.0063 (sub-additive). Roughly two-thirds of the materials’ banked solo contribution fails to survive composition with the clutter patches.

The three promotion decisions

All three asks are still open in-channel. The measured facts, in decision order:

flagaskedworth aloneworth stackedcall this report supports
clutter_patch paste → default05:40Z 08-13−0.157 AUROC (0.713→0.556), 100/100carries the whole stackflip first, or alone — this is the payload
arm_photometrics="v1" → default02:1xZ 08-14−0.015 (0.713→0.698), CI-excl-0absorbed next to clutter (CI straddles 0 at n=100)safe to stack, zero measured cost — but don’t price it as additive
mount material fix → default(rider on photometrics)frame-null solorides the material stack at zero measured costflip with photometrics or not at all

The ordering matters for honesty, not safety: stacking everything is measured safe (no regression anywhere, wrist included), but the value claim belongs to clutter. If the material flags’ stacked worth needs pinning before a flip, a bigger-n read resolves the −0.55e-07 marginal — priced on request, not queued.

What remains, and what it costs

The stack lands at 0.552; the pipeline floor is 0.328. The remaining ~0.22 AUROC lives in the rendered arm itself — and the texture refutations say it is not albedo. The surviving hypothesis is geometry/relief and light transport: print-layer relief, specular structure, soft self-shadowing — things the classic fixed-function renderer cannot express without a normal-map/PBR path. That is a renderer-upgrade decision, priced separately if the sim-to-real gap ever justifies it; the screen’s measured advice is that nothing cheaper than that is left on the table.

Whole-screen ledger: nine pre-registered reads over three days, every render CPU, ~0.2 GPU-h total in embed passes — all of it alongside live GRPO training runs on the same GPU. Two clean negatives banked (texture ×2), one confound identified and dodged (armless-OOD), one assumed fact converted to a measured one (wrist-neutral), and a promotion case reduced from “three flags, unknown interactions” to “one payload plus two free riders.”

Artifacts

  • Lead chart: ladder + paired reads (appearance_report_chart.py, banked JSONs only)
  • Every underlying read: frozen analysis JSONs + charts + frame strips linked from the reports ledger entries cited above; every pre-reg on the posts index.