Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Reports

Every panel eval dumps a self-contained HTML report (headline tables, per-repo breakdowns, worst-frame galleries) plus a JSON that the frozen results instruments consume. The HTML reports and the frozen analysis JSONs are hosted on the dedicated fontaine-reports Space (moved off this Space 2026-08-10, owner request — the ~10 MB self-contained report files were the bulk of this Space’s storage); this page indexes them. Posts link the specific reports behind their numbers.

Owner-side reports

  • AR-pretrained trunks for flow decoders (interim, 2026-08-05) — the paired two-arm stage-2 phase behind the −2.7 MAE AR-adaptation number: same expert/init/seed/data order, trunk stock vs AR-pretrained; Δ−2.69 (−20%) at step 2,500, ~8× the probe noise floor. Shared by the owner 2026-08-06; the direct motivation for the Molmo2 AR-first amendment.

Naming: eval__<run>__<checkpoint>__<panel+sampler>heun30 = Heun 30-step, 1nfe_euler1 = single Euler step (1 expert eval), drawsN = mean-of-N ensembling, stable/stablekey = stable noise keying (#18.2), unmarked = legacy index keying. panel_curated_v0_k4l2 is the v1 25,800-frame panel; panel_v2 is the dedup-hardened revision.

SnapFlow 1-NFE distillation (results)

Flow teacher bijou_flow_artrunk_h1024_40k_ddp2 @80k

Flow teacher @40k (arch-batch control)

AR mainline bijou_arb_rcond_100k_ddp4 @100k

Box batch 40k AR arms (results)

State-dropout arm C fontaine_arb_rcond_statedrop80_40k_1xh100 @40k (results)

Molmo2 AR trunk fontaine_molmo2_ar_40k_ddp4 @40k (results)

Molmo2 AR trunk fontaine_molmo2_ar_60k_ddp4 @60k (results · fields panel)

Molmo2-ER trunk fontaine_molmo2_er_60k_ddp4 @15k (mid-training, owner-requested)

MolmoAct2-SO100_101 out-of-band panel (pre-reg · plan/deep-read)

  • 3-policy side-by-side HTML report — owner spec 11:59Z 08-10: flow teacher 80k (top-10-tickets + stable-key + heun-30 original) vs the released MolmoAct2 SO-100/101 fine-tune vs state-copy, same 25,800 frames, matched 30-step/1.0 s window, 32-frame gallery with 4 policies overlaid per joint; rendered 14:25Z 08-10
  • frozen matched-window reads JSON (molmoact2_panel_reads.py) — matched-window chunk MAE core frames (willnorris/bbox-2 excluded, owner amendment 13:14Z): flow-teacher top-10-tickets 3.90 / state-copy 8.32 / MolmoAct2 13.87 pooled, 16.97 clean-633 vs 7.00 contaminated-245 (beats state-copy only on repos in its own fine-tune mixture, −0.75; trails the flow teacher by +3.29 [+3.11, +3.48] even there)
  • contamination repo list — 245/878 panel repos in their SO100_SO101_MOLMOACT2 mixture (7,996 frames, 5,332 core), derived live from their repo file
  • sweep metadata — their predict_action end-to-end, bf16, 10-step Euler, seed = concat index; 25,800 frames at 352 f/min, ~1.3 GPU-h total

er_60k ENDPOINT @60000 — THE ER decision read (pre-reg): ER init WINS both legs

  • endpoint eval report — chained in-unit panel (rc 13:28Z 08-11, ~153/155 GPU-h run total): fast path 5.7782/1.9898 core — the best banked trunk number to date; narrated arm 5.83 (+0.055 pairing, 45% win); aux holding 0.915 / progress MAE 0.060 / event 0.858 / visible 0.822
  • paired decision reads JSON (er15k_panel_reads.py, key bijou@60000) — vs 40k endpoint (6.0079) pooled −0.2297 [CI95 −0.281, −0.154] BELOW-BASELINE; vs 60k-cont (5.8602) pooled −0.0821 [CI95 −0.126, −0.025] BELOW-BASELINE, CI excludes zero = the pre-registered decision read: the ER-init trunk beats both banked baselines at matched panel class; state-copy integrity byte-match ×3. Rung trajectory 15k +1.52 → 35k +0.28 → 55k −0.18 → 60k −0.23 vs the 40k endpoint. Rig-data effect read at endpoint: NOT split-compatible — the panel contains no owner-rig repos (checked against the npz repo_id identity), recorded as skipped per the pre-reg’s if-clause
  • Weights: step_060000 (weights-only, hub-uploaded 12:44Z in 42.0 s, commit 4ed3dd0)

er_60k @60000 events one-off (owner request 12:44Z 08-11, record-only)

What events does the model actually see? Generated event strings vs the weak judge labels on the 8,987 judge-labeled panel frames, via the new --dump-generations instrument (commit 7f43c54 — main-arm generations retained under explicit --generate).

  • standalone report — 13-class model×gt confusion (incl. none/none), per-class P/R, and 136 image cards across hit / class-swap / miss / false-alarm galleries + probe examples (repo-diverse selection)
  • Headline: both-none 7,238 · hits 333 · swaps 129 · misses 683 · false alarms 604. On the 1,145 gt-event frames the model speaks on 40%, but class-agrees 72% when it does; exact-string match 3.6% (same event, different words)
  • constrained-probe JSON — on the misses, a 1-step none-ban re-decode (frame’s own generated prefix replayed; unbanned replay reproduced none bit-exact 679/683): forced guess lands the gt class 63% → the dominant miss mode is saw-it-under-threshold, not blindness (idle 86% / release-place 80% / occlusion 72% / blur 62%; camera quirks 10% and episode markers 0% are the genuinely-not-encoded tail)
  • confusion JSON · dump-pass eval json · per-frame generations dump (25,800 rows, ~1.55/4 GPU-h). Instrument oracle: presence acc 0.8568 vs banked 0.8582 — Δ 13 frames, inside the documented cross-world-size bf16 batch-composition band (banked ran 4-way on the box)

er_60k @55000 owner-requested panel, standard both-arms (pre-reg, record-only)

  • standard eval report — the @55000 read (rc=0 12:00Z 08-11, ~2.2/8 GPU-h): fast path 5.8269/2.0172 core + narrated arm 5.869 (+0.039 pairing, 46% win — same ~0.04–0.05 narration-cost class as 15k/35k); aux vs weak labels at full n≈8,987: holding acc 0.915→0.920, progress MAE 0.065→0.060, event acc 0.875→0.858, visible acc 0.823→0.822
  • class-matched paired reads JSON (er15k_panel_reads.py, key bijou@55000) — vs 40k endpoint (6.0079) pooled −0.1810 [CI95 −0.232, −0.105] BELOW-BASELINE (first such read for the ER trunk; @35000 was +0.281 above), vs 60k-cont (5.8602) −0.0334 [−0.078, +0.024] CI-SPANS-0 = parity at 92% training; state-copy integrity byte-match ×3; record-only — the @60000 endpoint panel decides
  • Weights: step_055000 (weights-only, hub-uploaded 09:4xZ in 42.9 s)

er_60k @35000 owner-requested panel, standard both-arms (pre-reg, record-only — SUPERSEDES the aux-arm read below)

  • standard eval report — the complete 35k read (rc=0 00:41Z 08-11, ~2.2/8 GPU-h): fast path 6.2892/2.3746 core + narrated arm 6.342 (+0.047 pairing, 44% win — narration costs the same ~0.05-class as at 15k); aux vs weak labels at full n≈8,987, ALL FOUR improved from 15k: holding acc 0.899→0.915, progress MAE 0.075→0.065, event acc 0.862→0.875, visible acc 0.704→0.823
  • class-matched paired reads JSON (er15k_panel_reads.py, key bijou@35000) — vs 40k endpoint (6.0079) pooled +0.2813 [CI95 +0.199, +0.337], vs 60k-cont (5.8602) +0.4290 [+0.353, +0.467]; ABOVE-BASELINE at 58% training, the 15k gap (+1.52) ~82% closed; state-copy integrity byte-match ×3

er_60k @35000 aux-narrated arm (superseded by the standard read above)

  • aux-narrated eval report — owner request 20:47Z 08-10: --generate subgoal holding progress event visible (actions follow the model’s own generated aux lines); core 6.3425/2.3770 at 58% training (er15k narrated-class was 7.601), win-rate 77% vs state-copy, Q3 condition sensitivity 1.62
  • paired reads JSON — vs 40k endpoint +0.335 [+0.247, +0.387], vs 60k-cont +0.482 [+0.399, +0.517]; cross-class caveat (narrated arm vs fast-path baselines) — the standard both-arms eval relaunched same-session supersedes these with class-matched reads when it lands (~01:0xZ 08-11)
  • Weights: step_035000 (weights-only, hub-uploaded 20:5xZ in 42.4 s)

MolmoAct2 SO-101 rig fine-tune (pre-reg · runbook · results)

  • anchor-rung HTML reportrig_ft_r1 (AE-only, 2000 steps, ~2.7/12 GPU-h): rung curve zero-shot 28.95 → 3.23@2000 vs state-copy 9.08 on the 240 rig anchor frames; per-timestep curves, motion-corr small multiples, 8-frame strided trajectory gallery. Pre-reg PASS at every gate; reads are train-frame sanity (contaminated by construction — the real eval is on-rig rollouts, runbook §3–4)
  • Frozen reads: zero-shot/preflight · step 500 · step 1000 · step 1500 · step 2000 (molmoact2_rig_preflight.py --model <rung>, identical 240 rows)
  • Weights on the hub: molmoact2_so101_rig_r1_step2000 — AE + resized-embedding delta vs the released checkpoint (trunk deduplicated, 704/707 tensors sha-verified byte-identical); serve-ready dir stays local at ~/checkpoints/molmoact2-so101-rig-r1-step2000-hf

Golden-ticket noise screen (close-out · visual report)

Frozen-trunk flow experts @10k, panel_v2 (attach memo · tiny results)

Both experts sit on the hard-frozen 60k trunk (sha e6ed783b), scored on the panel_v2 k4l2 plan (15,056 core frames pooled) — numbers compare within this section, not with the v1 scoreboard.

100-seed sim policy eval (pre-reg · results)

Five arms × seeds 0–99 in the v0 SO-101 sim (sysid’d servos), paired design; primary metric = boat→disk progress (cm). 0/500 successes; the engagement/direction split is the finding.

  • HTML report + video gallery — per-arm tables, paired CIs, four charts, best/median/worst clips per arm (+ the er60k reach-but-miss money shots)
  • frozen analysis JSON (sim100_reads.py: gates, summaries, paired bootstrap reads, ordering read auto-skipped — rung arms killed by the phase-2 amendment)

Contact-shadow pass v4 — the composite’s missing shadow, fitted and gated (lit page, 08-13)

The v3 composite’s pasted arm casts no shadow on the real plate — the one physics law every real frame obeys that no composite frame did. Leg (a) measured the real arm’s own shadow from 200 frames × 25 bank episodes (frame ÷ episode-plate darkening vs the sim-replayed silhouette slid along candidate light directions): real and directional — contrast +0.091 CI95 [0.081, 0.100] vs ring control, zenith 30° / azimuth 112.5° (85% bootstrap stability), strength 0.392, softness σ 24 px. render_style="v4" = v3 + the fitted shadow multiply-darkening the top plate (shared projector sim/shadow.py, 12 oracles; wrist bit-identical to v3). Paired encoder gate (seeds 0..99, fresh both arms — the banked v3 anchor 0.673 predates the bracket flip; fresh v3 reads 0.721): top 5-NN AUROC 0.721 → 0.715, and the paired per-seed read is decisive — Δknn5 −1.04e-07 CI95 [−1.53e-07, −5.6e-08], 66/100 seeds closer, ~10% of the remaining top-cam knn5 excess closed. Wrist 100/100 tied. GO recorded; default stays v3 pending the sim100 amendment-5 owner call. ~0.04 GPU-h. For scale: v1 scene −0.049, v2 inpainting −0.103, v3 content −0.100, v4 shadows −0.006 — the tail is thinning.

Fitted wrist lens — cubemap render path + gate: the fit’s center term double-counts the pose, the curve-only refit passes (lit page, 08-13)

The deployed wrist warp assumed an ideal equidistant lens centered at the image midpoint; the plumb-line fit on the 150 pinned real frames (leg (a), 08-13 01:4xZ) measured the real module off-center (22 px left / 14 px down, ~5σ) with stronger peripheral compression (−12.8 px at the corner, CI-excludes-0). Leg (b) landed the render path that can draw ANY lens: the wrist source is a pinhole cubemap around the camera axis (output→face map precomputed, so runtime is one bilinear gather; only referenced faces render; face focal matched to the deployed source so A/Bs read geometry, not sharpness; the camera-riding headlight is re-pointed at the base axis per face — without that, face boundaries carry a shading seam, caught by the rotated-cubemap oracle at mean|Δ| 6.77). Gate read (pre-reg 03:27Z, 20 seeds × 5 draws, er60k trunk, control 0.560): full fit 0.667 FAIL — and a labeled post-hoc center-only arm reads 0.672, reproducing the whole regression. The 08-12 wrist pose re-tune was fit to real frames under the deployed lens, so it already absorbed the principal-point offset (~2.6° yaw-equivalent); bolting the fitted center on top applies it twice. The curve-only refit (k2 +0.101, k4 −0.036) passes: 0.523 ≤ the 0.548 gate, paired Δknn5 −7.6e-07 CI95 [−8.5e-07, −6.8e-07], 96/100 frames closer — ~7× the contact-shadow GO effect, and cost-neutral (single face covers the frame: 73 vs 70 ms/tick). lens_model="fitted" now pins the curve-only params; default stays equidistant pending the sim100 amendment-6 owner call. Top cam bit-identical across all arms (0.713 — now the frontier number). Full-fit center use is parked behind a joint pose+lens refit (sim-joint-pose-lens-refit, owner-held). ~0.04 GPU-h total (4 probe arms).

Appearance screen CONSOLIDATED (report, 08-14)

The whole top-cam appearance programme in one chart-led post, written for the three pending promotion decisions: nine pre-registered reads, ~0.2 GPU-h total — clutter patches carry the removable share (0.713 → 0.556), materials are free riders, texture refuted twice, wrist-neutral, stack 0.5521 sub-additive; the remaining gap to the 0.328 pipeline floor is arm geometry/relief (renderer-class, priced separately). The individual entries below remain the frozen records.

Top-cam gap decomposition — the whole 0.713 residue lives in the rendered foreground pixels (08-13)

With the wrist at 0.523 under the curve-only fitted lens, the frontier is the top cam’s 5-NN AUROC 0.713 — unmoved by every lens arm. The decomposition probe asked WHERE it lives before spending on any lever: five ablation arms on the pinned 20×5 reset-render harness (numpy backend, fresh in-run v3 baseline reproduces the banked anchor exactly at 0.713). The verdict is unambiguous. Swapping the rendered foreground for real dynamic pixels — real arm/boat/hand pixels mined from bank-episode frames, re-lit source→target episode via the bank affines, pasted on a different drawn plate by the exact production arithmetic — collapses the read to 0.328, at the clean-repo anchor (0.283) and below the 0.5 null: the compositing pipeline (plates, mask feathering, grade, noise) can reach real-level; only the pixels it pastes are wrong. The arithmetic residue read agrees from the other side: fg→plate (0.869) vs plate-only (0.865) differ by +0.004 AUROC (paired Δknn5 +2.3e-07 CI95 [1.4e-07, 3.3e-07] — real but ~5% of the armless shift and far under the +0.05 registered bar), so edge-weighting + foreground-PSF carry ~nothing. Two labeled surprises: (1) the fg→plate/plate-only arms read FARTHER than v3 (0/100 frames closer, +4.5e-06) — an armless frame is itself OOD (every real frame has the arm somewhere), so registered rule 1’s first clause misfired for a measured, explainable reason; the real-fg arm carries the localization instead. (2) The shadow-band crop read is near-ceiling (v3 0.989, v4 0.988, crop clean anchor 0.304): within the table region the encoder separates sim from real almost perfectly, and the fitted shadow does NOT close it locally (paired crop Δ CI includes 0) — but the registered box grew to cover most of the lower frame (89:480, 81:640), i.e. it includes the rendered arm itself, so it localizes the signal to “the region containing the pasted render”, consistent with the real-fg verdict rather than a separate shadow story. v4’s full-frame paired read replicated the shadow gate on the 20×5 protocol (−8.3e-08 CI [−1.34e-07, −3.1e-08], 66/100 closer). Decision (registered rule): the next leg is foreground appearance — and the sample frames name the prime suspect: the untextured gray clutter stand-ins (cylinder mug, white disk) sit next to photoreal plates; queued as sim-foreground-appearance-pass with a content-split leg (clutter vs arm vs benchy, keeping the rest rendered to dodge the armless confound) before any material work. ~0.02 GPU-h embeds; renders CPU.

Foreground content split — the clutter stand-ins (~5% of pixels) carry the removable share (08-13)

Leg (a) of the appearance pass asked WHICH rendered class carries the 0.713: arm bodies (96 geoms, ~7.1% of pixels), benchy (341, ~0.1%), the clutter stand-ins mouse/mug/laptop/pcb (~5.1%), or the disk (~0.5%, split out of “clutter” as the always-rendered named suspect). One production v3 instance was hooked at _composite, so every slot yields all 10 arms — v3, plate-only, no_(class), only_(class) — through the exact production arithmetic with a segmentation-restricted mask: same physics, same drawn plate, same sensor noise (RNG-state restore), making the paired Δ exactly the class’s visible-pixel effect (in-run oracle: hooked v3 bit-exact == the production observation, all 100 slots; fresh v3 read 0.7127, inside the registered abort band). Removing the clutter stand-ins alone collapses the read 0.713 → 0.576 (paired Δknn5 −1.73e-06 CI95 [−1.92e-06, −1.54e-06], 99/100 frames closer) — the unique class past the registered ±0.05 material bar: no_disk −0.006 and no_benchy −0.002 are CI-excl-0 but immaterial, and no_arm reads +0.113 WORSE, the armless-content confound the decomposition labeled (every real frame has the arm). The keep-only duals all pull toward real when added to the bare plate (only_arm 0.654, only_clutter 0.824, only_benchy/only_disk 0.832 vs plate-only 0.866), so no class is rendered badly enough to overwhelm its own content benefit — the ranking rests on the removal direction, which is also the honest one for clutter (real episodes genuinely vary clutter presence; the bank plates are mined clutter-free). Registered primary rule fires: leg (b) target = clutter appearance (real-crop textures or plate-sourced patches for the gray untextured stand-ins). Ceiling note, registered before leg (b): no_clutter’s 0.576 still sits far above the real-fg anchor 0.328, and the arm carries most of that remainder (only_arm 0.654 vs the real-content direction ~0.33) — clutter alone cannot close the gap, it is just the best ROI per rendered pixel. Renders CPU (~5 min), embeds 12 groups ~0.02 GPU-h.

Foreground appearance fix — real-crop clutter patches beat the removal ceiling, gate PASS (08-13)

Legs (b)+(c) of the appearance pass (pre-reg in-channel 05:23Z) executed the registered follow-up: replace the untextured gray stand-ins with real-pixel crops pasted into the plate. make_clutter_crops.py mined per-object RGBA crops from the bank episodes’ naive per-pixel medians (source episode = largest measured blob; alpha = the feathered static-novelty mask vs the gain/bias-corrected global plate — the same statistic the bank pass localized the objects with; recomputed areas bit-match the manifest), normalized to global-plate lighting. clutter_patch.py pastes them at the drawn poses by inverse warp through the verified analytic fisheye model (target pixel → object-height plane → rigid drawn→mined transform → source pixel, bilinear), so translation, yaw jitter and the fisheye’s local scale all ride the camera model; the active episode’s affine grades the patch exactly like the rendered foreground; the fixed_canonical pcb pastes at its real measured location (identity). Zero extra appearance-RNG draws — slots pair 1:1 with production v3. The leg (a) harness then read three arms off one hooked instance: patched 0.556 vs v3 0.713 (ΔAUROC −0.157, paired Δknn5 −2.02e-06 CI95 [−2.21e-06, −1.83e-06], 100/100 slots closer) — the registered −0.05 gate passes at 3× the bar, and patched lands 0.020 BELOW the no_clutter removal ceiling 0.576 (75/100 closer, CI-excl-0): real-looking clutter beats clutter-free plates, as the real reference (clutter present in 15–77% of episodes) predicts. Integrity: in-run v3 0.7127 inside the abort band, no_clutter 0.5764 reproduces leg (a) within the registered ±0.01, hooked-v3 bit-exact all 100 slots, clean anchor 0.283 unchanged. Promotion of the patch paste into production v3/v4 is an owner call (asked in-channel 05:40Z); the remaining ceiling to real-fg 0.328 is the arm’s ~7% of pixels — a separate future item. Renders CPU (~4 min), embeds 5 groups ~0.02 GPU-h.

Wrist-view read of the arm material fixes — wrist-neutral: the two-flag stack moves ~230 raw px and the CI straddles zero (pre-reg, 08-14)

The wrist-side fact the two pending promotion asks (photometrics + mount) assumed rather than measured. Both flags are model-level material writes, so the wrist camera — inches from the recolored surfaces, its frame a RAW render (no composite) — sees them directly. Two paired production instances, 20 seeds × 5 draws, settled resets, er_60k knn5 probe, both cameras; gates all green (in-run TOP 0.713 dead-center; WRIST 0.561 in the registered [0.50, 0.60] reset band; qpos bit-equal ×100; changed-px tripwire quiet at 0.56% max). PRIMARY: paired wrist Δknn5 −1.39e-08, CI95 [−4.53, +1.73]e-08 straddles zero (46/100) — wrist-neutral; AUROC 0.561 → 0.560. The mechanism is visibility: at the home pose the wrist camera sees ~230 raw px of graded surface (servo 208 / PLA 21 / mount 1), so there is nearly nothing for the encoder to read — no regression (the texture failure mode did not fire), no gain. The top rider replicated the mount read’s combo delta bit-for-bit (−1.4937e-07, CI [−2.451, −0.570]e-07, 0.713 → 0.702) — production reset() observations and the _composite hook path produce identical frames: the hook was bit-exact. Registered limitation stands: the 0.828 ROLLOUT-pose wrist gap (gripper filling the frame mid-manipulation) is a different, still-open fact — needs banked trajectories or fresh rollouts, priced separately. Renders CPU (~9 min), embeds 8 groups ~0.02 GPU-h.

Arm micro-texture — a clean negative: statistically-matched grain reads MORE fake, both registered CIs above zero (pre-reg, 08-14)

The registered residual branch of the photometric close, executed and decisively refuted — the cheap kind of negative result. The graded arm is locally FLAT vs real (PLA print-layer local contrast 8.36 vs 4.66; servo glint tail p97 205.6 vs 125.2), so a composite-stage micro-texture (opt-in arm_texture="v1", deterministic static fields from a private pinned RNG, zero shared-stream draws, applied under seg masks before the production remap/blur/noise; 6 test oracles + init checks) was fitted THROUGH the composite to the mined real statistics: PLA local contrast landed 8.24 vs real 8.36, servo 10.46 vs 9.22, glint tail ~20% closed, photometric guard loss improved on both populations. The registered 20×5 read, all gates green (in-run v3_photo 0.698 dead-center, anchors exact): PRIMARY v3_tex vs v3_photo +9.33e-07 CI95 [+8.27, +10.42]e-07 entirely ABOVE zero, 3/100 closer, AUROC 0.698 → 0.751; MECHANISM only_links_tex +1.30e-06 CI95 [+1.22, +1.38]e-06, 0/100 closer, 0.652 → 0.740 — the texture undoes most of the grade’s gain. Reading: the pooled per-pixel statistics moved toward real while the encoder moved away — the probe sees spatial structure, not marginal statistics; screen-fixed band-limited grain reads as blotchy mottling (the zoom strip shows it), not as anisotropic, surface-tracking, shading-coupled print ridges. Composite-stage stats-matching is the wrong instrument class for texture; the branch dies in one session at ~0.02 GPU-h. Disposition per the frozen rule: no promotion ask; sim-arm-surface-texture-mjspec (true UV-mapped surface texture via the recompile path, physics-preservation oracles as its bar) queued as the escalation, not auto-run; the photometric grade (0.698/0.652) remains the arm-appearance frontier.

Arm SURFACE texture (mjSpec) — the SECOND refutation: true surface-tracking bands still read MORE fake (pre-reg + results, 08-14)

The micro-texture refutation’s registered escalation, executed and refuted in one session. arm_texture="v2" bakes a quasi-periodic layer-line texture INTO the 18 PLA link materials via an mjSpec recompile — bands live in OBJECT space and track the surface, the exact property the first refutation demanded. Physics hard bar 11/11 oracles green (every model field bit-equal, qpos bit-equal incl. a 60-tick excursion); zero-clip tanh generator with grade-preserving mean compensation; registered reflection rider (the texture legitimately shows in the tabletop’s 0.02-reflectance mirror of the arm — and is then fully absorbed by the PSF blur: composited max |Δ| 0). Fit honesty: period 32 frozen at the plausibility bound (lc response monotonic — fine bands die in the blur chain), amplitude CAPPED at the 0.42 no-clip headroom → realized PLA local contrast 6.43 of real 8.36 (grade-only 4.66): the albedo-modulation channel closes ~41% of the quadrature gap and cannot close the rest. The registered 20×5 read, all gates green (in-run v3_photo 0.698 dead-center): PRIMARY v3_surf vs v3_photo +3.07e-07 CI95 [+2.42, +3.71]e-07 entirely ABOVE zero, 14/100 closer, AUROC 0.698 → 0.718; MECHANISM only_links_surf +1.98e-07 CI95 [+1.36, +2.59]e-07, 27/100, 0.652 → 0.671 — about a third of the micro-texture’s harm, but confidently fake-ward. Coherence was NOT the missing ingredient. Diagnostics: the cube shrink-wrap renders sunburst fans on several faces (not clean layers), and the bands are pure albedo modulation while real print layers are RELIEF — shading/specular structure the classic renderer cannot express without a normal-map path. The arm-texture direction is COLD at this abstraction level; the graded arm (0.698/0.652) stays the production frontier; no further texture rung auto-queued.

Camera-mount material split — mechanism lands (93/100), whole-frame null: the part is fixed but too small to move the frame read (pre-reg, 08-14)

The arm-split’s per-pixel worst offender, measured and fixed — with a split verdict the pre-reg’s decision rule adjudicates cleanly. The mount (the wrist camera’s white 3D-printed bracket) shared a material with a black gripper piece; the fix first made the material mount-exclusive via a byte-identical detach (the gripper geom drops to matid=-1 with the color copied — the shipped material carries exactly mjv’s material-less defaults; oracle-pinned), then mined the real bracket at recorded poses. The white part can’t darkness-snap, so its mask rode the dark gripper/wrist per-body locks plus a brightness guard: 81/156 frames, 91k px — the real mount region reads neutral light gray [123, 120, 125], luma p50 121 vs the recolor-black composite’s 55. Fit through the production composite chose the same specular ceiling as both link populations (spec 1.0, shin 0.1; albedo 0.455/0.430/0.431), loss 177188 → 9028, composited medians dead-on real. The registered 20×5 read (in-run v3 0.713 dead-center; bridges reproduce the arm-split anchors exactly): MECHANISM PASSES decisively — only_mount_v1 −1.03e-06 CI95 [−1.16, −0.90]e-06, AUROC 0.821 → 0.793, 93/100 closer, and against the bare plate the graded mount reads −2.67e-06 with 100/100 closer — with the right color, mount presence now beats absence (the no_mount amputation confound, reversed). But PRIMARY FAILS — v3_mount vs v3 CI95 [−0.07, +1.42]e-07 includes zero, 45/100, AUROC 0.713 → 0.713: at ~0.66% of pixels the fixed part is below the whole-frame read’s detection floor. Per the frozen rule: no promotion ask for the mount flag alone. Record-only rider: the two-flag stack (mount + photometrics, what the pending promotion asks would flip together) reads 0.713 → 0.702, CI95 [−2.45, −0.57]e-07 entirely below zero (61/100) — the photometrics carries it; the mount flag rides at zero measured frame-level cost if the owner flips both. Amendment 1 logged pre-read: the locality oracle’s bit-equality was amended to a bound — the tabletop’s 0.02 reflectance mirrors any arm color change (≤24 px, ≤5 counts measured across all 200 oracle slots vs the 3000 px / 6 count bound). Renders CPU (3 sequential instances), embeds 8 groups ~0.02 GPU-h on the R1-A-freed GPU.

The execution of the arm-split verdict. Instead of guessing a better arm color, the real arm’s pixels were MEASURED: the sim posed at the recorded joints of 142 real v2 frames, its silhouette projected through the production fisheye onto them (per-body FFT darkness-snap ±60 px absorbs the tens-of-px registration offset; ring + absolute darkness guards, wrist excluded for its dark distractors), pooling 436k printed-PLA and 77k servo-casing pixels. The real black hardware is brighter than the flat recolor (median luma 66 vs 54), cool-cast [60, 66, 83], and 16–18% glints — sim rendered 5%/0%. The missing term was shine, not paint. Albedo solved per channel through the production composite, specular × shininess by grid: both populations chose the specular ceiling (spec 1.0, shin 0.1); fit loss ↓8.5× (PLA) / 2.3× (servo). Landed as opt-in arm_photometrics="v1" (default byte-identical, zero RNG draws, 5 oracles). The registered 20×5 read, all gates green (in-run v3 0.713 dead-center): PRIMARY v3_photo −2.22e-07 CI95 [−3.08e-07, −1.38e-07] entirely below 0, AUROC 0.713 → 0.698 (72/100 closer); MECHANISM only_links_photo −7.37e-07 CI95 [−8.35, −6.42]e-07, 0.705 → 0.652 (96/100) — the graded links alone now match the no_mount amputation best (0.654) without removing anything. Residuals for the registered texture follow-up: print-layer local contrast (real 8.4 vs graded 4.7) and the servo glint tail (p97 206 vs 125). Rider finding: the camera mount is WHITE in reality, black in sim, and its material is shared with the gripper wrist-roll piece — a mount fix needs a material split first. Production-default promotion pends the owner go. Renders CPU (~9 min, two sequential instances), embeds 7 groups ~0.02 GPU-h alongside R1-A.

The arm-class follow-on to the content split: the rendered arm (~7.1% of pixels) is the biggest remaining rendered class after the clutter patches (patched 0.556 ≫ real-fg 0.328), so WHICH arm sub-part carries it? Same hooked harness — one production v3 instance, _composite re-run per segmentation subset with RNG-state restore, 20 seeds × 5 draws — over two exact partitions of the 96 arm-class geoms: gripper+jaw (46 geoms, 0.3% px) / links base→wrist (44, 6.1%) / camera mounts (6, 0.7%), and follower (48, 3.6%) / leader (48, 3.5%). All gates green: in-run v3 0.713 (band ±0.005), bridges plate_only 0.866 / only_arm 0.654 / no_arm 0.825 all inside their registered ±0.02. The registered rule names LINKS: 88% of the whole arm’s keep-only paired delta (only_links −4.63e-06 of −5.26e-06, CI-excl-0; only_links alone reads 0.705 vs plate-only 0.866 — nearly the full v3 0.713). Gripper 26% and mount 31% sit below both the 60% and 35% thresholds. The instance axis is sub-additive — only_follower −4.05e-06 and only_leader −4.14e-06 each carry ~77–79% alone — so the encoder saturates on either instance and a fix must treat both arms. Record-only but striking: no_mount is the ONLY removal that moves v3 TOWARD real (0.713 → 0.654, 97/100 frames closer, CI-excl-0) despite the absence-OOD confound that makes no_arm read +0.113 WORSE — the six camera-mount geoms are per-pixel the most sim-distinctive thing in the frame, a cheap rider for the photometric rung. Follow-on queued: sim-arm-photometric-links (links material fix, both instances, mount-retexture rider). Renders CPU (~7 min), embeds 16 groups ~0.03 GPU-h.

20-seed behavioral spot-check under v3 (pre-reg, results, 08-12)

Same 20 seeds, physics bit-identical (spawn rows byte-matched in-run), only the rendering changed v0 -> v3. teacher80k improves +0.97 cm paired [CI95 +0.16, +1.81] — the only CI-excludes-zero read, direction flipped toward the disk; er60k (-0.07) and snap30k (+0.06) null. Visual familiarity moves the arm that engages. Amendment: GPU compositor (owner-approved) — 371 -> 94 ms/tick, probe reads preserved (0.669/0.113/0.544). ~1.3 GPU-h (gate 3).

Sim content diversity v3 — plate bank + clutter draws (pre-reg, results, 08-12)

Per-reset content variation for the v2 composite: a bank of 26 per-episode clean plates (each carrying its real episode’s lighting; ghost-free by inlier-median mining) + clutter presence/pose draws from the measured real between-episode spread. Registered bar (top k std/mean ≥ 0.15 AND 5-NN AUROC ≤ 0.790) MISSED on the spread leg (0.038 → 0.114) while the AUROC leg fell 0.773 → 0.673 (k-ratio 1.02× — top composites inside the real spread, best top-cam read yet). Wrist bit-identical to v2 (0.548). Default stays v2 per the registered flip rule; flip put to the owner. Record-only: the real disk wanders 8–29 cm × ±19 cm between episodes. ~0.08 GPU-h (gate 0.3).

Sim wrist-cam periphery re-tune (pre-reg, results, 08-12)

One runtime pose change in _repose_wrist_cam: camera moved from the wrist top behind the gripper to over the jaw base (≈10 cm forward, 55°→65° down) — under the 72° fisheye source the old pose filled the bottom ~40% of frame with gripper-body mass the real camera never sees. Registered bar (wrist 5-NN AUROC ≤ 0.786) smashed on the first candidate: 0.900 → 0.548, k-ratio 0.97× — sim wrist frames sit inside the real embedding spread. Guard green (top 0.773 bit-identical); 20×5 sensitivity 0.550. Per-episode wrist-plate axis retired. ~0.04 GPU-h (gate 0.2).

Sim visual matching v2 — real-frame inpainting (pre-reg, results, 08-12)

Real clean plates (per-pixel median over the 26 reference-half episodes; A/B pixel-disjointness verified in video-frame indices) composited under segmentation-masked rendered dynamic content (arms, benchy, disk + on-table clutter whose real twins move between episodes). Registered bar (top-cam 5-NN AUROC ≤ 0.790) MET: 0.890 (v0) → 0.876 (v1) → 0.773; overfit tripwire clear. Wrist composite regressed (0.951 vs 0.900 — the plate is cross-episode mush) so the shipped render_style="v2" (new default) keeps the v1 wrist path. Homogeneity unchanged (~4% vs 45% k std/mean) — content variation stays the diversity lever. ~0.06 GPU-h (gate 0.3).

Sim visual matching v1 — appearance pass + probe re-reads (pre-reg, results, 08-12)

Scene rebuild (real table texture, clutter layout, wrist-cam re-pose, fisheye remap, color grade, sensor emulation, per-reset appearance jitter) shipped as render_style="v1"; physics oracle-pinned bit-identical. Registered bar (top-cam 5-NN AUROC 0.890 → ≤0.790 on the reset-render probe) missed: best 0.874, final 0.876. Wrist responded to the camera re-pose (0.835 → 0.786 scene-only) then regressed under fisheye+grade (0.900). Sim stays ~10× too homogeneous at the encoder; lighting jitter moves per-seed distance only ~3%. Named next lever: real-frame inpainting (SIMPLER-RT style).

Encoder OOD probe — sim-vs-real at the policy’s eyes (rides the sim100 pre-reg, owner ask 01:11Z 08-12)

Sim frames (banked er60k-arm rollouts) vs real rig frames through the frozen er_60k vision trunk, per camera. Measured gap: top-cam 5-NN AUROC 0.885 / gap ratio 1.54× (wrist 0.828 / 1.33×); the clean-repo control lands inside the real spread (AUROC 0.26) so the shift is sim-specific. Sim is at the edge of the real manifold, not off it — the baseline the visual-matching lever must move.

  • frozen analysis JSON (sim_encoder_ood_probe.py: pinned frame selection, centroid-cosine primary + 5-NN secondary, AUROC/gap-ratio reads, per-frame distances)
  • distance strip chart — per camera × metric, three groups; sim’s tight blob at the real distribution’s right tail is the whole story in one look

Grasp-SFT route C fontaine_grasp_sft_joint_corrected @2000 (amendment · chain page)

  • flow-head unseen-100 eval report — owner request 08:25Z 08-16: 44/100 successes on unseen seeds 0–99 (euler-10) vs base 9 / corrupt-table stage-C 28 — A §5 verdict TABLE_FIX_POSITIVE (44 > 28+3, overlap band moot); anchor bar, per-seed spawn→final strip, 4-clip gallery, full table; rendered 08:5xZ 08-16 from the banked leg json
  • Remaining probe legs (flow-train memorization read, token-unseen vs R2 bar ≥20, token-base anchor) land ~12:3xZ 08-16; consolidated verdicts JSON analysis__grasp_sft_joint_probes.json + report to follow
  • standard 256-sample eval report (json) — owner request 09:06Z 08-16, stage-C train256 protocol reproduced (state-copy anchors bitwise 9.3562/9.8678): joint chunk MAE 3.24 vs corrupt-table stage-C 12.56 (which sat WORSE than state-copy 9.36 — the wrist_roll clamp); ~3.9× tighter fit on the same 256 demo frames with the corrected box
  • Weights: molmoact2_grasp_sft_joint_corrected_step2000 (weights-only, corrected table baked)

Grasp-SFT v1 grasp_sft_v1_joint_8xa100 @3000 (results · flow isolation · drift saga)

  • 3-leg sim100 chain panel — the chain that dated the collapse and separated the heads (14:17:56Z 08-17, ~6.2/12 GPU-h): step500 flow 4/100 / step500 token 16/100 / endpoint token under the serving fix b779ba4 14/100 vs probe flow 44 and endpoint flow 5 — token ~flat across training while flow never leaves the floor ⇒ the mis-fit normalization table poisons the flow targets, not the shared trunk; anchors bar, head-asymmetry slopegraph, per-seed strips, combined table, 9-clip gallery
  • frozen chain summary JSON (sft_v1_chain_report.py — headline numbers reproduce from the banked leg JSONs)
  • endpoint flow-head unseen-100 report5/100 vs probe 44 (per-seed data log-reconstructed after the box wipe; see the results page’s integrity note)
  • Weights: grasp_sft_v1_joint_step3000 (weights-only, byte-verified post-upload)

Grasp-SFT v2 drift discriminator grasp_sft_v2_demosonly_1gpu_disc @1000 (pre-reg · verdict)

The first non-drifting v2-corpus checkpoint (verdict HEALTHY 00:42Z 08-18 ⇒ distributed path convicted), panel-reported per the standing HTML-reports rule.

  • browsable HTML eval report (json) — current-stack eval on the probe-matched pins (demos holdout 0.1 / split-seed 0 / 256 samples seed 0 / chunk 30 / euler-10 / batch 12), 32 charted frames: chunk MAE 5.763 vs state-copy 7.671 (paired −1.95); reproduces the old-stack parity read 5.7626 to 3 decimals — the in-train probe’s 5.8989 is the known ×1.024 probe-vs-eval instrument shift from the verdict post. wrist_roll 12.31 stays the worst motor (3.5× state-copy’s 3.99), the residue the --per-dataset-flow-norm rerun targets
  • flow-head unseen-100 sim report (json) — the demosonly-v2 grasp cell of the isolation grid (04:19Z 08-18, ~2.2 GPU-h): 11/100 successes on unseen seeds 0–99 (euler-10, 30 s episodes) vs probe 44 / v1-endpoint 5 / base 9; mean progress 2.04 cm (probe: 3.86), 64/100 moved, 0 strikes, 7/11 success seeds shared with the probe. Healthy training + honest stats + demos-only corpus does NOT restore probe-level grasping — sits at the top edge of the broken class’s CI (~2–11), inside the pdnorm pre-reg’s own 11–19 ambiguous band (calibration note recorded in the draft pre-launch). Worn row: merged demos-native table via the default fallback (the leg json’s stats_repo_id field carries the rig lookup key — the pre-fix record semantics; fixed for future legs in bba4a45, which post-dates this leg’s launch)
  • k4l2 panel_v2 leg (json) — the baseline side of the pdnorm pre-reg’s paired panel read (04:57Z 08-18, ~0.5 GPU-h; euler-10 draws-1 stable, chunk 30, batch 32, protocol pinned in eval_disc1000_k4l2_panel.sh; npz pairing substrate local under reports/): chunk MAE 58.14 vs state-copy 8.37 — 7× worse than state-copy, 0% win rate on the community panel, from a checkpoint that beats state-copy on its own demos holdout (5.76 vs 7.67 above). The demosonly-v2 checkpoint is a narrow specialist; worst motors shoulder_lift 104 / elbow_flex 99 / wrist_roll 71. Interpretation was hedged pending the instrument audit (which row panel items wore: genuine forgetting vs the demos-recomputed table’s windows at serving) — resolved by the wear audit below: about half window, half collapse. Calibration fact recorded pre-launch: at baseline 58.14 the pre-reg’s +0.05 panel guard is near-vacuous as framed. Relaunch note: attempt 1 (batch 12/workers 8) was input-starved (66 f/min, projected 5.7 GPU-h) and killed at 4.7 min per the first-poll rule; r2 (batch 32/workers 20) ran at 96% util.
  • panel-row wear audit — what the 58.14 is made of (disc1000_row_audit.py, oracle tests/test_disc1000_row_audit.py; anchors reproduced to 1e-3). Wear fact: the checkpoint records the MERGED scheme (normalization: "q01q99", per_dataset_flow_norm: false), so the eval never consults per-dataset rows at all — every panel item wore the recomputed-at-launch demos-only global table; “community repos missing from the table” never arises, there is no lookup to miss. Decomposition: 85.8% of core truth elements have ≥1 joint outside the worn box, but the box FLOOR is only 14.40 of the 58.14, and predictions are NOT edge-saturated (≤0.3% on arm joints) — the wear hurts through the affine re-expression, not the clamp. Re-wearing the exact same normalized predictions through honest per-repo rows (fit on the panel’s own truth, 838 repos) halves the row to 27.40; through the released source table, 54.40 (also the wrong window). But the re-worn model is WORSE than a constant repo-box-midpoint null (25.15) — output-wear-corrected, the checkpoint carries no usable signal on community data; its raw predictions sit 22.6 from the constant demos action mean while truth sits 58.2 away (collapse to the demos prior). Verdict: ~half the 58.14 is serving-window re-expression, the residual ~half is genuine model failure on community inputs (state-side crush + forgetting, not separable post-hoc — the state input stayed binned through the demos table; a --molmo-norm-style re-run would separate them if it ever matters). Reading the pdnorm endpoint against this baseline: wear-corrected reference points are 27.40 (re-worn disc-1000) and 25.15 (midpoint null); the real bar stays state-copy’s 8.37
  • released-checkpoint k4l2 panel row (json) — the pre-SFT released checkpoint (molmoact2-so101-released, the SFT lineage’s init) through the pinned panel protocol (08:22Z 08-18, ~0.45 GPU-h, record-only PRE-GO; euler-10 draws-1 stable, chunk 30, batch 32/workers 20 at ~1173 f/min 95–100% util; wearing its OWN released source table — q01q99 global stats, no per-dataset rows, the default path is its honest wear): chunk MAE 25.89 on the 15,056 core frames, 9% win rate vs state-copy (8.3678 reproduced ≈ banked 8.37, anchor green); first_mae already 21.99 (global misprediction, not chunk-horizon drift); worst motors shoulder_lift 68.9 / elbow_flex 43.1 — the same two that dominate the SFT row’s 104/99. Read (frozen in the queue item pre-launch): 25.89 is AT the 25.15 midpoint null → community competence was never in reach for this lineage; SFT had ~nothing real to destroy on this panel, so the wear audit’s “genuine collapse” half reads as collapse-to-demos-prior of an already-at-null model, not forgetting of once-held competence. Wear-mismatch caveat DISSOLVED 09:xxZ 08-18 by the honest-wear re-expression below: same-wear released row 27.14 vs SFT 27.40
  • released-row honest-wear re-expression — the released row above re-worn through the SAME honest per-repo rows the disc-1000 27.40 reference wears (released_row_rewear.py, sibling of the wear audit; oracle tests/test_released_row_rewear.py; CPU, from the banked npz — no model re-run, output-side only). Anchors green (25.8924 / 8.3678 reproduced), inversion round-trip worst 1.5e-05 deg, midpoint-null identity anchor confirms the honest rows are byte-identical to the SFT audit’s (panels element-identical, null 25.154476 on both sides). Same-wear read: released 27.14 vs SFT 27.40 (Δ +0.26) — wear held fixed, SFT ended within noise of where it started, and both rows are slightly WORSE than the 25.15 repo-midpoint null; per-joint the re-wear trades shoulder_pan / wrist errors up for shoulder_lift 68.9→66.1 / elbow_flex 43.1→36.2 down, same two dominant motors. The anchor-ladder released rung is now the same-wear 27.14 (own-table 25.89 kept in the note); released-vs-endpoint stays the informative comparison on GO, now wear-consistent end to end
  • paired per-seed read: probe vs disc-1000 — retro shakedown of the frozen paired-read instrument (sim100_paired_read.py, oracle tests/test_sim100_paired_read.py; the pdnorm endpoint’s registered non-gating read vs this baseline, frozen pre-data): probe 44 vs disc-1000 11 = +33 successes, bootstrap CI95 [22, 44]; discordant seeds 37 probe-only vs 4 disc-only (McNemar exact p ≈ 1.0e-7); paired progress delta +3.57 cm [2.66, 4.46], 80% per-seed win rate. The instrument cleanly separates the healthy/broken classes on banked data — at the pdnorm endpoint it reads against disc-1000’s 100 episodes with these same constants (seed-0 bootstrap, 10k resamples). Rendered: the flow-unseen report above now carries this read as a “Paired read” section (delta tiles + McNemar discordant-seed chart, with the 11–19 ambiguous-band note) via grasp_sft_joint_unseen_report.py --paired-json (oracle tests/test_grasp_sft_joint_unseen_report.py, 4cfefae) — the pdnorm endpoint report gets the same section from its own frozen paired json
  • Weights: grasp_sft_v2_demosonly_1gpu_disc (steps 500 + 1000, weights-only, banked 00:5xZ 08-18)

Cross-family analyses

New reports land here as their evals finish; if a number in a post has no link yet, its report predates this page — ask and it gets pushed.