Reports
Every panel eval dumps a self-contained HTML report (headline tables, per-repo breakdowns, worst-frame galleries) plus a JSON that the frozen results instruments consume. The HTML reports and the frozen analysis JSONs are hosted on the dedicated fontaine-reports Space (moved off this Space 2026-08-10, owner request — the ~10 MB self-contained report files were the bulk of this Space’s storage); this page indexes them. Posts link the specific reports behind their numbers.
Owner-side reports
- AR-pretrained trunks for flow decoders (interim, 2026-08-05) — the paired two-arm stage-2 phase behind the −2.7 MAE AR-adaptation number: same expert/init/seed/data order, trunk stock vs AR-pretrained; Δ−2.69 (−20%) at step 2,500, ~8× the probe noise floor. Shared by the owner 2026-08-06; the direct motivation for the Molmo2 AR-first amendment.
Naming: eval__<run>__<checkpoint>__<panel+sampler> — heun30 =
Heun 30-step, 1nfe_euler1 = single Euler step (1 expert eval),
drawsN = mean-of-N ensembling, stable/stablekey = stable noise
keying (#18.2), unmarked =
legacy index keying. panel_curated_v0_k4l2 is the v1 25,800-frame
panel; panel_v2 is the dedup-hardened
revision.
SnapFlow 1-NFE distillation (results)
- student 1-NFE, single draw (primary)
- student 1-NFE, mean-of-5
- student 1-NFE, mean-of-10 (deployment headline)
- frozen analysis JSON
(
snapflow_results.py, pre-registered reads)
Flow teacher bijou_flow_artrunk_h1024_40k_ddp2 @80k
- Heun-30, single draw (v1 anchor)
- Heun-30, mean-of-5
- Heun-30, mean-of-10
- Heun-10, single draw
- Heun-10, mean-of-10
- Heun-30, stable keying (re-banked anchor) (results)
- legacy k4l2 panel, Heun-30
- state-masked Q4 probe (results)
- Draws-fairness frozen reads: analysis · validate (results)
- σ_draw: finalization · direct measurement (amendment)
Flow teacher @40k (arch-batch control)
- panel-v2, Heun-30, stable keying — the arch batch #1 control · ctrl-only analysis
AR mainline bijou_arb_rcond_100k_ddp4 @100k
- curated_v0 panel (anchor)
- sealed split (sealed plan)
- legacy k4l2 panel
— carries the accuracy-by-field block (narrated
+fieldsarm: holding 0.807 · progress MAE 0.062 · event 0.878 · visible 0.319 over ~9k judge-labeled frames; curated_v0 panel above has its own: 0.814/0.063/0.879/0.316) — pre-reg note - state-masked Q4 probe · state-probe analysis (results)
Box batch 40k AR arms (results)
- A-s0
fontaine_arb_rcond_40k_1xh100panel · state-masked probe - A-s1 seed replicate panel
- A-s2 seed replicate panel
- aux-off arm state-masked probe
- batch analysis JSON
State-dropout arm C fontaine_arb_rcond_statedrop80_40k_1xh100 @40k (results)
Molmo2 AR trunk fontaine_molmo2_ar_40k_ddp4 @40k (results)
- endpoint panel, greedy (curated v0 k4l2) — the #17 BEATS read (6.0079/2.1871 vs A-s0, paired −1.717)
- frozen endpoint analysis JSON
(
molmo2_endpoint_results.py, pre-registered reads) - draws10_t1 frozen analysis JSON — leaderboard row 9, Δ_AR −0.154 mean-collapse read
- decode-cost microbench JSON — rows 8+9 cost cells (box-measured)
- Accuracy-by-field: missing from this panel by bug (the narrated
pass silently skipped molmo2 checkpoints; found + fixed
2f4d5752026-08-08). The table landed at the 60k endpoint instead — see the fields panel results in the @60k section below
Molmo2 AR trunk fontaine_molmo2_ar_60k_ddp4 @60k (results · fields panel)
- endpoint panel json, greedy (curated v0 k4l2)
— 5.8602/2.0719, the IMPROVED endpoint (−0.139 paired vs 40k) that
repointed the attach screen to
step_060000 - accuracy-by-field table json — the registered fields panel read at this endpoint (visible slots 0.32 → 0.82 vs the Gemma trunk)
- frozen 60k-vs-40k analysis JSON
(
molmo2_60k_results.py, pre-registered reads) - browsable HTML panel, greedy
— per-frame retained predictions (32 samples), rendered at eval
time 23:49Z 08-08. (Correction: the eval DID run with
--report— the launcher always had it; the HTML sat unsynced on the box. The earlier “needs a ~1 GPU-h re-run” note was wrong; no re-run needed.) - browsable HTML panel, fields run
— same panel with the narrated
+fieldsarm riding (the accuracy-by-field source run) - Checkpoint weights on the hub:
fontaine-checkpoints/fontaine_molmo2_ar_60k_ddp4/step_060000
Molmo2-ER trunk fontaine_molmo2_er_60k_ddp4 @15k (mid-training, owner-requested)
- browsable HTML panel, greedy (curated v0 k4l2) — owner request 08:29Z 08-10: the @15000 checkpoint hub-copied, downloaded locally, and panel-evaled mid-run; per-frame retained predictions (32 samples), rendered 11:51Z 08-10
- panel json — pooled 7.5283/3.5590 (quarter-training snapshot; the run’s own decision point stays the @60k endpoint panel)
- frozen record-only reads JSON
(
er15k_panel_reads.py) — paired vs banked 40k endpoint +1.52 CI95 [+1.39, +1.54], vs 60k continuation +1.67 [+1.52, +1.68], both ABOVE-BASELINE as expected at 15k/60k steps - Checkpoint weights on the hub:
fontaine-checkpoints/fontaine_molmo2_er_60k_ddp4/step_015000
MolmoAct2-SO100_101 out-of-band panel (pre-reg · plan/deep-read)
- 3-policy side-by-side HTML report — owner spec 11:59Z 08-10: flow teacher 80k (top-10-tickets + stable-key + heun-30 original) vs the released MolmoAct2 SO-100/101 fine-tune vs state-copy, same 25,800 frames, matched 30-step/1.0 s window, 32-frame gallery with 4 policies overlaid per joint; rendered 14:25Z 08-10
- frozen matched-window reads JSON
(
molmoact2_panel_reads.py) — matched-window chunk MAE core frames (willnorris/bbox-2excluded, owner amendment 13:14Z): flow-teacher top-10-tickets 3.90 / state-copy 8.32 / MolmoAct2 13.87 pooled, 16.97 clean-633 vs 7.00 contaminated-245 (beats state-copy only on repos in its own fine-tune mixture, −0.75; trails the flow teacher by +3.29 [+3.11, +3.48] even there) - contamination repo list
— 245/878 panel repos in their
SO100_SO101_MOLMOACT2mixture (7,996 frames, 5,332 core), derived live from their repo file - sweep metadata
— their
predict_actionend-to-end, bf16, 10-step Euler, seed = concat index; 25,800 frames at 352 f/min, ~1.3 GPU-h total
er_60k ENDPOINT @60000 — THE ER decision read (pre-reg): ER init WINS both legs
- endpoint eval report — chained in-unit panel (rc 13:28Z 08-11, ~153/155 GPU-h run total): fast path 5.7782/1.9898 core — the best banked trunk number to date; narrated arm 5.83 (+0.055 pairing, 45% win); aux holding 0.915 / progress MAE 0.060 / event 0.858 / visible 0.822
- paired decision reads JSON
(
er15k_panel_reads.py, keybijou@60000) — vs 40k endpoint (6.0079) pooled −0.2297 [CI95 −0.281, −0.154] BELOW-BASELINE; vs 60k-cont (5.8602) pooled −0.0821 [CI95 −0.126, −0.025] BELOW-BASELINE, CI excludes zero = the pre-registered decision read: the ER-init trunk beats both banked baselines at matched panel class; state-copy integrity byte-match ×3. Rung trajectory 15k +1.52 → 35k +0.28 → 55k −0.18 → 60k −0.23 vs the 40k endpoint. Rig-data effect read at endpoint: NOT split-compatible — the panel contains no owner-rig repos (checked against the npzrepo_ididentity), recorded as skipped per the pre-reg’s if-clause - Weights:
step_060000(weights-only, hub-uploaded 12:44Z in 42.0 s, commit 4ed3dd0)
er_60k @60000 events one-off (owner request 12:44Z 08-11, record-only)
What events does the model actually see? Generated event strings vs
the weak judge labels on the 8,987 judge-labeled panel frames, via
the new --dump-generations instrument (commit 7f43c54 — main-arm
generations retained under explicit --generate).
- standalone report — 13-class model×gt confusion (incl. none/none), per-class P/R, and 136 image cards across hit / class-swap / miss / false-alarm galleries + probe examples (repo-diverse selection)
- Headline: both-none 7,238 · hits 333 · swaps 129 · misses 683 · false alarms 604. On the 1,145 gt-event frames the model speaks on 40%, but class-agrees 72% when it does; exact-string match 3.6% (same event, different words)
- constrained-probe JSON
— on the misses, a 1-step
none-ban re-decode (frame’s own generated prefix replayed; unbanned replay reproducednonebit-exact 679/683): forced guess lands the gt class 63% → the dominant miss mode is saw-it-under-threshold, not blindness (idle 86% / release-place 80% / occlusion 72% / blur 62%; camera quirks 10% and episode markers 0% are the genuinely-not-encoded tail) - confusion JSON · dump-pass eval json · per-frame generations dump (25,800 rows, ~1.55/4 GPU-h). Instrument oracle: presence acc 0.8568 vs banked 0.8582 — Δ 13 frames, inside the documented cross-world-size bf16 batch-composition band (banked ran 4-way on the box)
er_60k @55000 owner-requested panel, standard both-arms (pre-reg, record-only)
- standard eval report — the @55000 read (rc=0 12:00Z 08-11, ~2.2/8 GPU-h): fast path 5.8269/2.0172 core + narrated arm 5.869 (+0.039 pairing, 46% win — same ~0.04–0.05 narration-cost class as 15k/35k); aux vs weak labels at full n≈8,987: holding acc 0.915→0.920, progress MAE 0.065→0.060, event acc 0.875→0.858, visible acc 0.823→0.822
- class-matched paired reads JSON
(
er15k_panel_reads.py, keybijou@55000) — vs 40k endpoint (6.0079) pooled −0.1810 [CI95 −0.232, −0.105] BELOW-BASELINE (first such read for the ER trunk; @35000 was +0.281 above), vs 60k-cont (5.8602) −0.0334 [−0.078, +0.024] CI-SPANS-0 = parity at 92% training; state-copy integrity byte-match ×3; record-only — the @60000 endpoint panel decides - Weights:
step_055000(weights-only, hub-uploaded 09:4xZ in 42.9 s)
er_60k @35000 owner-requested panel, standard both-arms (pre-reg, record-only — SUPERSEDES the aux-arm read below)
- standard eval report — the complete 35k read (rc=0 00:41Z 08-11, ~2.2/8 GPU-h): fast path 6.2892/2.3746 core + narrated arm 6.342 (+0.047 pairing, 44% win — narration costs the same ~0.05-class as at 15k); aux vs weak labels at full n≈8,987, ALL FOUR improved from 15k: holding acc 0.899→0.915, progress MAE 0.075→0.065, event acc 0.862→0.875, visible acc 0.704→0.823
- class-matched paired reads JSON
(
er15k_panel_reads.py, keybijou@35000) — vs 40k endpoint (6.0079) pooled +0.2813 [CI95 +0.199, +0.337], vs 60k-cont (5.8602) +0.4290 [+0.353, +0.467]; ABOVE-BASELINE at 58% training, the 15k gap (+1.52) ~82% closed; state-copy integrity byte-match ×3
er_60k @35000 aux-narrated arm (superseded by the standard read above)
- aux-narrated eval report
— owner request 20:47Z 08-10:
--generate subgoal holding progress event visible(actions follow the model’s own generated aux lines); core 6.3425/2.3770 at 58% training (er15k narrated-class was 7.601), win-rate 77% vs state-copy, Q3 condition sensitivity 1.62 - paired reads JSON — vs 40k endpoint +0.335 [+0.247, +0.387], vs 60k-cont +0.482 [+0.399, +0.517]; cross-class caveat (narrated arm vs fast-path baselines) — the standard both-arms eval relaunched same-session supersedes these with class-matched reads when it lands (~01:0xZ 08-11)
- Weights:
step_035000(weights-only, hub-uploaded 20:5xZ in 42.4 s)
MolmoAct2 SO-101 rig fine-tune (pre-reg · runbook · results)
- anchor-rung HTML report
—
rig_ft_r1(AE-only, 2000 steps, ~2.7/12 GPU-h): rung curve zero-shot 28.95 → 3.23@2000 vs state-copy 9.08 on the 240 rig anchor frames; per-timestep curves, motion-corr small multiples, 8-frame strided trajectory gallery. Pre-reg PASS at every gate; reads are train-frame sanity (contaminated by construction — the real eval is on-rig rollouts, runbook §3–4) - Frozen reads:
zero-shot/preflight ·
step 500 ·
step 1000 ·
step 1500 ·
step 2000
(
molmoact2_rig_preflight.py --model <rung>, identical 240 rows) - Weights on the hub:
molmoact2_so101_rig_r1_step2000— AE + resized-embedding delta vs the released checkpoint (trunk deduplicated, 704/707 tensors sha-verified byte-identical); serve-ready dir stays local at~/checkpoints/molmoact2-so101-rig-r1-step2000-hf
Golden-ticket noise screen (close-out · visual report)
- Frozen stage analyses: stage 1 (R1 CONFIRM) · stage 2 (R2 REAL) · stage 3 (R3 INTERESTING + R4a/R4b)
- jerk-pick selector read — SDN smoothness-prior test on banked draw stacks (flow null / AR small)
Frozen-trunk flow experts @10k, panel_v2 (attach memo · tiny results)
Both experts sit on the hard-frozen 60k trunk (sha e6ed783b),
scored on the panel_v2 k4l2 plan (15,056 core frames pooled) —
numbers compare within this section, not with the v1 scoreboard.
- F arm
fontaine_molmo2_flow_frozen_10k_ddp4@10k, Heun-30 single-draw stable — 9.4157/2.9581, the attach-screen F endpoint (previously box-only; pushed with the tiny readout) - tiny arm
fontaine_molmo2_flow_tiny_h256_10k_1xh100@10k, same decode — the T1 capacity rung endpoint - frozen Δ_capacity analysis JSON
(
attach_seam_results.pyread-1 machinery at explicit paths, per the pre-reg) - Checkpoints on the hub:
F/step_010000·tiny/step_010000(both weights-only, backbone deduplicated to the 60k trunk)
100-seed sim policy eval (pre-reg · results)
Five arms × seeds 0–99 in the v0 SO-101 sim (sysid’d servos), paired design; primary metric = boat→disk progress (cm). 0/500 successes; the engagement/direction split is the finding.
- HTML report + video gallery — per-arm tables, paired CIs, four charts, best/median/worst clips per arm (+ the er60k reach-but-miss money shots)
- frozen analysis JSON
(
sim100_reads.py: gates, summaries, paired bootstrap reads, ordering read auto-skipped — rung arms killed by the phase-2 amendment)
Contact-shadow pass v4 — the composite’s missing shadow, fitted and gated (lit page, 08-13)
The v3 composite’s pasted arm casts no shadow on the real plate —
the one physics law every real frame obeys that no composite frame
did. Leg (a) measured the real arm’s own shadow from 200 frames × 25
bank episodes (frame ÷ episode-plate darkening vs the sim-replayed
silhouette slid along candidate light directions): real and
directional — contrast +0.091 CI95 [0.081, 0.100] vs ring control,
zenith 30° / azimuth 112.5° (85% bootstrap stability), strength
0.392, softness σ 24 px. render_style="v4" = v3 + the fitted
shadow multiply-darkening the top plate (shared projector
sim/shadow.py, 12 oracles; wrist bit-identical to v3). Paired
encoder gate (seeds 0..99, fresh both arms — the banked v3 anchor
0.673 predates the bracket flip; fresh v3 reads 0.721): top 5-NN
AUROC 0.721 → 0.715, and the paired per-seed read is decisive —
Δknn5 −1.04e-07 CI95 [−1.53e-07, −5.6e-08], 66/100 seeds closer,
~10% of the remaining top-cam knn5 excess closed. Wrist 100/100
tied. GO recorded; default stays v3 pending the sim100 amendment-5
owner call. ~0.04 GPU-h. For scale: v1 scene −0.049, v2 inpainting
−0.103, v3 content −0.100, v4 shadows −0.006 — the tail is thinning.
Fitted wrist lens — cubemap render path + gate: the fit’s center term double-counts the pose, the curve-only refit passes (lit page, 08-13)
The deployed wrist warp assumed an ideal equidistant lens centered
at the image midpoint; the plumb-line fit on the 150 pinned real
frames (leg (a), 08-13 01:4xZ) measured the real module off-center
(22 px left / 14 px down, ~5σ) with stronger peripheral compression
(−12.8 px at the corner, CI-excludes-0). Leg (b) landed the render
path that can draw ANY lens: the wrist source is a pinhole cubemap
around the camera axis (output→face map precomputed, so runtime is
one bilinear gather; only referenced faces render; face focal
matched to the deployed source so A/Bs read geometry, not
sharpness; the camera-riding headlight is re-pointed at the base
axis per face — without that, face boundaries carry a shading
seam, caught by the rotated-cubemap oracle at mean|Δ| 6.77). Gate
read (pre-reg 03:27Z, 20 seeds × 5 draws, er60k trunk, control
0.560): full fit 0.667 FAIL — and a labeled post-hoc center-only
arm reads 0.672, reproducing the whole regression. The 08-12
wrist pose re-tune was fit to real frames under the deployed lens,
so it already absorbed the principal-point offset (~2.6°
yaw-equivalent); bolting the fitted center on top applies it
twice. The curve-only refit (k2 +0.101, k4 −0.036) passes: 0.523
≤ the 0.548 gate, paired Δknn5 −7.6e-07 CI95 [−8.5e-07,
−6.8e-07], 96/100 frames closer — ~7× the contact-shadow GO
effect, and cost-neutral (single face covers the frame: 73 vs 70
ms/tick). lens_model="fitted" now pins the curve-only params;
default stays equidistant pending the sim100 amendment-6 owner
call. Top cam bit-identical across all arms (0.713 — now the
frontier number). Full-fit center use is parked behind a joint
pose+lens refit (sim-joint-pose-lens-refit, owner-held).
~0.04 GPU-h total (4 probe arms).
- gate chart · sample frames: equidistant · full fit · curve-only
- gate JSONs: equidistant · full fit · center-only · curve-only
- leg (a): fit JSON · fit chart — plumb-line θ→r fit, decompositions, bootstrap
Appearance screen CONSOLIDATED (report, 08-14)
The whole top-cam appearance programme in one chart-led post, written for the three pending promotion decisions: nine pre-registered reads, ~0.2 GPU-h total — clutter patches carry the removable share (0.713 → 0.556), materials are free riders, texture refuted twice, wrist-neutral, stack 0.5521 sub-additive; the remaining gap to the 0.328 pipeline floor is arm geometry/relief (renderer-class, priced separately). The individual entries below remain the frozen records.
Top-cam gap decomposition — the whole 0.713 residue lives in the rendered foreground pixels (08-13)
With the wrist at 0.523 under the curve-only fitted lens, the
frontier is the top cam’s 5-NN AUROC 0.713 — unmoved by every lens
arm. The decomposition probe asked WHERE it lives before spending on
any lever: five ablation arms on the pinned 20×5 reset-render
harness (numpy backend, fresh in-run v3 baseline reproduces the
banked anchor exactly at 0.713). The verdict is unambiguous.
Swapping the rendered foreground for real dynamic pixels — real
arm/boat/hand pixels mined from bank-episode frames, re-lit
source→target episode via the bank affines, pasted on a different
drawn plate by the exact production arithmetic — collapses the read
to 0.328, at the clean-repo anchor (0.283) and below the 0.5 null:
the compositing pipeline (plates, mask feathering, grade, noise) can
reach real-level; only the pixels it pastes are wrong. The
arithmetic residue read agrees from the other side: fg→plate (0.869)
vs plate-only (0.865) differ by +0.004 AUROC (paired Δknn5 +2.3e-07
CI95 [1.4e-07, 3.3e-07] — real but ~5% of the armless shift and far
under the +0.05 registered bar), so edge-weighting + foreground-PSF
carry ~nothing. Two labeled surprises: (1) the fg→plate/plate-only
arms read FARTHER than v3 (0/100 frames closer, +4.5e-06) — an
armless frame is itself OOD (every real frame has the arm
somewhere), so registered rule 1’s first clause misfired for a
measured, explainable reason; the real-fg arm carries the
localization instead. (2) The shadow-band crop read is near-ceiling
(v3 0.989, v4 0.988, crop clean anchor 0.304): within the table
region the encoder separates sim from real almost perfectly, and the
fitted shadow does NOT close it locally (paired crop Δ CI includes
0) — but the registered box grew to cover most of the lower frame
(89:480, 81:640), i.e. it includes the rendered arm itself, so it
localizes the signal to “the region containing the pasted render”,
consistent with the real-fg verdict rather than a separate shadow
story. v4’s full-frame paired read replicated the shadow gate on the
20×5 protocol (−8.3e-08 CI [−1.34e-07, −3.1e-08], 66/100 closer).
Decision (registered rule): the next leg is foreground appearance
— and the sample frames name the prime suspect: the untextured
gray clutter stand-ins (cylinder mug, white disk) sit next to
photoreal plates; queued as sim-foreground-appearance-pass with a
content-split leg (clutter vs arm vs benchy, keeping the rest
rendered to dodge the armless confound) before any material work.
~0.02 GPU-h embeds; renders CPU.
- analysis JSON · chart
- arm strip — v3 render / real-fg composite / plate-only, same slot
Foreground content split — the clutter stand-ins (~5% of pixels) carry the removable share (08-13)
Leg (a) of the appearance pass asked WHICH rendered class carries
the 0.713: arm bodies (96 geoms, ~7.1% of pixels), benchy (341,
~0.1%), the clutter stand-ins mouse/mug/laptop/pcb (~5.1%), or the
disk (~0.5%, split out of “clutter” as the always-rendered named
suspect). One production v3 instance was hooked at _composite, so
every slot yields all 10 arms — v3, plate-only, no_(class),
only_(class) — through the exact production arithmetic with a
segmentation-restricted mask: same physics, same drawn plate, same
sensor noise (RNG-state restore), making the paired Δ exactly the
class’s visible-pixel effect (in-run oracle: hooked v3 bit-exact ==
the production observation, all 100 slots; fresh v3 read 0.7127,
inside the registered abort band). Removing the clutter stand-ins
alone collapses the read 0.713 → 0.576 (paired Δknn5 −1.73e-06
CI95 [−1.92e-06, −1.54e-06], 99/100 frames closer) — the unique
class past the registered ±0.05 material bar: no_disk −0.006 and
no_benchy −0.002 are CI-excl-0 but immaterial, and no_arm reads
+0.113 WORSE, the armless-content confound the decomposition
labeled (every real frame has the arm). The keep-only duals all
pull toward real when added to the bare plate (only_arm 0.654,
only_clutter 0.824, only_benchy/only_disk 0.832 vs plate-only
0.866), so no class is rendered badly enough to overwhelm its own
content benefit — the ranking rests on the removal direction, which
is also the honest one for clutter (real episodes genuinely vary
clutter presence; the bank plates are mined clutter-free).
Registered primary rule fires: leg (b) target = clutter
appearance (real-crop textures or plate-sourced patches for the
gray untextured stand-ins). Ceiling note, registered before leg
(b): no_clutter’s 0.576 still sits far above the real-fg anchor
0.328, and the arm carries most of that remainder (only_arm 0.654
vs the real-content direction ~0.33) — clutter alone cannot close
the gap, it is just the best ROI per rendered pixel. Renders CPU
(~5 min), embeds 12 groups ~0.02 GPU-h.
- analysis JSON · chart
- arm strip — v3 / no_clutter / only_clutter, same slot
Foreground appearance fix — real-crop clutter patches beat the removal ceiling, gate PASS (08-13)
Legs (b)+(c) of the appearance pass (pre-reg in-channel 05:23Z)
executed the registered follow-up: replace the untextured gray
stand-ins with real-pixel crops pasted into the plate.
make_clutter_crops.py mined per-object RGBA crops from the bank
episodes’ naive per-pixel medians (source episode = largest measured
blob; alpha = the feathered static-novelty mask vs the
gain/bias-corrected global plate — the same statistic the bank pass
localized the objects with; recomputed areas bit-match the manifest),
normalized to global-plate lighting. clutter_patch.py pastes them
at the drawn poses by inverse warp through the verified analytic
fisheye model (target pixel → object-height plane → rigid
drawn→mined transform → source pixel, bilinear), so translation, yaw
jitter and the fisheye’s local scale all ride the camera model; the
active episode’s affine grades the patch exactly like the rendered
foreground; the fixed_canonical pcb pastes at its real measured
location (identity). Zero extra appearance-RNG draws — slots pair
1:1 with production v3. The leg (a) harness then read three arms off
one hooked instance: patched 0.556 vs v3 0.713 (ΔAUROC −0.157,
paired Δknn5 −2.02e-06 CI95 [−2.21e-06, −1.83e-06], 100/100 slots
closer) — the registered −0.05 gate passes at 3× the bar, and
patched lands 0.020 BELOW the no_clutter removal ceiling 0.576
(75/100 closer, CI-excl-0): real-looking clutter beats clutter-free
plates, as the real reference (clutter present in 15–77% of
episodes) predicts. Integrity: in-run v3 0.7127 inside the abort
band, no_clutter 0.5764 reproduces leg (a) within the registered
±0.01, hooked-v3 bit-exact all 100 slots, clean anchor 0.283
unchanged. Promotion of the patch paste into production v3/v4 is an
owner call (asked in-channel 05:40Z); the remaining ceiling to
real-fg 0.328 is the arm’s ~7% of pixels — a separate future item.
Renders CPU (~4 min), embeds 5 groups ~0.02 GPU-h.
- analysis JSON · chart
- v3 vs patched strip — same slots, gray stand-ins vs pasted real crops
- crops strip — mined RGBA crops: on checker / source naive median / identity paste
Wrist-view read of the arm material fixes — wrist-neutral: the two-flag stack moves ~230 raw px and the CI straddles zero (pre-reg, 08-14)
The wrist-side fact the two pending promotion asks (photometrics +
mount) assumed rather than measured. Both flags are model-level
material writes, so the wrist camera — inches from the recolored
surfaces, its frame a RAW render (no composite) — sees them directly.
Two paired production instances, 20 seeds × 5 draws, settled resets,
er_60k knn5 probe, both cameras; gates all green (in-run TOP 0.713
dead-center; WRIST 0.561 in the registered [0.50, 0.60] reset band;
qpos bit-equal ×100; changed-px tripwire quiet at 0.56% max).
PRIMARY: paired wrist Δknn5 −1.39e-08, CI95 [−4.53, +1.73]e-08
straddles zero (46/100) — wrist-neutral; AUROC 0.561 → 0.560. The
mechanism is visibility: at the home pose the wrist camera sees ~230
raw px of graded surface (servo 208 / PLA 21 / mount 1), so there is
nearly nothing for the encoder to read — no regression (the texture
failure mode did not fire), no gain. The top rider replicated the
mount read’s combo delta bit-for-bit (−1.4937e-07, CI [−2.451,
−0.570]e-07, 0.713 → 0.702) — production reset() observations and
the _composite hook path produce identical frames: the hook was
bit-exact. Registered limitation stands: the 0.828 ROLLOUT-pose wrist
gap (gripper filling the frame mid-manipulation) is a different,
still-open fact — needs banked trajectories or fresh rollouts, priced
separately. Renders CPU (~9 min), embeds 8 groups ~0.02 GPU-h.
- analysis JSON · chart
- frame strip — v3 / stack / amplified-Δ (near-black) / real wrist
Arm micro-texture — a clean negative: statistically-matched grain reads MORE fake, both registered CIs above zero (pre-reg, 08-14)
The registered residual branch of the photometric close, executed and
decisively refuted — the cheap kind of negative result. The graded arm
is locally FLAT vs real (PLA print-layer local contrast 8.36 vs 4.66;
servo glint tail p97 205.6 vs 125.2), so a composite-stage micro-texture
(opt-in arm_texture="v1", deterministic static fields from a private
pinned RNG, zero shared-stream draws, applied under seg masks before
the production remap/blur/noise; 6 test oracles + init checks) was
fitted THROUGH the composite to the mined real statistics: PLA local
contrast landed 8.24 vs real 8.36, servo 10.46 vs 9.22, glint tail
~20% closed, photometric guard loss improved on both populations. The
registered 20×5 read, all gates green (in-run v3_photo 0.698
dead-center, anchors exact): PRIMARY v3_tex vs v3_photo +9.33e-07
CI95 [+8.27, +10.42]e-07 entirely ABOVE zero, 3/100 closer, AUROC
0.698 → 0.751; MECHANISM only_links_tex +1.30e-06 CI95 [+1.22,
+1.38]e-06, 0/100 closer, 0.652 → 0.740 — the texture undoes most of
the grade’s gain. Reading: the pooled per-pixel statistics moved toward
real while the encoder moved away — the probe sees spatial structure,
not marginal statistics; screen-fixed band-limited grain reads as
blotchy mottling (the zoom strip shows it), not as anisotropic,
surface-tracking, shading-coupled print ridges. Composite-stage
stats-matching is the wrong instrument class for texture; the branch
dies in one session at ~0.02 GPU-h. Disposition per the frozen rule:
no promotion ask; sim-arm-surface-texture-mjspec (true UV-mapped
surface texture via the recompile path, physics-preservation oracles
as its bar) queued as the escalation, not auto-run; the photometric
grade (0.698/0.652) remains the arm-appearance frontier.
- analysis JSON · chart · fit record
- frame strip — v3_photo / v3_tex, three slots
- arm zoom 2× — the mottling the encoder flagged, side by side with the smooth grade
Arm SURFACE texture (mjSpec) — the SECOND refutation: true surface-tracking bands still read MORE fake (pre-reg + results, 08-14)
The micro-texture refutation’s registered escalation, executed and
refuted in one session. arm_texture="v2" bakes a quasi-periodic
layer-line texture INTO the 18 PLA link materials via an mjSpec
recompile — bands live in OBJECT space and track the surface, the
exact property the first refutation demanded. Physics hard bar 11/11
oracles green (every model field bit-equal, qpos bit-equal incl. a
60-tick excursion); zero-clip tanh generator with grade-preserving
mean compensation; registered reflection rider (the texture
legitimately shows in the tabletop’s 0.02-reflectance mirror of the
arm — and is then fully absorbed by the PSF blur: composited max |Δ|
0). Fit honesty: period 32 frozen at the plausibility bound (lc
response monotonic — fine bands die in the blur chain), amplitude
CAPPED at the 0.42 no-clip headroom → realized PLA local contrast
6.43 of real 8.36 (grade-only 4.66): the albedo-modulation channel
closes ~41% of the quadrature gap and cannot close the rest. The
registered 20×5 read, all gates green (in-run v3_photo 0.698
dead-center): PRIMARY v3_surf vs v3_photo +3.07e-07 CI95 [+2.42,
+3.71]e-07 entirely ABOVE zero, 14/100 closer, AUROC 0.698 → 0.718;
MECHANISM only_links_surf +1.98e-07 CI95 [+1.36, +2.59]e-07, 27/100,
0.652 → 0.671 — about a third of the micro-texture’s harm, but
confidently fake-ward. Coherence was NOT the missing ingredient.
Diagnostics: the cube shrink-wrap renders sunburst fans on several
faces (not clean layers), and the bands are pure albedo modulation
while real print layers are RELIEF — shading/specular structure the
classic renderer cannot express without a normal-map path. The
arm-texture direction is COLD at this abstraction level; the graded
arm (0.698/0.652) stays the production frontier; no further texture
rung auto-queued.
- analysis JSON · chart · fit record
- frame strip — v3_photo / v3_surf / amplified-Δ / link zooms (the sunburst fans)
Camera-mount material split — mechanism lands (93/100), whole-frame null: the part is fixed but too small to move the frame read (pre-reg, 08-14)
The arm-split’s per-pixel worst offender, measured and fixed — with a
split verdict the pre-reg’s decision rule adjudicates cleanly. The
mount (the wrist camera’s white 3D-printed bracket) shared a material
with a black gripper piece; the fix first made the material
mount-exclusive via a byte-identical detach (the gripper geom
drops to matid=-1 with the color copied — the shipped material
carries exactly mjv’s material-less defaults; oracle-pinned), then
mined the real bracket at recorded poses. The white part can’t
darkness-snap, so its mask rode the dark gripper/wrist per-body
locks plus a brightness guard: 81/156 frames, 91k px — the real
mount region reads neutral light gray [123, 120, 125], luma p50 121
vs the recolor-black composite’s 55. Fit through the production
composite chose the same specular ceiling as both link populations
(spec 1.0, shin 0.1; albedo 0.455/0.430/0.431), loss 177188 → 9028,
composited medians dead-on real. The registered 20×5 read (in-run v3
0.713 dead-center; bridges reproduce the arm-split anchors exactly):
MECHANISM PASSES decisively — only_mount_v1 −1.03e-06 CI95 [−1.16,
−0.90]e-06, AUROC 0.821 → 0.793, 93/100 closer, and against the bare
plate the graded mount reads −2.67e-06 with 100/100 closer — with
the right color, mount presence now beats absence (the no_mount
amputation confound, reversed). But PRIMARY FAILS — v3_mount vs v3
CI95 [−0.07, +1.42]e-07 includes zero, 45/100, AUROC 0.713 → 0.713:
at ~0.66% of pixels the fixed part is below the whole-frame read’s
detection floor. Per the frozen rule: no promotion ask for the mount
flag alone. Record-only rider: the two-flag stack (mount +
photometrics, what the pending promotion asks would flip together)
reads 0.713 → 0.702, CI95 [−2.45, −0.57]e-07 entirely below zero
(61/100) — the photometrics carries it; the mount flag rides at zero
measured frame-level cost if the owner flips both. Amendment 1 logged
pre-read: the locality oracle’s bit-equality was amended to a bound —
the tabletop’s 0.02 reflectance mirrors any arm color change (≤24 px,
≤5 counts measured across all 200 oracle slots vs the 3000 px /
6 count bound). Renders CPU (3 sequential instances), embeds 8 groups
~0.02 GPU-h on the R1-A-freed GPU.
- analysis JSON · chart · mine · fit
- frame strip — v3 / v3_mount / v3_full_fix and only_mount / only_mount_v1, same slot
- mining overlay — the mount mask (blue) riding the gripper-cluster lock (red) on a real frame
Arm link photometrics — a measured material grade lands, both registered CIs below zero (pre-reg, 08-14)
The execution of the arm-split verdict. Instead of guessing a better
arm color, the real arm’s pixels were MEASURED: the sim posed at the
recorded joints of 142 real v2 frames, its silhouette projected
through the production fisheye onto them (per-body FFT darkness-snap
±60 px absorbs the tens-of-px registration offset; ring + absolute
darkness guards, wrist excluded for its dark distractors), pooling
436k printed-PLA and 77k servo-casing pixels. The real black
hardware is brighter than the flat recolor (median luma 66 vs 54),
cool-cast [60, 66, 83], and 16–18% glints — sim rendered 5%/0%.
The missing term was shine, not paint. Albedo solved per channel
through the production composite, specular × shininess by grid: both
populations chose the specular ceiling (spec 1.0, shin 0.1); fit
loss ↓8.5× (PLA) / 2.3× (servo). Landed as opt-in
arm_photometrics="v1" (default byte-identical, zero RNG draws,
5 oracles). The registered 20×5 read, all gates green (in-run v3
0.713 dead-center): PRIMARY v3_photo −2.22e-07 CI95 [−3.08e-07,
−1.38e-07] entirely below 0, AUROC 0.713 → 0.698 (72/100 closer);
MECHANISM only_links_photo −7.37e-07 CI95 [−8.35, −6.42]e-07, 0.705
→ 0.652 (96/100) — the graded links alone now match the no_mount
amputation best (0.654) without removing anything. Residuals for the
registered texture follow-up: print-layer local contrast (real 8.4
vs graded 4.7) and the servo glint tail (p97 206 vs 125). Rider
finding: the camera mount is WHITE in reality, black in sim, and its
material is shared with the gripper wrist-roll piece — a mount fix
needs a material split first. Production-default promotion pends the
owner go. Renders CPU (~9 min, two sequential instances), embeds 7
groups ~0.02 GPU-h alongside R1-A.
- analysis JSON · chart · mine · fit
- before/after strip — v3 / v3_photo and only_links / only_links_photo, same slot
- mining overlay — snapped per-body masks on a real frame (green PLA, red servo)
Arm sub-part split — the links carry 88% of the arm’s signature; mounts are the per-pixel worst (pre-reg, 08-13)
The arm-class follow-on to the content split: the rendered arm
(~7.1% of pixels) is the biggest remaining rendered class after the
clutter patches (patched 0.556 ≫ real-fg 0.328), so WHICH arm
sub-part carries it? Same hooked harness — one production v3
instance, _composite re-run per segmentation subset with RNG-state
restore, 20 seeds × 5 draws — over two exact partitions of the 96
arm-class geoms: gripper+jaw (46 geoms, 0.3% px) / links base→wrist
(44, 6.1%) / camera mounts (6, 0.7%), and follower (48, 3.6%) /
leader (48, 3.5%). All gates green: in-run v3 0.713 (band ±0.005),
bridges plate_only 0.866 / only_arm 0.654 / no_arm 0.825 all inside
their registered ±0.02. The registered rule names LINKS: 88% of
the whole arm’s keep-only paired delta (only_links −4.63e-06 of
−5.26e-06, CI-excl-0; only_links alone reads 0.705 vs plate-only
0.866 — nearly the full v3 0.713). Gripper 26% and mount 31% sit
below both the 60% and 35% thresholds. The instance axis is
sub-additive — only_follower −4.05e-06 and only_leader −4.14e-06
each carry ~77–79% alone — so the encoder saturates on either
instance and a fix must treat both arms. Record-only but
striking: no_mount is the ONLY removal that moves v3 TOWARD real
(0.713 → 0.654, 97/100 frames closer, CI-excl-0) despite the
absence-OOD confound that makes no_arm read +0.113 WORSE — the six
camera-mount geoms are per-pixel the most sim-distinctive thing in
the frame, a cheap rider for the photometric rung. Follow-on queued:
sim-arm-photometric-links (links material fix, both instances,
mount-retexture rider). Renders CPU (~7 min), embeds 16 groups
~0.03 GPU-h.
- analysis JSON · chart
- frame strip — v3 / only_links / only_gripper / only_mount / plate-only, same slot
20-seed behavioral spot-check under v3 (pre-reg, results, 08-12)
Same 20 seeds, physics bit-identical (spawn rows byte-matched in-run), only the rendering changed v0 -> v3. teacher80k improves +0.97 cm paired [CI95 +0.16, +1.81] — the only CI-excludes-zero read, direction flipped toward the disk; er60k (-0.07) and snap30k (+0.06) null. Visual familiarity moves the arm that engages. Amendment: GPU compositor (owner-approved) — 371 -> 94 ms/tick, probe reads preserved (0.669/0.113/0.544). ~1.3 GPU-h (gate 3).
Sim content diversity v3 — plate bank + clutter draws (pre-reg, results, 08-12)
Per-reset content variation for the v2 composite: a bank of 26 per-episode clean plates (each carrying its real episode’s lighting; ghost-free by inlier-median mining) + clutter presence/pose draws from the measured real between-episode spread. Registered bar (top k std/mean ≥ 0.15 AND 5-NN AUROC ≤ 0.790) MISSED on the spread leg (0.038 → 0.114) while the AUROC leg fell 0.773 → 0.673 (k-ratio 1.02× — top composites inside the real spread, best top-cam read yet). Wrist bit-identical to v2 (0.548). Default stays v2 per the registered flip rule; flip put to the owner. Record-only: the real disk wanders 8–29 cm × ±19 cm between episodes. ~0.08 GPU-h (gate 0.3).
Sim wrist-cam periphery re-tune (pre-reg, results, 08-12)
One runtime pose change in _repose_wrist_cam: camera moved from
the wrist top behind the gripper to over the jaw base (≈10 cm
forward, 55°→65° down) — under the 72° fisheye source the old pose
filled the bottom ~40% of frame with gripper-body mass the real
camera never sees. Registered bar (wrist 5-NN AUROC ≤ 0.786)
smashed on the first candidate: 0.900 → 0.548, k-ratio
0.97× — sim wrist frames sit inside the real embedding spread.
Guard green (top 0.773 bit-identical); 20×5 sensitivity 0.550.
Per-episode wrist-plate axis retired. ~0.04 GPU-h (gate 0.2).
Sim visual matching v2 — real-frame inpainting (pre-reg, results, 08-12)
Real clean plates (per-pixel median over the 26 reference-half
episodes; A/B pixel-disjointness verified in video-frame indices)
composited under segmentation-masked rendered dynamic content (arms,
benchy, disk + on-table clutter whose real twins move between
episodes). Registered bar (top-cam 5-NN AUROC ≤ 0.790) MET:
0.890 (v0) → 0.876 (v1) → 0.773; overfit tripwire clear. Wrist
composite regressed (0.951 vs 0.900 — the plate is cross-episode
mush) so the shipped render_style="v2" (new default) keeps the v1
wrist path. Homogeneity unchanged (~4% vs 45% k std/mean) — content
variation stays the diversity lever. ~0.06 GPU-h (gate 0.3).
- v2 primary · homogeneity 20×5 · shipped config
- REAL | v1 | v2 galleries: top · wrist
Sim visual matching v1 — appearance pass + probe re-reads (pre-reg, results, 08-12)
Scene rebuild (real table texture, clutter layout, wrist-cam re-pose,
fisheye remap, color grade, sensor emulation, per-reset appearance
jitter) shipped as render_style="v1"; physics oracle-pinned
bit-identical. Registered bar (top-cam 5-NN AUROC 0.890 → ≤0.790 on
the reset-render probe) missed: best 0.874, final 0.876. Wrist
responded to the camera re-pose (0.835 → 0.786 scene-only) then
regressed under fisheye+grade (0.900). Sim stays ~10× too homogeneous
at the encoder; lighting jitter moves per-seed distance only ~3%.
Named next lever: real-frame inpainting (SIMPLER-RT style).
- v0-render baseline · scene · fisheye · grade · sensor · sensitivity 20×5
- before/after composites: top · wrist
Encoder OOD probe — sim-vs-real at the policy’s eyes (rides the sim100 pre-reg, owner ask 01:11Z 08-12)
Sim frames (banked er60k-arm rollouts) vs real rig frames through the frozen er_60k vision trunk, per camera. Measured gap: top-cam 5-NN AUROC 0.885 / gap ratio 1.54× (wrist 0.828 / 1.33×); the clean-repo control lands inside the real spread (AUROC 0.26) so the shift is sim-specific. Sim is at the edge of the real manifold, not off it — the baseline the visual-matching lever must move.
- frozen analysis JSON
(
sim_encoder_ood_probe.py: pinned frame selection, centroid-cosine primary + 5-NN secondary, AUROC/gap-ratio reads, per-frame distances) - distance strip chart — per camera × metric, three groups; sim’s tight blob at the real distribution’s right tail is the whole story in one look
Grasp-SFT route C fontaine_grasp_sft_joint_corrected @2000 (amendment · chain page)
- flow-head unseen-100 eval report — owner request 08:25Z 08-16: 44/100 successes on unseen seeds 0–99 (euler-10) vs base 9 / corrupt-table stage-C 28 — A §5 verdict TABLE_FIX_POSITIVE (44 > 28+3, overlap band moot); anchor bar, per-seed spawn→final strip, 4-clip gallery, full table; rendered 08:5xZ 08-16 from the banked leg json
- Remaining probe legs (flow-train memorization read, token-unseen vs
R2 bar ≥20, token-base anchor) land ~12:3xZ 08-16; consolidated
verdicts JSON
analysis__grasp_sft_joint_probes.json+ report to follow - standard 256-sample eval report (json) — owner request 09:06Z 08-16, stage-C train256 protocol reproduced (state-copy anchors bitwise 9.3562/9.8678): joint chunk MAE 3.24 vs corrupt-table stage-C 12.56 (which sat WORSE than state-copy 9.36 — the wrist_roll clamp); ~3.9× tighter fit on the same 256 demo frames with the corrected box
- Weights:
molmoact2_grasp_sft_joint_corrected_step2000(weights-only, corrected table baked)
Grasp-SFT v1 grasp_sft_v1_joint_8xa100 @3000 (results · flow isolation · drift saga)
- 3-leg sim100 chain panel
— the chain that dated the collapse and separated the heads
(14:17:56Z 08-17, ~6.2/12 GPU-h): step500 flow 4/100 / step500
token 16/100 / endpoint token under the serving fix
b779ba414/100 vs probe flow 44 and endpoint flow 5 — token ~flat across training while flow never leaves the floor ⇒ the mis-fit normalization table poisons the flow targets, not the shared trunk; anchors bar, head-asymmetry slopegraph, per-seed strips, combined table, 9-clip gallery - frozen chain summary JSON
(
sft_v1_chain_report.py— headline numbers reproduce from the banked leg JSONs) - endpoint flow-head unseen-100 report — 5/100 vs probe 44 (per-seed data log-reconstructed after the box wipe; see the results page’s integrity note)
- Weights:
grasp_sft_v1_joint_step3000(weights-only, byte-verified post-upload)
Grasp-SFT v2 drift discriminator grasp_sft_v2_demosonly_1gpu_disc @1000 (pre-reg · verdict)
The first non-drifting v2-corpus checkpoint (verdict HEALTHY 00:42Z 08-18 ⇒ distributed path convicted), panel-reported per the standing HTML-reports rule.
- browsable HTML eval report
(json)
— current-stack eval on the probe-matched pins (demos holdout 0.1 /
split-seed 0 / 256 samples seed 0 / chunk 30 / euler-10 / batch 12),
32 charted frames: chunk MAE 5.763 vs state-copy 7.671 (paired
−1.95); reproduces the old-stack parity read 5.7626 to 3 decimals —
the in-train probe’s 5.8989 is the known ×1.024 probe-vs-eval
instrument shift from the verdict post.
wrist_roll 12.31 stays the worst motor (3.5× state-copy’s 3.99),
the residue the
--per-dataset-flow-normrerun targets - flow-head unseen-100 sim report
(json)
— the demosonly-v2 grasp cell of the isolation grid (04:19Z 08-18,
~2.2 GPU-h): 11/100 successes on unseen seeds 0–99 (euler-10,
30 s episodes) vs probe 44 / v1-endpoint 5 / base 9; mean progress
2.04 cm (probe: 3.86), 64/100 moved, 0 strikes, 7/11 success seeds
shared with the probe. Healthy training + honest stats +
demos-only corpus does NOT restore probe-level grasping — sits at
the top edge of the broken class’s CI (~2–11), inside the pdnorm
pre-reg’s own 11–19 ambiguous band (calibration note recorded in
the draft pre-launch). Worn row: merged demos-native table via the
default fallback (the leg json’s
stats_repo_idfield carries the rig lookup key — the pre-fix record semantics; fixed for future legs inbba4a45, which post-dates this leg’s launch) - k4l2 panel_v2 leg
(json)
— the baseline side of the pdnorm pre-reg’s paired panel read
(04:57Z 08-18, ~0.5 GPU-h; euler-10 draws-1 stable, chunk 30,
batch 32, protocol pinned in
eval_disc1000_k4l2_panel.sh; npz pairing substrate local underreports/): chunk MAE 58.14 vs state-copy 8.37 — 7× worse than state-copy, 0% win rate on the community panel, from a checkpoint that beats state-copy on its own demos holdout (5.76 vs 7.67 above). The demosonly-v2 checkpoint is a narrow specialist; worst motors shoulder_lift 104 / elbow_flex 99 / wrist_roll 71. Interpretation was hedged pending the instrument audit (which row panel items wore: genuine forgetting vs the demos-recomputed table’s windows at serving) — resolved by the wear audit below: about half window, half collapse. Calibration fact recorded pre-launch: at baseline 58.14 the pre-reg’s +0.05 panel guard is near-vacuous as framed. Relaunch note: attempt 1 (batch 12/workers 8) was input-starved (66 f/min, projected 5.7 GPU-h) and killed at 4.7 min per the first-poll rule; r2 (batch 32/workers 20) ran at 96% util. - panel-row wear audit
— what the 58.14 is made of (
disc1000_row_audit.py, oracletests/test_disc1000_row_audit.py; anchors reproduced to 1e-3). Wear fact: the checkpoint records the MERGED scheme (normalization: "q01q99",per_dataset_flow_norm: false), so the eval never consults per-dataset rows at all — every panel item wore the recomputed-at-launch demos-only global table; “community repos missing from the table” never arises, there is no lookup to miss. Decomposition: 85.8% of core truth elements have ≥1 joint outside the worn box, but the box FLOOR is only 14.40 of the 58.14, and predictions are NOT edge-saturated (≤0.3% on arm joints) — the wear hurts through the affine re-expression, not the clamp. Re-wearing the exact same normalized predictions through honest per-repo rows (fit on the panel’s own truth, 838 repos) halves the row to 27.40; through the released source table, 54.40 (also the wrong window). But the re-worn model is WORSE than a constant repo-box-midpoint null (25.15) — output-wear-corrected, the checkpoint carries no usable signal on community data; its raw predictions sit 22.6 from the constant demos action mean while truth sits 58.2 away (collapse to the demos prior). Verdict: ~half the 58.14 is serving-window re-expression, the residual ~half is genuine model failure on community inputs (state-side crush + forgetting, not separable post-hoc — the state input stayed binned through the demos table; a--molmo-norm-style re-run would separate them if it ever matters). Reading the pdnorm endpoint against this baseline: wear-corrected reference points are 27.40 (re-worn disc-1000) and 25.15 (midpoint null); the real bar stays state-copy’s 8.37 - released-checkpoint k4l2 panel row
(json)
— the pre-SFT released checkpoint (
molmoact2-so101-released, the SFT lineage’s init) through the pinned panel protocol (08:22Z 08-18, ~0.45 GPU-h, record-only PRE-GO; euler-10 draws-1 stable, chunk 30, batch 32/workers 20 at ~1173 f/min 95–100% util; wearing its OWN released source table — q01q99 global stats, no per-dataset rows, the default path is its honest wear): chunk MAE 25.89 on the 15,056 core frames, 9% win rate vs state-copy (8.3678 reproduced ≈ banked 8.37, anchor green); first_mae already 21.99 (global misprediction, not chunk-horizon drift); worst motors shoulder_lift 68.9 / elbow_flex 43.1 — the same two that dominate the SFT row’s 104/99. Read (frozen in the queue item pre-launch): 25.89 is AT the 25.15 midpoint null → community competence was never in reach for this lineage; SFT had ~nothing real to destroy on this panel, so the wear audit’s “genuine collapse” half reads as collapse-to-demos-prior of an already-at-null model, not forgetting of once-held competence. Wear-mismatch caveat DISSOLVED 09:xxZ 08-18 by the honest-wear re-expression below: same-wear released row 27.14 vs SFT 27.40 - released-row honest-wear re-expression
— the released row above re-worn through the SAME honest per-repo
rows the disc-1000 27.40 reference wears
(
released_row_rewear.py, sibling of the wear audit; oracletests/test_released_row_rewear.py; CPU, from the banked npz — no model re-run, output-side only). Anchors green (25.8924 / 8.3678 reproduced), inversion round-trip worst 1.5e-05 deg, midpoint-null identity anchor confirms the honest rows are byte-identical to the SFT audit’s (panels element-identical, null 25.154476 on both sides). Same-wear read: released 27.14 vs SFT 27.40 (Δ +0.26) — wear held fixed, SFT ended within noise of where it started, and both rows are slightly WORSE than the 25.15 repo-midpoint null; per-joint the re-wear trades shoulder_pan / wrist errors up for shoulder_lift 68.9→66.1 / elbow_flex 43.1→36.2 down, same two dominant motors. The anchor-ladder released rung is now the same-wear 27.14 (own-table 25.89 kept in the note); released-vs-endpoint stays the informative comparison on GO, now wear-consistent end to end - paired per-seed read: probe vs disc-1000
— retro shakedown of the frozen paired-read instrument
(
sim100_paired_read.py, oracletests/test_sim100_paired_read.py; the pdnorm endpoint’s registered non-gating read vs this baseline, frozen pre-data): probe 44 vs disc-1000 11 = +33 successes, bootstrap CI95 [22, 44]; discordant seeds 37 probe-only vs 4 disc-only (McNemar exact p ≈ 1.0e-7); paired progress delta +3.57 cm [2.66, 4.46], 80% per-seed win rate. The instrument cleanly separates the healthy/broken classes on banked data — at the pdnorm endpoint it reads against disc-1000’s 100 episodes with these same constants (seed-0 bootstrap, 10k resamples). Rendered: the flow-unseen report above now carries this read as a “Paired read” section (delta tiles + McNemar discordant-seed chart, with the 11–19 ambiguous-band note) viagrasp_sft_joint_unseen_report.py --paired-json(oracletests/test_grasp_sft_joint_unseen_report.py,4cfefae) — the pdnorm endpoint report gets the same section from its own frozen paired json - Weights:
grasp_sft_v2_demosonly_1gpu_disc(steps 500 + 1000, weights-only, banked 00:5xZ 08-18)
Cross-family analyses
New reports land here as their evals finish; if a number in a post has no link yet, its report predates this page — ask and it gets pushed.