Papers
Reviews of the papers behind the literature slices. One page per
paper — or per tight theme cluster where the papers only make sense
together — each covering four things: what the paper contributes,
what experiments it actually ran, what transfers to our setup
and what doesn’t, and which idea or experimental arm it fed
(the #N references point into Ideas).
These pages are written for a reader with less context than the
research log assumes. The one-line hooks in ideas.md remain the
index into the backlog; the page here is the record of what was
actually read and why it mattered. Standing rule (owner,
2026-08-07): every literature slice lands its Papers page in the
same session it is banked. Standing rule (owner, 2026-08-08
18:57Z): every page opens with “The paper in plain words” — a
short, jargon-free summary of what the paper is about and what it
showed, before the dense analysis.
Pages
| Page | Papers | Fed |
|---|---|---|
| π0.5 + Knowledge Insulation | 2504.16054, 2505.23705 | #4 attachment arms, #6 self-subgoal probe, #16 north star, #5 |
| LabVLA | 2606.13578 | #4 — the KI-joint arm is the field’s incumbent |
| Q-VGM | 2606.08015 | #4 — the frozen arm keeps an offline-RL escalation path |
| Test-time selection for VLAs | 2510.05681, 2605.01194, 2602.12281, 2506.17811, 2605.25547, 2607.03751, 2605.28527 | #19 — the six selection flavors behind the best-of-10 ceiling gate |
| Self-Certainty | 2502.18581 | #6 rung (b) — the frozen verifier-free scorer for subgoal-draws selection |
| Progress from logits | 2602.19313, 2605.28231 | #6 rung-(b) escalation routing — history fixes phase (zero-shot, our trunk family); masked-contrast prerequisite verified met; SC scorer cell survives |
| SnapFlow | 2604.05656 | #12 — replicated on our stack; 1-NFE distillation adopted-signal |
| The seam debate | 2604.16067, 2605.30877 | #4 — the escalation branches on both sides of the F/K screen |
| Encoder winners don’t transfer | 2606.14153 | #4 — the scale-transfer caveat on reading Δ_seam |
| Hierarchy & subgoals | 2606.10267, 2607.04816 | #6 — design constraints + escalation map for the self-subgoal probe |
| The one-step menu | 2603.12480, 2603.01469, 2606.05737, 2603.14245 | #12 — the fallback menu SnapFlow made moot, and why it worked |
| Sampling beyond selection | 2603.15757, 2606.03847, 2510.12483 | #1, #19 — noise tickets, variance gates, the adopted ES column |
| The state shortcut | 2506.23944, 2509.18644, 2601.16667, 2602.12032, 2602.06575, 2606.22836 | #9, #11 — the crutch we measured, and the mis-banked p=0.8 citation |
| Grounding & conditioning placement | 2601.16207, 2509.04996, 2602.04208, 2506.01844 | #11 — the acuity-probe triangulation; arm B’s published baseline |
| Action tokenization | 2501.09747, 2512.04952 | #5 — the v3 refit spec + the learned-VQ falsifiers |
| Data & trunks | 2602.09722, 2604.23001, 2606.31382, 2607.10172 | #9, #16, #17, #18.7 — two skim-banked claims corrected loudly |
| The attachment frontier | 2603.10126, 2607.13429, NVIDIA WAM post | #4, #6, #17 — expert memory, the anchoring third recipe, π0.7 |
| LP-FT: the schedule with a matched control and a theorem | 2202.10054, 2405.16747 | #4 f-then-joint (third citation, first with matched control + theory); the a(t),b(t) compute framing |
| VLM4VLA: nine trunks + module freezing | 2601.03309 | #17 vision-unfreeze prior; #10/#17 proxy-collapse; NOT compute-matched caveat |
| APT — seam damage as initialization | 2606.12366 | #4 — why random-init experts wreck trunks; the F-then-joint escalation rung |
| ActionX — pre-train the expert, then unfreeze | fnbot.2026.1806605 | #4 — F-then-joint’s second same-shape citation (+38 Long over joint-from-scratch); #16 — expert-scoped RL pole |
| The initialization thread | 2605.25802, 2601.03309 | #17 trunk criterion, #4 — F’s frozen-vision caveat + the vision-first diagnostic |
| Checkpointing without stalling | CheckFreq, Gemini, 2406.10707, 2511.07035, 2605.17821, 2512.24511 | #18.9 async saves (landed e3bdc93) — design corroborated; pinned-buffer + save-frequency hooks banked |
| ELASTIC — adaptive test-time compute | 2606.31132 | #1 rung-3 candidate (dispersion-gated draw allocation), #19, #6 — R4b’s monotone-dispersion read is this paper’s premise, measured free on our panel |
| RoVer — a 0.2B learned verifier | 2510.10975 | #6 escalation rung (learned PRM if “scorer is the gap”; chunk-step caveat pre-registered ammunition), #19, #1 |
| Label-free selection signals | 2605.10158 (+ 2606.14084 re-read) | #6 scorer rung design constraint (score the SET, not the candidate — read the same session as the NO-SCORER verdict); masked-conditioning scorer sketch; jerkpick already banked on noise-space III |
| VLAFlow: training-objective bake-off | 2607.01586 | #4 attachment decision brief (stop-grad −26 pts, frozen-VLM trade table); #6 aux corroboration; NEW future-latent-alignment hook (#6/#17) |
| Q-guided flow critic | 2607.02092 | #6 escalation map third scorer shape (gradient guidance, flow-side); #1/#19 rung-3 note + uncertainty gate |
| FlowDAgger: latent-space DAgger | 2607.08877 | #16 rig few-intervention adaptation recipe (retention 0.88 vs SFT −0.94); #4 steer-window frozen-capital note; #19/#1 inversion-as-label-source |
| Hy-Embodied-0.5-VLA: the full stack | 2606.14409 | #16 — FlowPRO preference RL (weight-space pole, retention-unmeasured caveat) + H=50 Bézier chunk-stitch deployment lever; #4 joint-pole ledger entry under APT’s condition |
| Decode-time stochasticity | 2605.22493, 2605.29766, 2603.20538, 2605.30660, 2508.20072 | #19 — the dT read’s directional prior; 2nd strike on cheap probe selectors; q-token theory anchor |
| Offline validation: does the panel predict the robot? | 2606.29898, 2605.00066, 2405.05941, 2503.24278, 2602.12691 | #16 — raw-MSE proxy measured at ρ −0.61 (sign flips exist); critical-frame re-pooling rung banked; MMRV for future proxy audits |
| LAFM: learned prior libraries | 2606.23420 | #1 — the rung above the ticket screen on the noise-structure ladder; R4’s task-locality read reinterpreted; DSRL named as the next read if stage 1 CONFIRMs |
| Where should the words come from? HiRoC + VLA-Talker | 2608.05999, 2608.05738 | #6 — two fresh directional priors for tonight’s self-subgoal probe (cold-start misalignment; injected-vs-supervised language); #16 evidence-injection few-shot hook |
| Noise-space steering: the ladder above the ticket | 2506.15799, 2606.01151, 2606.13675 | #1 — DSRL read (the named next-read); LP-DS trust-region guard banked for any CEM escalation; #16 — FRS/DSBC 10-demo frozen-trunk rig lever |
| Noise-space steering II: execution + the human loop | 2606.19774, 2605.10821 | #22 — PAINT banked as the new training-free first arm (chunk-50 π₀, beats the TT-RTC fallback on cost); #16 — UniSteer rig lever #3 (corrections→noise, SFT-then-RL prior); #1 — locality probe noted, no gate change |
| Runtime plan verification: gate, refresh, recover | 2604.02965, 2510.16281, 2512.03913 | #6 — the escalation ladder above rung (a) priced (gate needs recovery; subgoal-draws width scaling); #22 — SV-VLA as a drift-monitor competitor; #19/#1 subgoal-draws bridge |
| Noise-space steering III: attribution + a judge-free selector | 2603.11642, 2606.14084 | #1 — three pre-reg priors for the per-dataset-tickets rung (interaction-dominated locality 39.4% vs 1.4% noise main effect; path-intact sampler is why the channel exists; boundary artifact = named panel-blind unknown of ticket 33); #19 — SDN’s jerk-pick selector queued as a free record-only read on banked draw stacks |
| The loss and the mask: CCE + FlexAttention | 2411.09009, FlexAttention docs | #2b, #18 — CE-memory escalation ladder w/ entry condition (row added retroactively 08-08: the page landed 08-08 without its table row) |
| VEGA: encoder-level 3D-aware alignment | 2605.10485 (+ FiT3D 2407.20229 context) | #17 vu5k — the aux-alignment third pole between freeze and thaw (interpretation lever + named cheap escalation); #11 placement echo; #6 aux-family sighting; Spatial Forcing 2510.12276 banked as a new radar hook |
| HyperVLA: hypernetwork inference | 2510.04898 | #17 trunk ledger — inference-efficiency pole (understand-once/execute-tiny) + the generated-update normalization design rule; #16 rig latency existence proof; MSE-vs-diffusion ablation explicitly NOT read onto AR-vs-flow |
| Async execution II: shrink, smooth, or train | 2603.19199, 2602.23901, 2605.19294 | #22 arm menu re-ranked (HAS-on-decode new rung 2; DEFLECT’s restart-corrected +1.6–2.3 pp; d≈18 still untested by anyone); #16 TTFA accounting + jerk instruments; #12 fourth pole (one-step head, many-step tail) |
| Spatial Forcing: convergence, not score | 2510.12276 | #17 — the aux pole’s second recipe (teacher×depth interaction: VGGT works at LLM-24, collapses at encoder); the 3.8× is a fewer-steps lever, teacher overhead unreported; #11 aux-family; SF may fit single-tower Molmo2 better than VEGA’s |
| RDT2: 10k hours of UMI + the F-shaped recipe | 2602.03310 | #4 F-pole ledger context (AR-first + frozen-trunk expert + distill, no joint stage) pre-Δ_seam; #16 hours-scale data premise + β≈0.23; #5 RVQ priced-first; #12 second 1-NFE production point |
| QDepth-VLA: predict quantized depth tokens | 2510.14836 | #11 aux-family third recipe (generative expert, monocular pseudo-labels); #17 — the only aux-spatial recipe needing no encoder seam (single-tower fallback); #5 quantized-beats-regression +3.9; the −2.9 loss vs −8.5 expert ablation split carried loudly |
| ForesightFlow: teaching the flow to score its own draws | 2606.04968 | #19/#1 — seventh selection flavor; the K-sweep evidence anchor (separate 500M critic FLAT K=1→5, self-scored +5.0 — selector shape > size, third strike on post-hoc probes); #12 — 1-NFE endpoint preview with measured ranking fidelity (τ 0.83); #16 — decoupled-AWFM weight-space recipe |
| Fewer layers than you think (CLP) | 2606.20246 | #17 — trunk-redundancy ledger opens (33–50% of finetuned-VLA depth is CKA twins; 8 of 16 DiT expert layers free); throughput accounting fourth lever class (fewer layers, train+inference, FLOP-count mechanism); #4 — prune-then-attach named sequel arm |
| SEAM: closing the chunk seam in noise space | 2607.04609 | #22 — cheapest bridging arm (closed-form, 1.01× vs RTC’s 1.22×, no training); #1 — the cross-chunk half of the boundary term the SDN read couldn’t see; boundary-incompatibility CPU read on banked npz banked as a free hook |
| Robot Critics that Sweat the Small Stuff | 2606.21572 | #19/#6 — trained-critic pole placed and PARKED (needs rollout labels + a video model; ceiling reads cap the payoff on our decodes); one more point that learning the judge is what makes judging work |
| Qwen-VLA: the early-fusion pole | 2605.30280 | #17 trunk ledger — early-fusion pole staked (Qwen3.5-4B + 1.15B single-stream DiT; OOD 76.9 vs π₀.₅ 41.5, no-fusion-ablation confound loud); #4 — F-then-joint production vote #2 (Stage I frozen-trunk expert warm-start) filed pre-Δ_seam; #19 τ=0.6 deploy sharpening; #16 embodiment prompts + data mixture |
| Observation aliasing: when the frame alone can’t tell you what to do | 2605.14712, 2605.14598 | fieldcond-subgoal-meta-report — NN-divergence frame-mining protocol + the delta-concentration chart as the report’s central claim; #6 — external baseline shape for the subgoal channel (frame-only 9% → intent-conditioned 45.8% on aliased states; DSSP’s strict floor-gap theorem); #11 — aliasing census banked as the entry condition for any history/memory arm |
| Correcting corrected weight decay | 2512.08217 | adamc-100k-live readout — grad-norm chart interpretive frame (flat norms expected, ~nil loss effect; head-exclusion partition validated twice; 10% LR floor on the recommended side; no-steady-state-at-100k caveat); ScionC radar-only |
| Z-1: unfreeze the trunk only when diagnostics say so | 2606.31846 | #4 fjoint rung — joint phase as diagnostic-gated conditional escalation (4th frozen-first vote); #16 post-SFT menu RL pole (+13.2 pts from 1,199 demos, GRPO over flow-SDE log-probs) |
| VLA-Corrector: a 40M drift monitor | 2607.01804 | #6 learned-verifier design constraints (residual target; decoupled external judge +14.8 pp); #22 event-triggered truncation datum (+11.65 of +15.65 pp is when to cut); #19 verifier-family sighting |
| π-StepNFT: step-wise critic-free RL | 2603.02083 | #16 RL-pole entry 4 — the pole’s first measured IND-vs-OOD trade (critic-free +11.1 OOD over PPO, −5.5 IND); #1 ticket-informed-exploration footnote |
| DFM-VLA: discrete tokens that get to change their mind | 2603.26320 | #17 head-axis fourth quadrant (commitment, not discreteness, is the expensive property); #5 MAAT metric-aligned embedding +4.4 pp datum; #16 low-data column (10%: 3.21 vs AR 1.71) |
| OneWM-VLA: a world model on one token per frame | 2605.07931 | #17 predictive-supervision pole, self-anchored variant (14.7M LoRA, no teacher; monotone bandwidth sweep; unsupervised scaffold < nothing); #11 dynamics-aux adjacency |
| HiF-VLA: codec motion vectors as temporal context | 2512.09928 | #11 history-arm candidate representation (MPEG-4 MVs + decode-stage AdaLN), behind the aliasing-census gate |
| Muon-SW: the AdamC correction, re-derived for Muon | 2607.23777 | adamc-100k-live readout — weight-norm chart expected shape (plateau-then-flat = correction working); λ ∝ η now derived 3 independent ways; alignment-cosine probe banked as free second opinion |
| AsyncVLA: re-noise the tokens you don’t trust | 2511.14148 | #17 commitment-axis datum 3 (within-model: revisability ≫ more denoise compute; coin-flip selector keeps 2/3 of gain); #6 verifier ledger (dense per-token ≫ outcome labels; relative-confidence blind spot); #22 negative placement (not async execution) |
| Silent failures: proprio vs vision observability | 2606.03134 | #16 bench constraint (telemetry success flags 32–48% false-positive in clean sim → exteroceptive label audit); #6 verifier ledger (modality > capacity; final-state exteroception carries the precision signal) |
| SA-VLA: spatially-aware flow-matching RL | 2602.00743 | #16 RL-pole entry 5 (naive sparse RL measured NEGATIVE, 77.5 vs 81.0 no-RL; protective-machinery framing; noise-parameterization taxonomy); #11/#17 aux-family fourth mode (frozen feature injection, erosion-proof under RL) |
| StreamVLA: completion-state gating | 2602.01100 | #6 phase-estimation constraint (completion-anchored gate sidesteps the measured mid-execution bottleneck; event-triggered refresh ≈ always-reason at half latency); #22 adjacency (re-reasons, never cuts the chunk) |
| Rollout-free eval: RoboWorld + PolaRiS | 2607.01060, 2512.16881 | #16 eval-substrate menu third tier (PolaRiS scan-to-sim priced, co-training load-bearing, DROID-only calibration; RoboWorld no artifact, judge unvalidated; rig-day scan rider banked); independently replicates our offline-validation read |
| FACTR 2: sensorless torque + force-informed sampling | 2606.12406 | #9 phase-weighted sampling candidate + zero-GPU contact-segmentation gate (Δq_d = action − state, free in every episode); #16 rig-day 10-min free-motion protocol note; current sensor load-bearing, +17% bundles conditioning, code unreleased |
| Is Diversity All You Need? | 2507.06219 | #9 velocity-debias lever (+15% ≈ 2.5× data, diffusion head, never operator-ablated; zero-GPU speed census → panel-MAE correlation → normalization arm chain); rig-relevance-filtering warning; Bridge V2 pilot demoted |
| H2R emergence: the human-video gate | 2512.22414 | #9 human-video lever parked with reopening condition (pays ~2× only atop diverse robot pretraining; base VLM ~zero); #17 embodied-trunk precondition — strengthens er_60k’s rationale; angle-A spares (CLAP/Motus/LingBot) gated off |
Retroactive backlog
The owner asked (2026-08-07) for retroactive pages covering every lit slice banked so far. Grouped by theme, most load-bearing first; landed in three work-session batches the same day. Cleared 2026-08-07 (batch 3): all 42 sources covered. The table stays as the per-paper index; from here the standing rule applies — every new lit slice lands its page in the same session.
The attachment seam (#4) — how to attach a flow expert to a pretrained trunk:
| Paper | arXiv | Status |
|---|---|---|
| π0.5 | 2504.16054 | ✅ page |
| Knowledge Insulation | 2505.23705 | ✅ page |
| LabVLA | 2606.13578 | ✅ page |
| Q-VGM | 2606.08015 | ✅ page |
| AEGIS (gradient asymmetry) | 2604.16067 | ✅ page |
| Wall-OSS-0.5 | 2605.30877 | ✅ page |
| Encoder winners don’t transfer across scale | 2606.14153 | ✅ page |
| AR-VLA (history-aware AR expert) | 2603.10126 | ✅ page |
| Representation anchoring | 2607.13429 | ✅ page |
| VLAFlow (objective bake-off) | 2607.01586 | ✅ page |
Test-time selection & sampling (#19, #1):
| Paper | arXiv | Status |
|---|---|---|
| MG-Select | 2510.05681 | ✅ page |
| VLA-ATTC | 2605.01194 | ✅ page |
| CoVer | 2602.12281 | ✅ page |
| RoboMonkey | 2506.17811 | ✅ page |
| TapSampling | 2605.25547 | ✅ page |
| Look Before You Leap | 2607.03751 | ✅ page |
| What Frozen VLAs Already Know About Success | 2605.28527 | ✅ page |
| Self-Certainty (best-of-N without a judge) | 2502.18581 | ✅ page |
| DVAC (variance-gated replanning) | 2606.03847 | ✅ page |
| Golden Ticket (noise search) | 2603.15757 | ✅ page |
| Energy Policy (energy-score training) | 2510.12483 | ✅ page |
| Guided Action Flow (Q-guided critic) | 2607.02092 | ✅ page |
| FlowDAgger (latent-space DAgger) | 2607.08877 | ✅ page |
One-step decoding & distillation (#12):
| Paper | arXiv | Status |
|---|---|---|
| SnapFlow | 2604.05656 | ✅ page |
| One-Step Flow Policy (OFP) | 2603.12480 | ✅ page |
| MeanFlow one-step VLA | 2603.01469 | ✅ page |
| Let It Be Simple | 2606.05737 | ✅ page |
| GoldenStart | 2603.14245 | ✅ page (screened out) |
Hierarchy & subgoals (#6):
| Paper | arXiv | Status |
|---|---|---|
| Hi-VLA (hierarchy design study) | 2606.10267 | ✅ page |
| CAC-VLA (gated latent-action conditioning) | 2607.04816 | ✅ page |
| π0.7 / world-action models | NVIDIA WAM post | ✅ page |
State shortcut & modality imbalance (#9, #11):
| Paper | arXiv | Status |
|---|---|---|
| Adapt Your Body (proprio masking p=0.8) | 2506.23944 | ✅ page (withdrawn paper) |
| State-free policy | 2509.18644 | ✅ page |
| ReViP (state-dominant bias) | 2601.16667 | ✅ page |
| GAP (phase-guided gradient scaling) | 2602.12032 | ✅ page |
| ThinkProprio | 2602.06575 | ✅ page |
| Cloak (visual EE masking) | 2606.22836 | ✅ page |
Grounding & conditioning placement (#11):
| Paper | arXiv | Status |
|---|---|---|
| IVRA (patch-affinity injection) | 2601.16207 | ✅ page |
| FLOWER (deep-layer pruning) | 2509.04996 | ✅ page |
| SCALE (adaptive temperatures — banked title was wrong) | 2602.04208 | ✅ page |
| SmolVLA (mid-stack conditioning) | 2506.01844 | ✅ page |
Data, tokenization & trunks (#5, #9, #16, #17, #18):
| Paper | arXiv | Status |
|---|---|---|
| FAST (local canon) | 2501.09747 | ✅ page |
| FASTer (learned VQ tokenizer) | 2512.04952 | ✅ page |
| Rethinking VLA scaling (negative transfer) | 2602.09722 | ✅ page |
| Data-engine survey | 2604.23001 | ✅ page |
| VLM-to-VLA parameter redundancy | 2606.31382 | ✅ page |
| LoRA-r32 fine-tuning study (π0 on UR5e) | 2607.10172 | ✅ page |
Vision-encoder freeze/unfreeze (#17, owner question 08-07):
| Paper | arXiv | Status |
|---|---|---|
| OpenVLA (vision-FT ablation) | 2406.09246 | ✅ page |
| MAPS (module-wise proximity scheduling) | 2511.19878 | ✅ page |
| Dual-encoder representation preservation | 2509.11417 | ✅ page |
| VEGA (encoder grounding alignment) | 2605.10485 | ✅ page |
| HyperVLA (hypernetwork inference) | 2510.04898 | ✅ page |
| ActionX (RL expert pre-training) | fnbot.2026.1806605 | ✅ page |
Unfreezing schedules under a compute budget (owner steering 08-09 10:38Z, a(t)/b(t) framing):
| Paper | arXiv | Status |
|---|---|---|
| LP-FT (feature distortion + two-phase schedule) | 2202.10054 | ✅ page |
| LP-FT mechanism via NTK (LLMs) | 2405.16747 | ✅ covered in page |
| VLM4VLA (9-trunk sweep, module freezing, proxy collapse) | 2601.03309 | ✅ page |
Smoothness / boundary family (fed by the 08-09 boundary-incompat read):
| Paper | arXiv | Status |
|---|---|---|
| SEAM (inference-side seam steering) | 2607.04609 | ✅ page |
| FAFM (training-side frequency-space smoothness) | 2606.20135 | ✅ page |
Data ingestion / heterogeneous collection (radar set 08-09):
| Paper | arXiv | Status |
|---|---|---|
| VISTA (UMI adaptation: fisheye VQA + physics validation) | 2606.04708 | ✅ page |
| LAFP (latent-action flow policy) | 2606.10517 | ✅ page |
| Flowing With Purpose (latent-action FM) | 2606.23420 | ✅ already covered: LAFM page (dup caught 08-09) |
Fresh sweep 0810 (adamc readout + fjoint sequencing):
| Paper | arXiv | Status |
|---|---|---|
| Correction of Decoupled Weight Decay (AdamC successor) | 2512.08217 | ✅ page |
| Z-1 (efficient GRPO for flow VLAs, selective joint training) | 2606.31846 | ✅ page |
Radar 0811 (banked hooks from the 0810 fresh sweep):
| Paper | arXiv | Status |
|---|---|---|
| TCFM (trajectory-consistent flow matching, RK4 decode) | 2605.08511 | ✅ page |
| RLDT (SVGD density-transport RL on flow policies) | 2606.08602 | ✅ page |
| FAN (feasible-action-neighborhood prior) | 2604.01570 | ✅ page |
| HiFlow (tokenization-free scale-wise AR-via-FM) | 2603.27281 | ✅ page |
| VLA-JEPA (latent world model) | 2602.10098 | ✅ page |
Radar 0812b (banked hooks from the 0811 refill sweep):
| Paper | arXiv | Status |
|---|---|---|
| VLA-Corrector (detect-and-correct inference, adaptive horizon) | 2607.01804 | ✅ page |
| π-StepNFT (step-wise negative-aware online RL for flow VLAs) | 2603.02083 | ✅ page |
| DFM-VLA (discrete flow matching iterative refinement) | 2603.26320 | ✅ page |
| One-Token-Per-Frame / OneWM-VLA (visual bandwidth in world models) | 2605.07931 | ✅ page |
| HiF-VLA (hindsight/insight/foresight motion representation) | 2512.09928 | ✅ page |
Radar 0814 (banked hooks from the 0813 refill sweep):
| Paper | arXiv | Status |
|---|---|---|
| Hyperball (Fantastic Pretraining Optimizers II, weight-norm equilibria) | 2606.16899 | ✅ page |
| Anytime Pretraining (horizon-free schedules + weight averaging) | 2602.03702 | ✅ page |
| VLA-FAIL (zero-failure-data detection: Mahalanobis + chunk consistency) | 2606.21386 | ✅ page |
| FPO (likelihood-free RFT of flow-matching VLAs, ICRA 2026) | 2510.09976 | ✅ page |
| X-Tokenizer (multimodal action tokenizer as auxiliary supervision) | 2606.14752 | ✅ page |
Radar 0815 (banked hooks from the 0814 refill sweep):
| Paper | arXiv | Status |
|---|---|---|
| Weight-norm criticality (loss spikes from decay+normalization driving scale-invariant norms below a critical floor) | 2607.21005 | ✅ page |
| Weibull weight-scale (three-force norm decomposition; spline recovery of alignment force from sparse checkpoints) | 2606.19367 | ✅ page |
| Decoupled Action Expert (5M MLP ≈ 244M U-Net; task knowledge fits in the conditioning pathway) | 2511.12101 | ✅ page |
| Foresight (learned failure detection over action-conditioned world-model latents, outcome labels only, conformal FPR band) | 2606.23085 | ✅ page |
| RedFlow (offline failure→correction RL for flow VLAs) | 2607.27782 | ✅ page |
| Weight decay improves LM plasticity (pretrain λ 0.5–1.0 beats 0.1 downstream; base loss under-predicts post-finetune quality) | 2602.11137 | ✅ page |
| Learning While Deploying (16-robot fleet offline-to-online RL; DIVL distributional critic + QAM flow-native extraction, frozen trunk) | 2605.00416 | ✅ page |
| FoMo-FD (inverse-transport nonconformity on a success-only flow world model; 96.6% detection @1.3% FA, wrist-cam-dependent) | 2607.27511 | ✅ page |
| VLA-GSE (spectral-init adapter-MoE from the frozen backbone’s SVD; init carries the gain, Gaussian-init lands below LoRA) | 2605.06175 | ✅ page |
| ActionCache (training-free retrieval cache over the flow decode; head-only speedups, trunk untouched — our bottleneck unaddressed) | 2607.06370 | ✅ page |
| MolmoAct2 (AI2 VLA on the Molmo2 trunk: Molmo2-ER backbone, 621M per-layer-KV flow expert, SO-100/101 checkpoint + curated 184h pool) | 2605.02881 | ✅ deep-dive post |
Radar 0817 (banked hooks from the 0816 refill sweep; MolmoAct2 slot satisfied by the owner deep dive above):
| Paper | arXiv | Status |
|---|---|---|
| ArmnetBench v0.1 (3-cell SO-101 arm farm; 2,518 human-scored rollouts over 7 policies × 12 tasks; 2,288 labeled failures released LeRobot-native) | 2607.24481 | ✅ page |
| SAFECAST (contrast-set rollouts for SAFE-style hidden-state failure probes; needs closed-loop re-executions + labeled failures — not offline; flow-policy cells below coin-flip) | 2608.04246 | ✅ page |
| Reflex (timestep-invariant trunk → exact KV reuse across denoising steps + async serving; 2.58× vs a soft baseline, stall 100%→0%) | 2607.14695 | ✅ page |
| Legato (guidance-aware flow objective makes chunk continuation native; −20% completion time vs RTC, smoothness ~flat) | 2602.12978 | ✅ page |
| Compression Gap (encoder gains propagate through continuous heads, blocked by an 80-bit FSQ codebook — tiny non-VLA models, single seed, mechanism asserted) | 2604.03191 | ✅ page |
Radar 0818 (banked hooks from the 0817 refill sweep; every hook needed corrections again):
| Paper | arXiv | Status |
|---|---|---|
| ATHENA (influence-function curation at π-0 3.3B scale — Kronecker projection + low-rank Hessian, 313× vs own dense baseline; rollout-anchored, 9.3h/6.9h corpora, no code) | 2606.16208 | ✅ page |
| ProbeAct (hook wrong both clauses: position regressor on 50k sim-oracle labels + hand-coded kinematic detection, zero detection metrics; trunk decodes position R²=0.968 while action head drifts) | 2606.09740 | ✅ page |
| Qwen-RobotManip (38,100h is ~65% re-rendered human video, ~7,800h real teleop; 5-stage offline state-action filter excluded 81% of RoboMIND-UR; nothing released) | 2606.17846 | ✅ page |
| Plasticity at scale (5M–314M LMs: scale delays, never prevents; onset T ∝ P^0.83; WD clause of the hook was a citation of 2602.11137; health proxies all fail to track onset) | 2606.24752 | ✅ page |
Radar 0819 (banked hooks from the 0818 fresh sweep — new angles: sim2real for SO-class arms, action-space design, VLM-trunk continual learning, cross-embodiment; 14/16 candidates survived the local corpus grep, spares banked in the queue item):
| Paper | arXiv | Status |
|---|---|---|
| Squint (SO-101 vendored into ManiSkill3 — NOT upstreamed — + MIT “SO-101 Task Set”, 8 envs, verified installable; single-task 16×16 wrist-cam visual SAC, 91.3% real vs 96.1% sim, ranking preserved; the rollout-substrate blocker is mechanically gone, visual world far-OOD so relative screens first) | 2602.21203 | ✅ page |
| Demystifying Action Space Design (single-arm AgileX, 13k rollouts, chunked flow policies included, code+data released and verified: chunk-wise delta-joint beats our absolute-joint cell +8.4pp, step-wise delta is the trap, execution-horizon interaction; cheapest justified arm = delta-joint retrain) | 2602.23408 | ✅ page |
| Benchmarking VLAs on SO-101 (320 real rollouts, 4 tasks × 4 policies × n=20; multi-label taxonomy despite its own single-label rule, execution labels saturate 91–100%; prize = 16 unlisted rollout_* LeRobot datasets on the author’s Hub account, unlabeled) | 2606.08881 | ✅ page |
| VLA continual-learning triangle (contradiction dissolves in the tables: all three show zero-replay sequential FT forgets catastrophically; episode replay ρ 0.02–0.2 @ 20% of batches fixes it at 3B full FT real-robot scale; “resistance” = better replay exchange rate from the VLM prior) | 2603.03818 + 2605.26820 + 2603.11653 | ✅ page |
Radar 0820 (banked hooks from the 0819 fresh sweep — new angles: world-model/video pretraining, extra sensing on low-cost arms, imitation scaling laws, eval methodology; 14/16 candidates survived the corpus grep — the two dups were papers we’d already deep-read, one independently re-converged on our banked offline-validation page):
| Paper | arXiv | Status |
|---|---|---|
| Rollout-free eval cluster: RoboWorld (r=0.989 vs RoboArena confirmed but n=8, no artifact released, GPT-4o judge never human-validated) + PolaRiS (r=0.9 over 24 policy-env points, MIT code live — but per-checkpoint co-training is load-bearing and calibration is DROID-only) | 2607.01060 + 2512.16881 | ✅ page |
| FACTR 2 (“no force sensor” hid a load-bearing 100 Hz current sensor; +17% bundles torque-as-observation with re-sampling, sampling-only never ablated; cheapest arm touched is a $2,500 Piper; but the load-bearing input Δq_d = action − state is free in our corpus) | 2606.12406 | ✅ page |
| Is Diversity All You Need? (“expert diversity hurts” was never operator-ablated — the evidence is the velocity-debias gain +15% ≈ 2.5× data, on a DIFFUSION action expert, so flow-head immunity is exactly what their setup contradicts; recipe unreleased; velocity spread is also an eval confound for chunk-MAE panels) | 2507.06219 | ✅ page |
| Emergence of human-to-robot transfer (π0.5+ego: human video ~doubles generalization but ONLY atop diverse robot pretraining; base-VLM init pays ~zero — we sit at the measured no-transfer corner; “threshold” partly our compression, no absolute units published; angle-A spares gated off) | 2512.22414 | ✅ page |
Radar 0821 (banked hooks from the 0820 refill sweep — angles: eval methodology (rich again), imitation scaling laws, extra sensing (audio/current), data curation for robot corpora (new angle, hot); 16 candidates abs-verified by the sweep, 12 survived the corpus grep — the four casualties were all papers we had ALREADY deep-read (MolmoAct2, ArmnetBench, CI-MSE, Compression Gap), a sign the sweep is converging on our own reading list):
| Paper | arXiv | Status |
|---|---|---|
| Quality over Quantity (influence curation anchored to 10–20 held-out demos, NOT rollouts — the offline pole we can compute against our panel; but every policy gain is on 40–50% author-injected failures, “per-episode weighting” was an overread — it’s hard top-N with strong budget sensitivity 36.7→86.7%; no code) | 2603.09056 | ✅ page |
| The Curse of Precision (log N ∝ 1/(P−c) confirmed as the model, R²>0.97 — but it’s a sim-only Franka-only FIT, tightest points extrapolated 23–65× beyond trained N; hook’s “not the task” wrong — low-randomization ablation moved c 2.35→1.00 mm; c needs rollout sweeps, so it’s a rig-phase instrument, not pre-computable) | 2607.23108 | ✅ page |
| NeuralActuator (cost floor broken: the third platform IS the SO-101 — force from Feetech load registers, no current sensor, torque via diffsim not calibration, MAE 0.47–0.73 N; hook’s “torque-from-current” wrong twice at our class; everything MIT-released incl. 3 SO-101 checkpoints + teleop code; corpus still can’t feed it — #9 gate stands, #16 rider shovel-ready) | 2607.11734 | ✅ page |
| GigaWorld-1 / WMBench (324K “rollouts” are human-graded world-model VIDEOS under replayed actions — no policy drives, real-robot ranking correlation defined but never reported; action-faithfulness>realism measured, partly definitional; big news the hook missed: full Apache-2.0 release of Nano 1.3B/Pro 5B + validated open VLM judge, and Ctrl-World is no longer the only released artifact) | 2607.02642 | ✅ page |
| SPARES (8, grep-clean 08-09): 2606.27375 ABC-130K open BC scaling substrate (3,500 h / 130K eps / 195 tasks + recipe sweeps); 2601.18723 Eval-Actions graded execution-quality labels (13K episodes, SRCC 0.81–0.84); 2603.13616 Beyond Binary Success anytime-valid sequential policy comparison (−70% eval burden); 2511.09958 Audio-VLA contact-mic template; 2512.08405 audio world models (flow-matching audio prediction); 2606.17598 MuseVLA frozen-trunk multimodal sensing; 2607.21588 AXIS community data engine (+5.8% from auto-QA); 2605.26349 episode-level teleop quality scoring | — | banked in queue |
Radar 0822 (banked hooks from the 0821 refill sweep — angles: curation still the richest vein (a coherent curation-metrics testbed cluster surfaced), eval methodology keyword-rich/citation-thin, scaling medium, motor-current sensing nearly rested; 18 candidates checked, 15 abs-page-verified, 12 survived — the 3 dups were all papers we had already deep-read, the sweep keeps converging on our own reading list):
| Paper | arXiv | Status |
|---|---|---|
| Ambient Diffusion Policy (MIT/Tedrake, RSS “It’s the demos” spotlight: suboptimal demos contribute only at high/low diffusion times, justified by a spectral power law in robot actions; up to +33% over naive co-training, purely offline — a #9 re-weighting lever on the flow/diffusion TIME axis, timestep-gated inclusion instead of hard filtering; check the spectral argument transfers to rectified flow + how the suboptimal split is designated) | 2606.12365 | hook banked |
| What Demonstration Curation Metrics Do to Your Policy (detection accuracy and policy quality sharply decoupled: best defect detector AUROC 0.804 → WORST curated policy 13.3%, weaker 0.638 detector nearly matches oracle 90.0 vs 93.3%; 5 of 7 metrics secretly exploit episode length — a direct confound warning for every #9 curation experiment and for our chunk-MAE panel; testbed released) | 2606.10229 | hook banked |
| Auditing Demonstration Curation Metrics (companion audit: action-only scorers catch noise/tremor/truncation but structural defects — wrong action at a key moment — are invisible to EVERY action-only metric, two actively prefer defective episodes; our positions-only corpus is exactly the failing feature space — the sharpest stress test for the label-free-selection-signals conclusions; check whether their “state” metrics need visual state) | 2606.05588 | hook banked |
| PhAIL (Franka FR3 open benchmark replacing binary-success-at-timeout with time-to-success CDFs: Human-Relative Throughput + bootstrap CIs + per-object KS tests, claims usable resolution at N≤30 rollouts/cell — THE statistical-protocol question for our rig-day rollout budget; needs human reference runs, check the resolution claim isn’t carried by the human anchor; dataset + reference implementation released) | 2605.29710 | hook banked |
| SPARES (8, abs-verified + grep-clean 08-10): 2607.04434 RoboDojo (42 sim + 18 real tasks, cloud real-eval, 30-policy leaderboard — the sim-vs-real alignment substrate); 2605.20774 VLA-REPLICA (low-cost reproducible real-world VLA bench — closest analogue to our #16 design); 2607.15330 Xiaomi-Robotics-1 (100K-h real-trajectory scaling report, checkpoints promised unconfirmed); 2606.15064 Phase-Localized Curation Does Not Help (negative result, same testbed family); 2606.20521 HumanScale (ego human video OUTPERFORMS robot data claim — candidate trigger for the h2r-lowdata-counterexample screen; check for a hidden diverse robot corpus in the alignment stage); 2603.05504 RoboPocket (AR visual-foresight targets demo collection at policy-weak regions, 2× data efficiency); 2606.30988 MuSe (post-hoc force-sensing attachment without forgetting — the #4/angle-B template); 2607.26047 S2A2 (spatial+spectral contact audio across ACT/DP/VQ-BeT/π0) | — | banked in queue |
Sim lane (owner directive 2026-08-11 17:07Z — sim + sim-to-real reading for the SO-101 eval substrate; the general lit pause stays for non-sim topics):
| Page | Papers | Fed |
|---|---|---|
| Sim-as-eval (SIMPLER MMRV 0.056/r 0.924 via sysid + green-screen, not photorealism; AutoEval’s 0/50-sim-vs-47/50-real caution on new policy families; SureSim paired-rectification CIs; continuous progress separates policies at up to 70% fewer trials; 2026 head-to-head: simulator choice moves Spearman 0.400↔0.700 on identical real evals) | 2405.05941, 2503.24278, 2510.04354, 2603.13616, 2606.10366, 2512.19562 | sim-policy-eval-100seeds protocol design; #16 |
| SO-101 sim landscape (census: no public SO-101 sim eval with a continuous metric exists — 2026’s SO-101 benchmarks are all real-world; LeIsaac binary-success/Isaac-heavy, so101-nexus beta, so-frame’s REAL|SIM|OVERLAY worth stealing, LIBERO’s frozen init-states = the 100-seed pattern; our menagerie model’s kp 998 vs TheRobotStudio’s kp 17.8 for the same servo) | 2602.21203, 2512.19562, 2606.08881, 2607.24481, 2605.20774 | sim-policy-eval-100seeds; publish-later option (EnvHub) |
| Contact fidelity (all four sim-review findings have documented mechanisms + named fixes: CoACD threshold-not-cap / SDF escape hatch that also fixes the CC-BY-ND per-machine asset hazard; priority override is spec — explicit contact pair + condim 4 + elliptic cones; SIMPLER ablation: controller sysid first-order for MMRV, friction values second-order; BAM ships an identified STS3215 model) | MuJoCo docs, 2205.02961, 2111.01391, 2410.08650, 2405.05941 | the pre-fix list ahead of the 100-seed pre-reg |
| Composite shadows (the paste-the-robot papers nearly all skip shadows and none measure them; ConCent’s recipe = random light + silhouette projection as randomization; Re³Sim ablation: foreground mesh→splat moves success 0.70→0.70 — scene, not foreground realism, binds; GreenAug-Rand beats generative backgrounds for training → randomize-in-training / match-in-eval split) | 2606.30268, 2503.14526, 2502.08645, 2407.07868 | composite-contact-shadows probe idea; sim-wrist-compositing design |
| Fisheye lens fitting (wrist fisheye 0.988 vs pinhole 0.181 real; policies overfit absolute pixel scale as a distance ruler — 220° lens transfer 0.0025→0.60 with Random Scale Augmentation; cubemap→equirect→any-lens MuJoCo pipeline removes our 72°-source ceiling and makes the real 130° module’s calibrated θ→r curve renderable) | 2603.02139 | fit-real-lens-model idea; v1 wrist render path upgrades |
| DR schedules (randomization width as success-throttled curriculum: DORAEMON entropy-max s.t. success ≥ α beats AutoDR 60% vs 26.7% real on 17-param Panda push; α=0.5 not 0.9 — the policy must be allowed to fail; one-scalar curriculum coefficient gets most of it; eval stays at the matched center, always) | 2311.01885, 1910.07113, 2505.05753, 2111.00956 | dr-schedule-for-sim-rl (conditional on GRPO probe); eval/train firewall rule |