Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Papers

Reviews of the papers behind the literature slices. One page per paper — or per tight theme cluster where the papers only make sense together — each covering four things: what the paper contributes, what experiments it actually ran, what transfers to our setup and what doesn’t, and which idea or experimental arm it fed (the #N references point into Ideas).

These pages are written for a reader with less context than the research log assumes. The one-line hooks in ideas.md remain the index into the backlog; the page here is the record of what was actually read and why it mattered. Standing rule (owner, 2026-08-07): every literature slice lands its Papers page in the same session it is banked. Standing rule (owner, 2026-08-08 18:57Z): every page opens with “The paper in plain words” — a short, jargon-free summary of what the paper is about and what it showed, before the dense analysis.

Pages

PagePapersFed
π0.5 + Knowledge Insulation2504.16054, 2505.23705#4 attachment arms, #6 self-subgoal probe, #16 north star, #5
LabVLA2606.13578#4 — the KI-joint arm is the field’s incumbent
Q-VGM2606.08015#4 — the frozen arm keeps an offline-RL escalation path
Test-time selection for VLAs2510.05681, 2605.01194, 2602.12281, 2506.17811, 2605.25547, 2607.03751, 2605.28527#19 — the six selection flavors behind the best-of-10 ceiling gate
Self-Certainty2502.18581#6 rung (b) — the frozen verifier-free scorer for subgoal-draws selection
Progress from logits2602.19313, 2605.28231#6 rung-(b) escalation routing — history fixes phase (zero-shot, our trunk family); masked-contrast prerequisite verified met; SC scorer cell survives
SnapFlow2604.05656#12 — replicated on our stack; 1-NFE distillation adopted-signal
The seam debate2604.16067, 2605.30877#4 — the escalation branches on both sides of the F/K screen
Encoder winners don’t transfer2606.14153#4 — the scale-transfer caveat on reading Δ_seam
Hierarchy & subgoals2606.10267, 2607.04816#6 — design constraints + escalation map for the self-subgoal probe
The one-step menu2603.12480, 2603.01469, 2606.05737, 2603.14245#12 — the fallback menu SnapFlow made moot, and why it worked
Sampling beyond selection2603.15757, 2606.03847, 2510.12483#1, #19 — noise tickets, variance gates, the adopted ES column
The state shortcut2506.23944, 2509.18644, 2601.16667, 2602.12032, 2602.06575, 2606.22836#9, #11 — the crutch we measured, and the mis-banked p=0.8 citation
Grounding & conditioning placement2601.16207, 2509.04996, 2602.04208, 2506.01844#11 — the acuity-probe triangulation; arm B’s published baseline
Action tokenization2501.09747, 2512.04952#5 — the v3 refit spec + the learned-VQ falsifiers
Data & trunks2602.09722, 2604.23001, 2606.31382, 2607.10172#9, #16, #17, #18.7 — two skim-banked claims corrected loudly
The attachment frontier2603.10126, 2607.13429, NVIDIA WAM post#4, #6, #17 — expert memory, the anchoring third recipe, π0.7
LP-FT: the schedule with a matched control and a theorem2202.10054, 2405.16747#4 f-then-joint (third citation, first with matched control + theory); the a(t),b(t) compute framing
VLM4VLA: nine trunks + module freezing2601.03309#17 vision-unfreeze prior; #10/#17 proxy-collapse; NOT compute-matched caveat
APT — seam damage as initialization2606.12366#4 — why random-init experts wreck trunks; the F-then-joint escalation rung
ActionX — pre-train the expert, then unfreezefnbot.2026.1806605#4 — F-then-joint’s second same-shape citation (+38 Long over joint-from-scratch); #16 — expert-scoped RL pole
The initialization thread2605.25802, 2601.03309#17 trunk criterion, #4 — F’s frozen-vision caveat + the vision-first diagnostic
Checkpointing without stallingCheckFreq, Gemini, 2406.10707, 2511.07035, 2605.17821, 2512.24511#18.9 async saves (landed e3bdc93) — design corroborated; pinned-buffer + save-frequency hooks banked
ELASTIC — adaptive test-time compute2606.31132#1 rung-3 candidate (dispersion-gated draw allocation), #19, #6 — R4b’s monotone-dispersion read is this paper’s premise, measured free on our panel
RoVer — a 0.2B learned verifier2510.10975#6 escalation rung (learned PRM if “scorer is the gap”; chunk-step caveat pre-registered ammunition), #19, #1
Label-free selection signals2605.10158 (+ 2606.14084 re-read)#6 scorer rung design constraint (score the SET, not the candidate — read the same session as the NO-SCORER verdict); masked-conditioning scorer sketch; jerkpick already banked on noise-space III
VLAFlow: training-objective bake-off2607.01586#4 attachment decision brief (stop-grad −26 pts, frozen-VLM trade table); #6 aux corroboration; NEW future-latent-alignment hook (#6/#17)
Q-guided flow critic2607.02092#6 escalation map third scorer shape (gradient guidance, flow-side); #1/#19 rung-3 note + uncertainty gate
FlowDAgger: latent-space DAgger2607.08877#16 rig few-intervention adaptation recipe (retention 0.88 vs SFT −0.94); #4 steer-window frozen-capital note; #19/#1 inversion-as-label-source
Hy-Embodied-0.5-VLA: the full stack2606.14409#16 — FlowPRO preference RL (weight-space pole, retention-unmeasured caveat) + H=50 Bézier chunk-stitch deployment lever; #4 joint-pole ledger entry under APT’s condition
Decode-time stochasticity2605.22493, 2605.29766, 2603.20538, 2605.30660, 2508.20072#19 — the dT read’s directional prior; 2nd strike on cheap probe selectors; q-token theory anchor
Offline validation: does the panel predict the robot?2606.29898, 2605.00066, 2405.05941, 2503.24278, 2602.12691#16 — raw-MSE proxy measured at ρ −0.61 (sign flips exist); critical-frame re-pooling rung banked; MMRV for future proxy audits
LAFM: learned prior libraries2606.23420#1 — the rung above the ticket screen on the noise-structure ladder; R4’s task-locality read reinterpreted; DSRL named as the next read if stage 1 CONFIRMs
Where should the words come from? HiRoC + VLA-Talker2608.05999, 2608.05738#6 — two fresh directional priors for tonight’s self-subgoal probe (cold-start misalignment; injected-vs-supervised language); #16 evidence-injection few-shot hook
Noise-space steering: the ladder above the ticket2506.15799, 2606.01151, 2606.13675#1 — DSRL read (the named next-read); LP-DS trust-region guard banked for any CEM escalation; #16 — FRS/DSBC 10-demo frozen-trunk rig lever
Noise-space steering II: execution + the human loop2606.19774, 2605.10821#22 — PAINT banked as the new training-free first arm (chunk-50 π₀, beats the TT-RTC fallback on cost); #16 — UniSteer rig lever #3 (corrections→noise, SFT-then-RL prior); #1 — locality probe noted, no gate change
Runtime plan verification: gate, refresh, recover2604.02965, 2510.16281, 2512.03913#6 — the escalation ladder above rung (a) priced (gate needs recovery; subgoal-draws width scaling); #22 — SV-VLA as a drift-monitor competitor; #19/#1 subgoal-draws bridge
Noise-space steering III: attribution + a judge-free selector2603.11642, 2606.14084#1 — three pre-reg priors for the per-dataset-tickets rung (interaction-dominated locality 39.4% vs 1.4% noise main effect; path-intact sampler is why the channel exists; boundary artifact = named panel-blind unknown of ticket 33); #19 — SDN’s jerk-pick selector queued as a free record-only read on banked draw stacks
The loss and the mask: CCE + FlexAttention2411.09009, FlexAttention docs#2b, #18 — CE-memory escalation ladder w/ entry condition (row added retroactively 08-08: the page landed 08-08 without its table row)
VEGA: encoder-level 3D-aware alignment2605.10485 (+ FiT3D 2407.20229 context)#17 vu5k — the aux-alignment third pole between freeze and thaw (interpretation lever + named cheap escalation); #11 placement echo; #6 aux-family sighting; Spatial Forcing 2510.12276 banked as a new radar hook
HyperVLA: hypernetwork inference2510.04898#17 trunk ledger — inference-efficiency pole (understand-once/execute-tiny) + the generated-update normalization design rule; #16 rig latency existence proof; MSE-vs-diffusion ablation explicitly NOT read onto AR-vs-flow
Async execution II: shrink, smooth, or train2603.19199, 2602.23901, 2605.19294#22 arm menu re-ranked (HAS-on-decode new rung 2; DEFLECT’s restart-corrected +1.6–2.3 pp; d≈18 still untested by anyone); #16 TTFA accounting + jerk instruments; #12 fourth pole (one-step head, many-step tail)
Spatial Forcing: convergence, not score2510.12276#17 — the aux pole’s second recipe (teacher×depth interaction: VGGT works at LLM-24, collapses at encoder); the 3.8× is a fewer-steps lever, teacher overhead unreported; #11 aux-family; SF may fit single-tower Molmo2 better than VEGA’s
RDT2: 10k hours of UMI + the F-shaped recipe2602.03310#4 F-pole ledger context (AR-first + frozen-trunk expert + distill, no joint stage) pre-Δ_seam; #16 hours-scale data premise + β≈0.23; #5 RVQ priced-first; #12 second 1-NFE production point
QDepth-VLA: predict quantized depth tokens2510.14836#11 aux-family third recipe (generative expert, monocular pseudo-labels); #17 — the only aux-spatial recipe needing no encoder seam (single-tower fallback); #5 quantized-beats-regression +3.9; the −2.9 loss vs −8.5 expert ablation split carried loudly
ForesightFlow: teaching the flow to score its own draws2606.04968#19/#1 — seventh selection flavor; the K-sweep evidence anchor (separate 500M critic FLAT K=1→5, self-scored +5.0 — selector shape > size, third strike on post-hoc probes); #12 — 1-NFE endpoint preview with measured ranking fidelity (τ 0.83); #16 — decoupled-AWFM weight-space recipe
Fewer layers than you think (CLP)2606.20246#17 — trunk-redundancy ledger opens (33–50% of finetuned-VLA depth is CKA twins; 8 of 16 DiT expert layers free); throughput accounting fourth lever class (fewer layers, train+inference, FLOP-count mechanism); #4 — prune-then-attach named sequel arm
SEAM: closing the chunk seam in noise space2607.04609#22 — cheapest bridging arm (closed-form, 1.01× vs RTC’s 1.22×, no training); #1 — the cross-chunk half of the boundary term the SDN read couldn’t see; boundary-incompatibility CPU read on banked npz banked as a free hook
Robot Critics that Sweat the Small Stuff2606.21572#19/#6 — trained-critic pole placed and PARKED (needs rollout labels + a video model; ceiling reads cap the payoff on our decodes); one more point that learning the judge is what makes judging work
Qwen-VLA: the early-fusion pole2605.30280#17 trunk ledger — early-fusion pole staked (Qwen3.5-4B + 1.15B single-stream DiT; OOD 76.9 vs π₀.₅ 41.5, no-fusion-ablation confound loud); #4 — F-then-joint production vote #2 (Stage I frozen-trunk expert warm-start) filed pre-Δ_seam; #19 τ=0.6 deploy sharpening; #16 embodiment prompts + data mixture
Observation aliasing: when the frame alone can’t tell you what to do2605.14712, 2605.14598fieldcond-subgoal-meta-report — NN-divergence frame-mining protocol + the delta-concentration chart as the report’s central claim; #6 — external baseline shape for the subgoal channel (frame-only 9% → intent-conditioned 45.8% on aliased states; DSSP’s strict floor-gap theorem); #11 — aliasing census banked as the entry condition for any history/memory arm
Correcting corrected weight decay2512.08217adamc-100k-live readout — grad-norm chart interpretive frame (flat norms expected, ~nil loss effect; head-exclusion partition validated twice; 10% LR floor on the recommended side; no-steady-state-at-100k caveat); ScionC radar-only
Z-1: unfreeze the trunk only when diagnostics say so2606.31846#4 fjoint rung — joint phase as diagnostic-gated conditional escalation (4th frozen-first vote); #16 post-SFT menu RL pole (+13.2 pts from 1,199 demos, GRPO over flow-SDE log-probs)
VLA-Corrector: a 40M drift monitor2607.01804#6 learned-verifier design constraints (residual target; decoupled external judge +14.8 pp); #22 event-triggered truncation datum (+11.65 of +15.65 pp is when to cut); #19 verifier-family sighting
π-StepNFT: step-wise critic-free RL2603.02083#16 RL-pole entry 4 — the pole’s first measured IND-vs-OOD trade (critic-free +11.1 OOD over PPO, −5.5 IND); #1 ticket-informed-exploration footnote
DFM-VLA: discrete tokens that get to change their mind2603.26320#17 head-axis fourth quadrant (commitment, not discreteness, is the expensive property); #5 MAAT metric-aligned embedding +4.4 pp datum; #16 low-data column (10%: 3.21 vs AR 1.71)
OneWM-VLA: a world model on one token per frame2605.07931#17 predictive-supervision pole, self-anchored variant (14.7M LoRA, no teacher; monotone bandwidth sweep; unsupervised scaffold < nothing); #11 dynamics-aux adjacency
HiF-VLA: codec motion vectors as temporal context2512.09928#11 history-arm candidate representation (MPEG-4 MVs + decode-stage AdaLN), behind the aliasing-census gate
Muon-SW: the AdamC correction, re-derived for Muon2607.23777adamc-100k-live readout — weight-norm chart expected shape (plateau-then-flat = correction working); λ ∝ η now derived 3 independent ways; alignment-cosine probe banked as free second opinion
AsyncVLA: re-noise the tokens you don’t trust2511.14148#17 commitment-axis datum 3 (within-model: revisability ≫ more denoise compute; coin-flip selector keeps 2/3 of gain); #6 verifier ledger (dense per-token ≫ outcome labels; relative-confidence blind spot); #22 negative placement (not async execution)
Silent failures: proprio vs vision observability2606.03134#16 bench constraint (telemetry success flags 32–48% false-positive in clean sim → exteroceptive label audit); #6 verifier ledger (modality > capacity; final-state exteroception carries the precision signal)
SA-VLA: spatially-aware flow-matching RL2602.00743#16 RL-pole entry 5 (naive sparse RL measured NEGATIVE, 77.5 vs 81.0 no-RL; protective-machinery framing; noise-parameterization taxonomy); #11/#17 aux-family fourth mode (frozen feature injection, erosion-proof under RL)
StreamVLA: completion-state gating2602.01100#6 phase-estimation constraint (completion-anchored gate sidesteps the measured mid-execution bottleneck; event-triggered refresh ≈ always-reason at half latency); #22 adjacency (re-reasons, never cuts the chunk)
Rollout-free eval: RoboWorld + PolaRiS2607.01060, 2512.16881#16 eval-substrate menu third tier (PolaRiS scan-to-sim priced, co-training load-bearing, DROID-only calibration; RoboWorld no artifact, judge unvalidated; rig-day scan rider banked); independently replicates our offline-validation read
FACTR 2: sensorless torque + force-informed sampling2606.12406#9 phase-weighted sampling candidate + zero-GPU contact-segmentation gate (Δq_d = action − state, free in every episode); #16 rig-day 10-min free-motion protocol note; current sensor load-bearing, +17% bundles conditioning, code unreleased
Is Diversity All You Need?2507.06219#9 velocity-debias lever (+15% ≈ 2.5× data, diffusion head, never operator-ablated; zero-GPU speed census → panel-MAE correlation → normalization arm chain); rig-relevance-filtering warning; Bridge V2 pilot demoted
H2R emergence: the human-video gate2512.22414#9 human-video lever parked with reopening condition (pays ~2× only atop diverse robot pretraining; base VLM ~zero); #17 embodied-trunk precondition — strengthens er_60k’s rationale; angle-A spares (CLAP/Motus/LingBot) gated off

Retroactive backlog

The owner asked (2026-08-07) for retroactive pages covering every lit slice banked so far. Grouped by theme, most load-bearing first; landed in three work-session batches the same day. Cleared 2026-08-07 (batch 3): all 42 sources covered. The table stays as the per-paper index; from here the standing rule applies — every new lit slice lands its page in the same session.

The attachment seam (#4) — how to attach a flow expert to a pretrained trunk:

PaperarXivStatus
π0.52504.16054page
Knowledge Insulation2505.23705page
LabVLA2606.13578page
Q-VGM2606.08015page
AEGIS (gradient asymmetry)2604.16067page
Wall-OSS-0.52605.30877page
Encoder winners don’t transfer across scale2606.14153page
AR-VLA (history-aware AR expert)2603.10126page
Representation anchoring2607.13429page
VLAFlow (objective bake-off)2607.01586page

Test-time selection & sampling (#19, #1):

PaperarXivStatus
MG-Select2510.05681page
VLA-ATTC2605.01194page
CoVer2602.12281page
RoboMonkey2506.17811page
TapSampling2605.25547page
Look Before You Leap2607.03751page
What Frozen VLAs Already Know About Success2605.28527page
Self-Certainty (best-of-N without a judge)2502.18581page
DVAC (variance-gated replanning)2606.03847page
Golden Ticket (noise search)2603.15757page
Energy Policy (energy-score training)2510.12483page
Guided Action Flow (Q-guided critic)2607.02092page
FlowDAgger (latent-space DAgger)2607.08877page

One-step decoding & distillation (#12):

PaperarXivStatus
SnapFlow2604.05656page
One-Step Flow Policy (OFP)2603.12480page
MeanFlow one-step VLA2603.01469page
Let It Be Simple2606.05737page
GoldenStart2603.14245page (screened out)

Hierarchy & subgoals (#6):

PaperarXivStatus
Hi-VLA (hierarchy design study)2606.10267page
CAC-VLA (gated latent-action conditioning)2607.04816page
π0.7 / world-action modelsNVIDIA WAM postpage

State shortcut & modality imbalance (#9, #11):

PaperarXivStatus
Adapt Your Body (proprio masking p=0.8)2506.23944page (withdrawn paper)
State-free policy2509.18644page
ReViP (state-dominant bias)2601.16667page
GAP (phase-guided gradient scaling)2602.12032page
ThinkProprio2602.06575page
Cloak (visual EE masking)2606.22836page

Grounding & conditioning placement (#11):

PaperarXivStatus
IVRA (patch-affinity injection)2601.16207page
FLOWER (deep-layer pruning)2509.04996page
SCALE (adaptive temperatures — banked title was wrong)2602.04208page
SmolVLA (mid-stack conditioning)2506.01844page

Data, tokenization & trunks (#5, #9, #16, #17, #18):

PaperarXivStatus
FAST (local canon)2501.09747page
FASTer (learned VQ tokenizer)2512.04952page
Rethinking VLA scaling (negative transfer)2602.09722page
Data-engine survey2604.23001page
VLM-to-VLA parameter redundancy2606.31382page
LoRA-r32 fine-tuning study (π0 on UR5e)2607.10172page

Vision-encoder freeze/unfreeze (#17, owner question 08-07):

PaperarXivStatus
OpenVLA (vision-FT ablation)2406.09246page
MAPS (module-wise proximity scheduling)2511.19878page
Dual-encoder representation preservation2509.11417page
VEGA (encoder grounding alignment)2605.10485page
HyperVLA (hypernetwork inference)2510.04898page
ActionX (RL expert pre-training)fnbot.2026.1806605page

Unfreezing schedules under a compute budget (owner steering 08-09 10:38Z, a(t)/b(t) framing):

PaperarXivStatus
LP-FT (feature distortion + two-phase schedule)2202.10054page
LP-FT mechanism via NTK (LLMs)2405.16747✅ covered in page
VLM4VLA (9-trunk sweep, module freezing, proxy collapse)2601.03309page

Smoothness / boundary family (fed by the 08-09 boundary-incompat read):

PaperarXivStatus
SEAM (inference-side seam steering)2607.04609page
FAFM (training-side frequency-space smoothness)2606.20135page

Data ingestion / heterogeneous collection (radar set 08-09):

PaperarXivStatus
VISTA (UMI adaptation: fisheye VQA + physics validation)2606.04708page
LAFP (latent-action flow policy)2606.10517page
Flowing With Purpose (latent-action FM)2606.23420✅ already covered: LAFM page (dup caught 08-09)

Fresh sweep 0810 (adamc readout + fjoint sequencing):

PaperarXivStatus
Correction of Decoupled Weight Decay (AdamC successor)2512.08217page
Z-1 (efficient GRPO for flow VLAs, selective joint training)2606.31846page

Radar 0811 (banked hooks from the 0810 fresh sweep):

PaperarXivStatus
TCFM (trajectory-consistent flow matching, RK4 decode)2605.08511page
RLDT (SVGD density-transport RL on flow policies)2606.08602page
FAN (feasible-action-neighborhood prior)2604.01570page
HiFlow (tokenization-free scale-wise AR-via-FM)2603.27281page
VLA-JEPA (latent world model)2602.10098page

Radar 0812b (banked hooks from the 0811 refill sweep):

PaperarXivStatus
VLA-Corrector (detect-and-correct inference, adaptive horizon)2607.01804page
π-StepNFT (step-wise negative-aware online RL for flow VLAs)2603.02083page
DFM-VLA (discrete flow matching iterative refinement)2603.26320page
One-Token-Per-Frame / OneWM-VLA (visual bandwidth in world models)2605.07931page
HiF-VLA (hindsight/insight/foresight motion representation)2512.09928page

Radar 0814 (banked hooks from the 0813 refill sweep):

PaperarXivStatus
Hyperball (Fantastic Pretraining Optimizers II, weight-norm equilibria)2606.16899page
Anytime Pretraining (horizon-free schedules + weight averaging)2602.03702page
VLA-FAIL (zero-failure-data detection: Mahalanobis + chunk consistency)2606.21386page
FPO (likelihood-free RFT of flow-matching VLAs, ICRA 2026)2510.09976page
X-Tokenizer (multimodal action tokenizer as auxiliary supervision)2606.14752page

Radar 0815 (banked hooks from the 0814 refill sweep):

PaperarXivStatus
Weight-norm criticality (loss spikes from decay+normalization driving scale-invariant norms below a critical floor)2607.21005page
Weibull weight-scale (three-force norm decomposition; spline recovery of alignment force from sparse checkpoints)2606.19367page
Decoupled Action Expert (5M MLP ≈ 244M U-Net; task knowledge fits in the conditioning pathway)2511.12101page
Foresight (learned failure detection over action-conditioned world-model latents, outcome labels only, conformal FPR band)2606.23085page
RedFlow (offline failure→correction RL for flow VLAs)2607.27782page
Weight decay improves LM plasticity (pretrain λ 0.5–1.0 beats 0.1 downstream; base loss under-predicts post-finetune quality)2602.11137page
Learning While Deploying (16-robot fleet offline-to-online RL; DIVL distributional critic + QAM flow-native extraction, frozen trunk)2605.00416page
FoMo-FD (inverse-transport nonconformity on a success-only flow world model; 96.6% detection @1.3% FA, wrist-cam-dependent)2607.27511page
VLA-GSE (spectral-init adapter-MoE from the frozen backbone’s SVD; init carries the gain, Gaussian-init lands below LoRA)2605.06175page
ActionCache (training-free retrieval cache over the flow decode; head-only speedups, trunk untouched — our bottleneck unaddressed)2607.06370page
MolmoAct2 (AI2 VLA on the Molmo2 trunk: Molmo2-ER backbone, 621M per-layer-KV flow expert, SO-100/101 checkpoint + curated 184h pool)2605.02881deep-dive post

Radar 0817 (banked hooks from the 0816 refill sweep; MolmoAct2 slot satisfied by the owner deep dive above):

PaperarXivStatus
ArmnetBench v0.1 (3-cell SO-101 arm farm; 2,518 human-scored rollouts over 7 policies × 12 tasks; 2,288 labeled failures released LeRobot-native)2607.24481page
SAFECAST (contrast-set rollouts for SAFE-style hidden-state failure probes; needs closed-loop re-executions + labeled failures — not offline; flow-policy cells below coin-flip)2608.04246page
Reflex (timestep-invariant trunk → exact KV reuse across denoising steps + async serving; 2.58× vs a soft baseline, stall 100%→0%)2607.14695page
Legato (guidance-aware flow objective makes chunk continuation native; −20% completion time vs RTC, smoothness ~flat)2602.12978page
Compression Gap (encoder gains propagate through continuous heads, blocked by an 80-bit FSQ codebook — tiny non-VLA models, single seed, mechanism asserted)2604.03191page

Radar 0818 (banked hooks from the 0817 refill sweep; every hook needed corrections again):

PaperarXivStatus
ATHENA (influence-function curation at π-0 3.3B scale — Kronecker projection + low-rank Hessian, 313× vs own dense baseline; rollout-anchored, 9.3h/6.9h corpora, no code)2606.16208page
ProbeAct (hook wrong both clauses: position regressor on 50k sim-oracle labels + hand-coded kinematic detection, zero detection metrics; trunk decodes position R²=0.968 while action head drifts)2606.09740page
Qwen-RobotManip (38,100h is ~65% re-rendered human video, ~7,800h real teleop; 5-stage offline state-action filter excluded 81% of RoboMIND-UR; nothing released)2606.17846page
Plasticity at scale (5M–314M LMs: scale delays, never prevents; onset T ∝ P^0.83; WD clause of the hook was a citation of 2602.11137; health proxies all fail to track onset)2606.24752page

Radar 0819 (banked hooks from the 0818 fresh sweep — new angles: sim2real for SO-class arms, action-space design, VLM-trunk continual learning, cross-embodiment; 14/16 candidates survived the local corpus grep, spares banked in the queue item):

PaperarXivStatus
Squint (SO-101 vendored into ManiSkill3 — NOT upstreamed — + MIT “SO-101 Task Set”, 8 envs, verified installable; single-task 16×16 wrist-cam visual SAC, 91.3% real vs 96.1% sim, ranking preserved; the rollout-substrate blocker is mechanically gone, visual world far-OOD so relative screens first)2602.21203page
Demystifying Action Space Design (single-arm AgileX, 13k rollouts, chunked flow policies included, code+data released and verified: chunk-wise delta-joint beats our absolute-joint cell +8.4pp, step-wise delta is the trap, execution-horizon interaction; cheapest justified arm = delta-joint retrain)2602.23408page
Benchmarking VLAs on SO-101 (320 real rollouts, 4 tasks × 4 policies × n=20; multi-label taxonomy despite its own single-label rule, execution labels saturate 91–100%; prize = 16 unlisted rollout_* LeRobot datasets on the author’s Hub account, unlabeled)2606.08881page
VLA continual-learning triangle (contradiction dissolves in the tables: all three show zero-replay sequential FT forgets catastrophically; episode replay ρ 0.02–0.2 @ 20% of batches fixes it at 3B full FT real-robot scale; “resistance” = better replay exchange rate from the VLM prior)2603.03818 + 2605.26820 + 2603.11653page

Radar 0820 (banked hooks from the 0819 fresh sweep — new angles: world-model/video pretraining, extra sensing on low-cost arms, imitation scaling laws, eval methodology; 14/16 candidates survived the corpus grep — the two dups were papers we’d already deep-read, one independently re-converged on our banked offline-validation page):

PaperarXivStatus
Rollout-free eval cluster: RoboWorld (r=0.989 vs RoboArena confirmed but n=8, no artifact released, GPT-4o judge never human-validated) + PolaRiS (r=0.9 over 24 policy-env points, MIT code live — but per-checkpoint co-training is load-bearing and calibration is DROID-only)2607.01060 + 2512.16881page
FACTR 2 (“no force sensor” hid a load-bearing 100 Hz current sensor; +17% bundles torque-as-observation with re-sampling, sampling-only never ablated; cheapest arm touched is a $2,500 Piper; but the load-bearing input Δq_d = action − state is free in our corpus)2606.12406page
Is Diversity All You Need? (“expert diversity hurts” was never operator-ablated — the evidence is the velocity-debias gain +15% ≈ 2.5× data, on a DIFFUSION action expert, so flow-head immunity is exactly what their setup contradicts; recipe unreleased; velocity spread is also an eval confound for chunk-MAE panels)2507.06219page
Emergence of human-to-robot transfer (π0.5+ego: human video ~doubles generalization but ONLY atop diverse robot pretraining; base-VLM init pays ~zero — we sit at the measured no-transfer corner; “threshold” partly our compression, no absolute units published; angle-A spares gated off)2512.22414page

Radar 0821 (banked hooks from the 0820 refill sweep — angles: eval methodology (rich again), imitation scaling laws, extra sensing (audio/current), data curation for robot corpora (new angle, hot); 16 candidates abs-verified by the sweep, 12 survived the corpus grep — the four casualties were all papers we had ALREADY deep-read (MolmoAct2, ArmnetBench, CI-MSE, Compression Gap), a sign the sweep is converging on our own reading list):

PaperarXivStatus
Quality over Quantity (influence curation anchored to 10–20 held-out demos, NOT rollouts — the offline pole we can compute against our panel; but every policy gain is on 40–50% author-injected failures, “per-episode weighting” was an overread — it’s hard top-N with strong budget sensitivity 36.7→86.7%; no code)2603.09056page
The Curse of Precision (log N ∝ 1/(P−c) confirmed as the model, R²>0.97 — but it’s a sim-only Franka-only FIT, tightest points extrapolated 23–65× beyond trained N; hook’s “not the task” wrong — low-randomization ablation moved c 2.35→1.00 mm; c needs rollout sweeps, so it’s a rig-phase instrument, not pre-computable)2607.23108page
NeuralActuator (cost floor broken: the third platform IS the SO-101 — force from Feetech load registers, no current sensor, torque via diffsim not calibration, MAE 0.47–0.73 N; hook’s “torque-from-current” wrong twice at our class; everything MIT-released incl. 3 SO-101 checkpoints + teleop code; corpus still can’t feed it — #9 gate stands, #16 rider shovel-ready)2607.11734page
GigaWorld-1 / WMBench (324K “rollouts” are human-graded world-model VIDEOS under replayed actions — no policy drives, real-robot ranking correlation defined but never reported; action-faithfulness>realism measured, partly definitional; big news the hook missed: full Apache-2.0 release of Nano 1.3B/Pro 5B + validated open VLM judge, and Ctrl-World is no longer the only released artifact)2607.02642page
SPARES (8, grep-clean 08-09): 2606.27375 ABC-130K open BC scaling substrate (3,500 h / 130K eps / 195 tasks + recipe sweeps); 2601.18723 Eval-Actions graded execution-quality labels (13K episodes, SRCC 0.81–0.84); 2603.13616 Beyond Binary Success anytime-valid sequential policy comparison (−70% eval burden); 2511.09958 Audio-VLA contact-mic template; 2512.08405 audio world models (flow-matching audio prediction); 2606.17598 MuseVLA frozen-trunk multimodal sensing; 2607.21588 AXIS community data engine (+5.8% from auto-QA); 2605.26349 episode-level teleop quality scoringbanked in queue

Radar 0822 (banked hooks from the 0821 refill sweep — angles: curation still the richest vein (a coherent curation-metrics testbed cluster surfaced), eval methodology keyword-rich/citation-thin, scaling medium, motor-current sensing nearly rested; 18 candidates checked, 15 abs-page-verified, 12 survived — the 3 dups were all papers we had already deep-read, the sweep keeps converging on our own reading list):

PaperarXivStatus
Ambient Diffusion Policy (MIT/Tedrake, RSS “It’s the demos” spotlight: suboptimal demos contribute only at high/low diffusion times, justified by a spectral power law in robot actions; up to +33% over naive co-training, purely offline — a #9 re-weighting lever on the flow/diffusion TIME axis, timestep-gated inclusion instead of hard filtering; check the spectral argument transfers to rectified flow + how the suboptimal split is designated)2606.12365hook banked
What Demonstration Curation Metrics Do to Your Policy (detection accuracy and policy quality sharply decoupled: best defect detector AUROC 0.804 → WORST curated policy 13.3%, weaker 0.638 detector nearly matches oracle 90.0 vs 93.3%; 5 of 7 metrics secretly exploit episode length — a direct confound warning for every #9 curation experiment and for our chunk-MAE panel; testbed released)2606.10229hook banked
Auditing Demonstration Curation Metrics (companion audit: action-only scorers catch noise/tremor/truncation but structural defects — wrong action at a key moment — are invisible to EVERY action-only metric, two actively prefer defective episodes; our positions-only corpus is exactly the failing feature space — the sharpest stress test for the label-free-selection-signals conclusions; check whether their “state” metrics need visual state)2606.05588hook banked
PhAIL (Franka FR3 open benchmark replacing binary-success-at-timeout with time-to-success CDFs: Human-Relative Throughput + bootstrap CIs + per-object KS tests, claims usable resolution at N≤30 rollouts/cell — THE statistical-protocol question for our rig-day rollout budget; needs human reference runs, check the resolution claim isn’t carried by the human anchor; dataset + reference implementation released)2605.29710hook banked
SPARES (8, abs-verified + grep-clean 08-10): 2607.04434 RoboDojo (42 sim + 18 real tasks, cloud real-eval, 30-policy leaderboard — the sim-vs-real alignment substrate); 2605.20774 VLA-REPLICA (low-cost reproducible real-world VLA bench — closest analogue to our #16 design); 2607.15330 Xiaomi-Robotics-1 (100K-h real-trajectory scaling report, checkpoints promised unconfirmed); 2606.15064 Phase-Localized Curation Does Not Help (negative result, same testbed family); 2606.20521 HumanScale (ego human video OUTPERFORMS robot data claim — candidate trigger for the h2r-lowdata-counterexample screen; check for a hidden diverse robot corpus in the alignment stage); 2603.05504 RoboPocket (AR visual-foresight targets demo collection at policy-weak regions, 2× data efficiency); 2606.30988 MuSe (post-hoc force-sensing attachment without forgetting — the #4/angle-B template); 2607.26047 S2A2 (spatial+spectral contact audio across ACT/DP/VQ-BeT/π0)banked in queue

Sim lane (owner directive 2026-08-11 17:07Z — sim + sim-to-real reading for the SO-101 eval substrate; the general lit pause stays for non-sim topics):

PagePapersFed
Sim-as-eval (SIMPLER MMRV 0.056/r 0.924 via sysid + green-screen, not photorealism; AutoEval’s 0/50-sim-vs-47/50-real caution on new policy families; SureSim paired-rectification CIs; continuous progress separates policies at up to 70% fewer trials; 2026 head-to-head: simulator choice moves Spearman 0.400↔0.700 on identical real evals)2405.05941, 2503.24278, 2510.04354, 2603.13616, 2606.10366, 2512.19562sim-policy-eval-100seeds protocol design; #16
SO-101 sim landscape (census: no public SO-101 sim eval with a continuous metric exists — 2026’s SO-101 benchmarks are all real-world; LeIsaac binary-success/Isaac-heavy, so101-nexus beta, so-frame’s REAL|SIM|OVERLAY worth stealing, LIBERO’s frozen init-states = the 100-seed pattern; our menagerie model’s kp 998 vs TheRobotStudio’s kp 17.8 for the same servo)2602.21203, 2512.19562, 2606.08881, 2607.24481, 2605.20774sim-policy-eval-100seeds; publish-later option (EnvHub)
Contact fidelity (all four sim-review findings have documented mechanisms + named fixes: CoACD threshold-not-cap / SDF escape hatch that also fixes the CC-BY-ND per-machine asset hazard; priority override is spec — explicit contact pair + condim 4 + elliptic cones; SIMPLER ablation: controller sysid first-order for MMRV, friction values second-order; BAM ships an identified STS3215 model)MuJoCo docs, 2205.02961, 2111.01391, 2410.08650, 2405.05941the pre-fix list ahead of the 100-seed pre-reg
Composite shadows (the paste-the-robot papers nearly all skip shadows and none measure them; ConCent’s recipe = random light + silhouette projection as randomization; Re³Sim ablation: foreground mesh→splat moves success 0.70→0.70 — scene, not foreground realism, binds; GreenAug-Rand beats generative backgrounds for training → randomize-in-training / match-in-eval split)2606.30268, 2503.14526, 2502.08645, 2407.07868composite-contact-shadows probe idea; sim-wrist-compositing design
Fisheye lens fitting (wrist fisheye 0.988 vs pinhole 0.181 real; policies overfit absolute pixel scale as a distance ruler — 220° lens transfer 0.0025→0.60 with Random Scale Augmentation; cubemap→equirect→any-lens MuJoCo pipeline removes our 72°-source ceiling and makes the real 130° module’s calibrated θ→r curve renderable)2603.02139fit-real-lens-model idea; v1 wrist render path upgrades
DR schedules (randomization width as success-throttled curriculum: DORAEMON entropy-max s.t. success ≥ α beats AutoDR 60% vs 26.7% real on 17-param Panda push; α=0.5 not 0.9 — the policy must be allowed to fail; one-scalar curriculum coefficient gets most of it; eval stays at the matched center, always)2311.01885, 1910.07113, 2505.05753, 2111.00956dr-schedule-for-sim-rl (conditional on GRPO probe); eval/train firewall rule