Lumen Research Digest — 2026-10-04
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. Previously featured work is excluded. Publications from the last 72 hours come first, with a strict seven-day maximum age.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
- Source: arXiv
- Published: 2026-10-01T17:35:10+00:00
- Why it matters: Adds better debugging hooks in 3D and visual generation.
- Summary: We introduce SkeleWAM, a compact WAM that represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. SkeleWAM is best read as better debugging hooks in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.02120
- PDF: https://arxiv.org/pdf/2610.02120
2. LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction
- Source: arXiv
- Published: 2026-10-01T15:25:35+00:00
- Why it matters: Adds an implementation framework in 3D and visual generation.
- Summary: Title: LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction Base summary: We present LiteReality-Agent, an agentic system for reconstructing real indoor environments as realistic, articulated, and simulation-ready 3D scenes…. Furthermore, as agent capabilities continue to improve rapidly, the system introduced by LiteReality-Agent remains a strong orchestration framework for future agents: it equips them with specialised tools, structured workflows, and robust verification…. LiteReality-Agent is best read as an implementation framework in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.01863
- PDF: https://arxiv.org/pdf/2610.01863
3. World Observer: Joint Actor-Observer Generation for Persistent World Modeling
- Source: arXiv
- Published: 2026-10-01T17:53:20+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for credible evaluation pressure.
- Summary: To evaluate out-of-view evolution, we further introduce world-space metrics and a benchmark spanning real and synthetic scenes. We ground the actor and observers by warping from a shared panoramic source for explicit geometric correspondence, and introduce an Observer Sink of high-resolution perspective references to restore fine appearance upon re-entry. World Observer is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.02162
- PDF: https://arxiv.org/pdf/2610.02162
4. Code Owns the Simulation, Jev Owns the Evaluation
- Source: arXiv
- Published: 2026-10-01T15:06:40+00:00
- Why it matters: Adds a stronger benchmark in agent workflows. Stands out for unusually strong scope and credible evaluation pressure.
- Summary: It fails when one call must both perform the simulation and evaluate based on it. However, it fails when the right option depends on simulation (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. Code Owns Simulation Jev Owns is best read as a stronger benchmark in agent workflows.
- Link: https://arxiv.org/abs/2610.01834
- PDF: https://arxiv.org/pdf/2610.01834
5. Continual Learning for 6-DoF Grasp Synthesis via Experience and Demonstrations
- Source: arXiv
- Published: 2026-10-01T08:36:06+00:00
- Why it matters: Adds an implementation framework in 3D and visual generation. Stands out for credible evaluation pressure.
- Summary: In this work, we present a continual-learning framework for single-view 6-DoF grasp synthesis for a parallel-jaw gripper in cluttered scenes. We show that our method matches the performance of existing 6-DoF grasping baselines even before adaptation, improves online on unseen objects from categories absent or underrepresented during training, and supports long-horizon continual learning with…. Continual Learning 6-DoF Grasp Synthesis is best read as an implementation framework in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.01301
- PDF: https://arxiv.org/pdf/2610.01301
6. Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
- Source: arXiv
- Published: 2026-10-01T12:45:01+00:00
- Why it matters: Adds a stronger benchmark in developer tooling.
- Summary: Title: Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs Base summary: Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet…. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. Before It Fades is best read as a stronger benchmark in developer tooling.
- Link: https://arxiv.org/abs/2610.01595
- PDF: https://arxiv.org/pdf/2610.01595
7. PAGER: Partial-to-global Alignment via Geometric and Relational Distillation
- Source: arXiv
- Published: 2026-10-01T12:42:20+00:00
- Why it matters: Adds new data infrastructure in 3D and visual generation.
- Summary: We show that this shift from globally learned 3D feature spaces to realistic partial observations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. We introduce PAGER, a label-free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. PAGER is best read as new data infrastructure in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.01589
- PDF: https://arxiv.org/pdf/2610.01589
Coverage notes
- Candidates considered: 5447
- Sources: scientific papers from official arXiv new-paper announcements, with the arXiv API as fallback. Published dates are original submissions verified on official arXiv abstract pages, not announcement or revision dates. Revisions and company news are excluded.
- Selection policy: never repeat featured work; prefer the last 72 hours; exclude publications older than seven days or with unknown dates. Fewer qualifying items means a shorter digest.