Lumen Research Digest — 2026-10-03
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. Previously featured work is excluded. Publications from the last 72 hours come first, with a strict seven-day maximum age.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
- Source: arXiv
- Published: 2026-10-01T17:59:50+00:00
- Why it matters: Adds a stronger benchmark in multimodal perception. Stands out for unusually strong scope.
- Summary: We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Guided Self-Improvement Embodied Agents is best read as a stronger benchmark in multimodal perception.
- Link: https://arxiv.org/abs/2610.02204
- PDF: https://arxiv.org/pdf/2610.02204
2. CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding
- Source: arXiv
- Published: 2026-09-30T15:12:10+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation.
- Summary: We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. Title: CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding Base summary: Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and…. CoVisco is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2609.39924
- PDF: https://arxiv.org/pdf/2609.39924
3. TouchTherm: Building Multimodal Digital Twins of Objects for Tactile and Thermal Rendering
- Source: arXiv
- Published: 2026-10-01T16:10:08+00:00
- Why it matters: Adds an implementation framework in 3D and visual generation.
- Summary: We present TouchTherm, a framework for constructing simulation-ready visuo-tactile-thermal object assets from real-world objects. Experiments on 20 objects show that the reconstructed micro-height fields preserve dominant surface structures and recover higher-frequency details beyond the coarse geometry, while the thermal fields achieve held-out surface-temperature MAEs of 0.465…. TouchTherm is best read as an implementation framework in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.01943
- PDF: https://arxiv.org/pdf/2610.01943
4. GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking
- Source: arXiv
- Published: 2026-10-01T14:20:38+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation.
- Summary: In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm. GenCOPE is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.01758
- PDF: https://arxiv.org/pdf/2610.01758
5. MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation
- Source: arXiv
- Published: 2026-10-01T13:52:12+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation.
- Summary: To overcome these limitations, we propose MEGA ( M esh E xtraction from GA ussians), a ``segment-then-mesh'' framework for extracting object-level, watertight meshes from complex 3DGS scenes. SVD treats the 3DGS model as a teacher, sampling diverse camera poses and rendering the corresponding views of each segmented object. MEGA is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.01707
- PDF: https://arxiv.org/pdf/2610.01707
6. Lens Flare Removal and Reconstruction
- Source: arXiv
- Published: 2026-09-30T11:27:28+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for credible evaluation pressure.
- Summary: We evaluate removal on an established benchmark and a new one for large reflective flares, quantify the flare/scene decomposition directly, and show that the pipeline is robust to errors in automatic light-source localization. To achieve this, we introduce a flare representation model that leverages the symmetry of lens flares about the camera's principal point. Lens Flare Removal Reconstruction is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2609.39527
- PDF: https://arxiv.org/pdf/2609.39527
7. FutureWorlds: Learning Robotic World Models from Alternative Futures
- Source: arXiv
- Published: 2026-10-01T04:04:31+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation.
- Summary: Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. FutureWorlds is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.01019
- PDF: https://arxiv.org/pdf/2610.01019
Coverage notes
- Candidates considered: 5447
- Sources: scientific papers from official arXiv new-paper announcements, with the arXiv API as fallback. Published dates are original submissions verified on official arXiv abstract pages, not announcement or revision dates. Revisions and company news are excluded.
- Selection policy: never repeat featured work; prefer the last 72 hours; exclude publications older than seven days or with unknown dates. Fewer qualifying items means a shorter digest.