Lumen Research Digest — 2026-10-06
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. Previously featured work is excluded. Publications from the last 72 hours come first, with a strict seven-day maximum age.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting
- Source: arXiv
- Published: 2026-10-04T15:07:47+00:00
- Why it matters: Adds an implementation framework in 3D and visual generation. Stands out for unusually strong scope and useful downstream control.
- Summary: We present Mobile-4DGS, a unified lightweight framework for high-fidelity real-time static and dynamic Gaussian rendering on mobile platforms. For compact appearance modeling, we introduce a Monte Carlo Specular Energy Aggregator that compresses high-order radiance residuals into the first-order Spherical Harmonics (SH), together with an Attribute-Conditioned SH Enhancement module whose predicted…. Mobile-4DGS is best read as an implementation framework in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.05289
- PDF: https://arxiv.org/pdf/2610.05289
2. VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding
- Source: arXiv
- Published: 2026-10-05T16:44:25+00:00
- Why it matters: Adds an implementation framework in agent workflows.
- Summary: We introduce VideoTapestry, a training-free multi-agent framework that adapts a preconstructed hierarchical video memory through coarse-to-fine, query-driven refinement. Title: VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding Base summary: Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across…. VideoTapestry is best read as an implementation framework in agent workflows.
- Link: https://arxiv.org/abs/2610.06672
- PDF: https://arxiv.org/pdf/2610.06672
3. Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants
- Source: arXiv
- Published: 2026-10-04T20:42:29+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for credible evaluation pressure.
- Summary: We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. Have I Scene Before Spatially is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.05526
- PDF: https://arxiv.org/pdf/2610.05526
4. Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos
- Source: arXiv
- Published: 2026-10-03T10:46:52+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for useful downstream control and credible evaluation pressure.
- Summary: To evaluate this capability, we introduce Video2World, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Title: Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos Base summary: Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily…. Video2World is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.04432
- PDF: https://arxiv.org/pdf/2610.04432
5. BrainTRACE: Tracing Longitudinal, Multimodal, and Volumetric Evidence in Brain MRI Clinical Reasoning
- Source: arXiv
- Published: 2026-10-05T15:52:37+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for credible evaluation pressure and for operational use cases.
- Summary: The benchmark is organized by five levels of clinical reasoning, from acquisition recognition to case-level synthesis, and by evidence demands covering longitudinal comparison, report-grounded references, multi-sequence integration, and volumetric spatial…. We introduce BrainTRACE, a report-grounded benchmark for evaluating whether vision-language models can trace the evidence structure required for longitudinal brain MRI interpretation. BrainTRACE is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.06571
- PDF: https://arxiv.org/pdf/2610.06571
6. Code2Games: Enabling Coding Agents for Gaming World Generation
- Source: arXiv
- Published: 2026-10-04T07:54:45+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for useful downstream control and credible evaluation pressure.
- Summary: To systematically evaluate gaming-world generation, we introduce the GameCode4D benchmark, which comprises ten fixed game prompts spanning different levels of scene and gameplay complexity. We propose Code2Games, an agentic framework that builds a structured gaming world upon a base Blender world generated from the same game intent. Code2Games is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.05033
- PDF: https://arxiv.org/pdf/2610.05033
7. ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image
- Source: arXiv
- Published: 2026-10-04T14:26:34+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for useful downstream control.
- Summary: For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from…. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. ArticuTable is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.05249
- PDF: https://arxiv.org/pdf/2610.05249
Coverage notes
- Candidates considered: 6818
- Sources: scientific papers from official arXiv new-paper announcements, with the arXiv API as fallback. Published dates are original submissions verified on official arXiv abstract pages, not announcement or revision dates. Revisions and company news are excluded.
- Selection policy: never repeat featured work; prefer the last 72 hours; exclude publications older than seven days or with unknown dates. Fewer qualifying items means a shorter digest.