Lumen Research Digest — 2026-09-25
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. Previously featured work is excluded. Publications from the last 72 hours come first, with a strict seven-day maximum age.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
- Source: arXiv
- Published: 2026-09-24T15:19:39+00:00
- Why it matters: Adds better debugging hooks in 3D and visual generation.
- Summary: We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. World Action Agent is best read as better debugging hooks in 3D and visual generation.
- Link: https://arxiv.org/abs/2609.29964
- PDF: https://arxiv.org/pdf/2609.29964
2. Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation
- Source: arXiv
- Published: 2026-09-24T15:01:24+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for unusually strong scope and credible evaluation pressure.
- Summary: Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation. Guided by these findings, we build Spatial-Nav-100K and fine-tune in two stages, i.e. first learning a shared spatial-navigation foundation, and then specializing each phase with the abilities it relies on. Beyond Spatial Benchmarks is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2609.29934
- PDF: https://arxiv.org/pdf/2609.29934
3. SWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering
- Source: arXiv
- Published: 2026-09-24T13:04:53+00:00
- Why it matters: Adds a stronger benchmark in agent workflows. Stands out for credible evaluation pressure.
- Summary: We present SWE-PolyVision, an executable benchmark of 92 real tasks from 36 open-source organizations, with 48 public tasks and 44 private holdouts. Two trace-linked Native Vision cases illustrate how complementary visual and textual clues can lead to source-localized, verified repairs; controlled interventions show that this conversion is not yet stable across inputs. SWE-PolyVision is best read as a stronger benchmark in agent workflows.
- Link: https://arxiv.org/abs/2609.29754
- PDF: https://arxiv.org/pdf/2609.29754
4. From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments
- Source: arXiv
- Published: 2026-09-24T06:19:04+00:00
- Why it matters: Adds an implementation framework in robotics and embodied perception. Stands out for credible evaluation pressure.
- Summary: To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions. We evaluate our method on a realistic Find-and-Place task, demonstrating its effectiveness in challenging environments where target objects must be actively discovered before manipulation. Agentic Embodied Manipulation Realistic Environments is best read as an implementation framework in robotics and embodied perception.
- Link: https://arxiv.org/abs/2609.29091
- PDF: https://arxiv.org/pdf/2609.29091
5. BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines
- Source: arXiv
- Published: 2026-09-23T08:55:17+00:00
- Why it matters: Adds an implementation framework in agent workflows. Stands out for for operational use cases.
- Summary: Title: BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines Base summary: DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems…. The framework decomposes the pipeline into six specialized AI agents, covering sample intake and quality control, alignment, variant calling, annotation, cross-stage monitoring, and reporting. BaseCamp --- Agentic AI Framework is best read as an implementation framework in agent workflows.
- Link: https://arxiv.org/abs/2609.28557
- PDF: https://arxiv.org/pdf/2609.28557
6. HarnessPAI: An Evolving Harness for Physical AI
- Source: arXiv
- Published: 2026-09-24T07:40:00+00:00
- Why it matters: Adds an implementation framework in robotics and embodied perception.
- Summary: We introduce HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface that organizes the underlying action primitive. Title: HarnessPAI: An Evolving Harness for Physical AI Base summary: Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. HarnessPAI is best read as an implementation framework in robotics and embodied perception.
- Link: https://arxiv.org/abs/2609.29166
- PDF: https://arxiv.org/pdf/2609.29166
7. Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning
- Source: arXiv
- Published: 2026-09-23T23:43:26+00:00
- Why it matters: Adds a stronger benchmark in developer tooling. Stands out for credible evaluation pressure.
- Summary: To evaluate robustness beyond accuracy or consistency on fixed test sets, we introduce VisFlip, a benchmark framework that constructs evaluation groups for a target model and visual variation setting to separately assess recovery of original predictions and…. VisFlip spans nine dataset-variation combinations across scientific reasoning, robot-scene understanding, and medical VQA, covering subtle visual variations common in each domain. Flip-Direction Steering Robust Vision-Language Reasoning is best read as a stronger benchmark in developer tooling.
- Link: https://arxiv.org/abs/2609.28851
- PDF: https://arxiv.org/pdf/2609.28851
Coverage notes
- Candidates considered: 616
- Sources: scientific papers from official arXiv new-paper announcements, with the arXiv API as fallback. Published dates are original submissions verified on official arXiv abstract pages, not announcement or revision dates. Revisions and company news are excluded.
- Selection policy: never repeat featured work; prefer the last 72 hours; exclude publications older than seven days or with unknown dates. Fewer qualifying items means a shorter digest.