Lumen Research Digest — 2026-10-05
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. Previously featured work is excluded. Publications from the last 72 hours come first, with a strict seven-day maximum age.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars
- Source: arXiv
- Published: 2026-10-02T17:04:22+00:00
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for unusually strong scope and for operational use cases.
- Summary: We present ManifoldSplat, the first end-toend framework for language-guided semantic shape editing of animatable 3D Gaussian Splatting avatars reconstructed from monocular videos. We introduce DeltaRegion, a per-region disentangled Conditional Variational Autoencoder (CVAE) delivering feedforward shape deltas, alongside a refining stage to recover view-consistent details. ManifoldSplat is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.03599
- PDF: https://arxiv.org/pdf/2610.03599
2. EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures
- Source: arXiv
- Published: 2026-10-02T14:43:42+00:00
- Why it matters: Adds a stronger benchmark in agent workflows. Stands out for useful downstream control.
- Summary: We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. Title: EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures Base summary: Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with…. EdgeAgent is best read as a stronger benchmark in agent workflows.
- Link: https://arxiv.org/abs/2610.03394
- PDF: https://arxiv.org/pdf/2610.03394
3. HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
- Source: arXiv
- Published: 2026-10-02T16:59:19+00:00
- Why it matters: Adds a stronger benchmark in agent workflows. Stands out for unusually strong scope and credible evaluation pressure.
- Summary: To evaluate both the scientific outputs and the decisions that produce them, we introduce the Hazard Weaver Benchmark, comprising 141 instances across seven single-hazard domains and four multi-hazard interaction classes. Extensive experiments on this benchmark show that HazardWeaver outperforms existing agent systems, with the largest gains on tasks with multiple eligible scientific routes. HazardWeaver is best read as a stronger benchmark in agent workflows.
- Link: https://arxiv.org/abs/2610.03591
- PDF: https://arxiv.org/pdf/2610.03591
4. CORNAV: Construction-Aware Reasoning for Robot Navigation on Active Worksites
- Source: arXiv
- Published: 2026-10-02T17:19:05+00:00
- Why it matters: Adds an implementation framework in 3D and visual generation.
- Summary: We present CORNAV, a blueprint-grounded, schedule-aware navigation framework that operates from 2D CAD drawings and project schedules without requiring a Building Information Model. CORNAV aligns architectural blueprints against hierarchical open-vocabulary 3D scene graphs to ground object queries, converts project schedules into time-varying navigation constraints, and validates requests through an LLM-based safety module that…. CORNAV is best read as an implementation framework in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.03622
- PDF: https://arxiv.org/pdf/2610.03622
5. Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning
- Source: arXiv
- Published: 2026-10-02T10:19:49+00:00
- Why it matters: Adds an implementation framework in agent workflows. Stands out for unusually strong scope and credible evaluation pressure.
- Summary: We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. Beyond Single Videos is best read as an implementation framework in agent workflows.
- Link: https://arxiv.org/abs/2610.03099
- PDF: https://arxiv.org/pdf/2610.03099
6. OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection
- Source: arXiv
- Published: 2026-10-02T08:49:01+00:00
- Why it matters: Adds an implementation framework in 3D and visual generation.
- Summary: We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. OmniAct3D is best read as an implementation framework in 3D and visual generation.
- Link: https://arxiv.org/abs/2610.03015
- PDF: https://arxiv.org/pdf/2610.03015
7. WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
- Source: arXiv
- Published: 2026-10-02T09:18:28+00:00
- Why it matters: Adds a stronger benchmark in multimodal perception. Stands out for credible evaluation pressure.
- Summary: Title: WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites Base summary: We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final…. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. WebFovea is best read as a stronger benchmark in multimodal perception.
- Link: https://arxiv.org/abs/2610.03036
- PDF: https://arxiv.org/pdf/2610.03036
Coverage notes
- Candidates considered: 6075
- Sources: scientific papers from official arXiv new-paper announcements, with the arXiv API as fallback. Published dates are original submissions verified on official arXiv abstract pages, not announcement or revision dates. Revisions and company news are excluded.
- Selection policy: never repeat featured work; prefer the last 72 hours; exclude publications older than seven days or with unknown dates. Fewer qualifying items means a shorter digest.