The easiest way to read a daily research digest is as a stack of disconnected papers. That is usually the least useful way to read it. The better move is to look for the technical directions that keep surfacing, the problems researchers are taking more seriously, and the kinds of systems that look increasingly deployable.

This brief is a synthesis of the digest rather than a direct dump of every item. The goal is to surface what matters for people building AI systems, workflow automation, internal assistants, and production infrastructure.

Where the structure showed up

The strongest signal in this digest is that multimodal work is becoming harder to separate from the orchestration layers around it. More of the useful progress is happening in the interfaces between perception, reasoning, tool use, and evaluation.

That matters because production systems are rarely judged on one capability in isolation. They are judged on whether the surrounding control surface turns model ability into repeatable behavior.

What builders should pay attention to

For teams shipping internal assistants or workflow systems, the practical gain is not just richer inputs. It is better system structure: clearer execution steps, tighter observation loops, and fewer hidden assumptions.

That points toward products that are narrower, better instrumented, and more explicit about how they operate when the environment gets messy.

Paper summaries

Below are the individual papers and a fuller summary of what each one is doing, what looks new, and why it may matter, followed by direct source links.

1. ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars

We present ManifoldSplat, the first end-toend framework for language-guided semantic shape editing of animatable 3D Gaussian Splatting avatars reconstructed from monocular videos. We introduce DeltaRegion, a per-region disentangled Conditional Variational Autoencoder (CVAE) delivering feedforward shape deltas, alongside a refining stage to recover view-consistent details. ManifoldSplat is best read as a stronger benchmark in 3D and visual generation.

Source link →

2. EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures

We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. Title: EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures Base summary: Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with…. EdgeAgent is best read as a stronger benchmark in agent workflows.

Source link →

3. HazardWeaver: Scientific Route Selection for Hazard Analysis Agents

To evaluate both the scientific outputs and the decisions that produce them, we introduce the Hazard Weaver Benchmark, comprising 141 instances across seven single-hazard domains and four multi-hazard interaction classes. Extensive experiments on this benchmark show that HazardWeaver outperforms existing agent systems, with the largest gains on tasks with multiple eligible scientific routes. HazardWeaver is best read as a stronger benchmark in agent workflows.

Source link →

4. CORNAV: Construction-Aware Reasoning for Robot Navigation on Active Worksites

We present CORNAV, a blueprint-grounded, schedule-aware navigation framework that operates from 2D CAD drawings and project schedules without requiring a Building Information Model. CORNAV aligns architectural blueprints against hierarchical open-vocabulary 3D scene graphs to ground object queries, converts project schedules into time-varying navigation constraints, and validates requests through an LLM-based safety module that…. CORNAV is best read as an implementation framework in 3D and visual generation.

Source link →

5. Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

We introduce AdsCVR, the first e-commerce cross-video reasoning benchmark, containing 2,483 videos and 6,110 question-answer pairs across six reasoning dimensions. We therefore propose AdSeek, an agentic framework that dynamically selects visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence acquisition. Beyond Single Videos is best read as an implementation framework in agent workflows.

Source link →

6. OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. OmniAct3D is best read as an implementation framework in 3D and visual generation.

Source link →

7. WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites

Title: WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites Base summary: We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final…. The challenge evaluates agents end to end on Protocol III of the WebRetriever benchmark (arXiv:2607.06118): starting from an entry URL on a live website, the agent must operate the site's own interface and return a verifiable answer. WebFovea is best read as a stronger benchmark in multimodal perception.

Source link →

References