The easiest way to read a daily research digest is as a stack of disconnected papers. That is usually the least useful way to read it. The better move is to look for the technical directions that keep surfacing, the problems researchers are taking more seriously, and the kinds of systems that look increasingly deployable.
This brief is a synthesis of the digest rather than a direct dump of every item. The goal is to surface what matters for people building AI systems, workflow automation, internal assistants, and production infrastructure.
Where the structure showed up
The strongest signal in this digest is that multimodal work is becoming harder to separate from the orchestration layers around it. More of the useful progress is happening in the interfaces between perception, reasoning, tool use, and evaluation.
That matters because production systems are rarely judged on one capability in isolation. They are judged on whether the surrounding control surface turns model ability into repeatable behavior.
What builders should pay attention to
For teams shipping internal assistants or workflow systems, the practical gain is not just richer inputs. It is better system structure: clearer execution steps, tighter observation loops, and fewer hidden assumptions.
That points toward products that are narrower, better instrumented, and more explicit about how they operate when the environment gets messy.
Paper summaries
Below are the individual papers and a fuller summary of what each one is doing, what looks new, and why it may matter, followed by direct source links.
1. CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation
We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning, demonstrations, external specialist perception or grasp modules, privileged scene state, or predefined task policies, agents should select visual evidence, form task-relevant 3D estimates, construct manipulation targets,…. CodeActionBench is best read as a stronger benchmark in 3D and visual generation.
2. Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos
We introduce Beyond3D, the first VQA benchmark to isolate this ability in dynamic egocentric video: every query targets an object that has been relocated and has since left the field of view. We benchmark nine general-purpose and spatially specialized VLMs. Benchmarking VLMs Out-of-Sight Spatiotemporal Reasoning is best read as a stronger benchmark in 3D and visual generation.
3. Recent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent Intelligence
Building on this analysis, we introduce a conceptual PhenoAgent framework that extends phenotyping beyond the estimation of isolated traits to evidence-based crop-state interpretation, uncertainty-aware reasoning, and management-oriented support. We also discuss challenges in dataset scarcity, annotation, benchmarking, model generalization, and explainability. Perspective Review Fragmented Multimodal Sensing is best read as new data infrastructure in agent workflows.
4. Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors
For long-range dependencies among unstructured anchors, we further introduce fixed-size memory slots with orthogonality-informed updates for accurate entropy-context modeling. Experiments show that SAGA achieves strong rate-distortion performance against GIFStream, with PSNR BD-rate reductions of 80.39% and 83.94% on Neu3D and MPEG MIV, respectively. Gaussian Splatting-based Volumetric Video Compression is best read as a concrete technical advance in 3D and visual generation.
5. Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment. To bridge high-level planning and real-world execution, we introduce Gaussian-aware final-state estimation through geometric sampling and state-based trajectory planning. Robot-GST is best read as a stronger benchmark in 3D and visual generation.
6. JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests…. Title: JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments Base summary: In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question…. JRDB-AVR is best read as a stronger benchmark in robotics and embodied perception.
7. GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing. GenNVS is best read as an implementation framework in 3D and visual generation.
References
- CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation
- Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos
- Recent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent Intelligence
- Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors
- Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
- JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
- GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior