Lumen Research Digest — 2026-04-25
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. This is ranked for novelty and likely significance rather than simply recency.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. Task-Driven Co-Design of Heterogeneous Multi-Robot Systems
- Source: arXiv
- Published: 2026-04-23T17:44:52Z
- Why it matters: Adds an implementation framework in agent workflows.
- Summary: In this work, we present a formal and compositional framework for the task-driven co-design of heterogeneous multi-robot systems. Building on a monotone co-design theory, we introduce general abstractions of robots, fleets, planners, executors, and evaluators as interconnected design problems with well-defined interfaces that are agnostic to both implementations and tasks. Task-Driven Co-Design Heterogeneous Multi-Robot Systems is best read as an implementation framework in agent workflows.
- Link: https://arxiv.org/abs/2604.21894v1
- PDF: https://arxiv.org/pdf/2604.21894v1
2. Plugins and skills
- Source: OpenAI
- Published: Thu, 23 Apr 2026 10:00:00 GMT
- Why it matters: Worth tracking as OpenAI pushes on agent workflows via a concrete technical advance.
- Summary: Title: Plugins and skills Base summary: Learn how to use Codex plugins and skills to connect tools, access data, and follow repeatable workflows to automate tasks and improve results. For example, a plugin might help Codex reference files in Google Drive, scan your email inbox, or work with information from another tool you use. Plugins skills is best read as a concrete technical advance in agent workflows.
- Link: https://openai.com/academy/codex-plugins-and-skills
3. AsgardBench: A benchmark for visually grounded interactive planning
- Source: Microsoft Research
- Published: Thu, 26 Mar 2026 19:02:53 +0000
- Why it matters: Worth tracking as Microsoft Research pushes on robotics and embodied perception via a stronger benchmark. Stands out for useful downstream control and credible evaluation pressure.
- Summary: This is the domain of embodied AI: systems Page title: AsgardBench: A benchmark for visually grounded interactive planning - Microsoft Research Page extract: AsgardBench evaluates whether embodied agents can revise their plans based on visual observations as…. Title: AsgardBench: A benchmark for visually grounded interactive planning Base summary: Imagine a robot tasked with cleaning a kitchen. AsgardBench is best read as a stronger benchmark in robotics and embodied perception.
- Link: https://www.microsoft.com/en-us/research/blog/asgardbench-a-benchmark-for-visually-grounded-interactive-planning/
4. Long-Horizon Manipulation via Trace-Conditioned VLA Planning
- Source: arXiv
- Published: 2026-04-23T17:59:04Z
- Why it matters: Adds better debugging hooks in robotics and embodied perception.
- Summary: We present LoHo-Manip, a modular framework that scales short-horizon VLA execution to long-horizon instruction following via a dedicated task-management VLM. The manager is decoupled from the executor and is invoked in a receding-horizon manner: given the current observation, it predicts a progress-aware remaining plan that combines (i) a subtask sequence with an explicit done + remaining split as lightweight…. Long-Horizon Manipulation via Trace-Conditioned VLA is best read as better debugging hooks in robotics and embodied perception.
- Link: https://arxiv.org/abs/2604.21924v1
- PDF: https://arxiv.org/pdf/2604.21924v1
5. Seeing Fast and Slow: Learning the Flow of Time in Videos
- Source: arXiv
- Published: 2026-04-23T17:59:57Z
- Why it matters: Adds new data infrastructure in multimodal perception. Stands out for unusually strong scope and useful downstream control.
- Summary: We then show that these learned temporal reasoning models enable us to curate the largest slow-motion video dataset to date from noisy in-the-wild sources. We first exploit the multimodal cues and temporal structure naturally present in videos to learn, in a self-supervised manner, to detect speed changes and estimate playback speed. Learning Flow Time Videos is best read as new data infrastructure in multimodal perception.
- Link: https://arxiv.org/abs/2604.21931v1
- PDF: https://arxiv.org/pdf/2604.21931v1
6. Working with Codex
- Source: OpenAI
- Published: Thu, 23 Apr 2026 10:00:00 GMT
- Why it matters: Worth tracking as OpenAI pushes on developer tooling via a concrete technical advance. Stands out for useful downstream control.
- Summary: Title: Working with Codex Base summary: Learn how to set up your Codex workspace, create threads and projects, manage files, and start completing tasks with step-by-step guidance. When you open Codex, you’ll see a few core elements: a sidebar menu, projects, settings, and a chat window. Working Codex is best read as a concrete technical advance in developer tooling.
- Link: https://openai.com/academy/working-with-codex
Coverage notes
- Candidates considered: 71
- Sources included: arXiv topic queries plus selected research/lab/blog feeds.
- Selection policy: novelty, likely downstream importance, technical substance, and recent coverage avoidance.