Lumen Research Digest — 2026-05-09
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. This is ranked for novelty and likely significance rather than simply recency.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. MedHorizon: Towards Long-context Medical Video Understanding in the Wild
- Source: arXiv
- Published: 2026-05-07T16:37:10Z
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for credible evaluation pressure.
- Summary: We introduce MedHorizon, an in-the-wild benchmark for long-context medical video understanding. MedHorizon preserves 759 hours of full-length clinical procedures and provides 1,253 evidence-grounded multiple-choice questionsthat jointly evaluate sparse evidence understanding and multi-hop clinical reasoning. MedHorizon is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2605.06537v1
- PDF: https://arxiv.org/pdf/2605.06537v1
2. Running Codex safely at OpenAI
- Source: OpenAI
- Published: Fri, 08 May 2026 12:30:00 GMT
- Why it matters: Worth tracking as OpenAI pushes on agent workflows via a concrete technical advance.
- Summary: Title: Running Codex safely at OpenAI Base summary: How OpenAI runs Codex securely with sandboxing, approvals, network policies, and agent-native telemetry to support safe and compliant coding agent adoption. Security teams need ways to govern how agents operate: what they can access, when human approval is required, which systems they can interact with, and what telemetry exists to explain their behavior. Running Codex safely OpenAI is best read as a concrete technical advance in agent workflows.
- Link: https://openai.com/index/running-codex-safely
3. Building realistic electric transmission grid dataset at scale: a pipeline from open dataset
- Source: Microsoft Research
- Published: Fri, 08 May 2026 19:53:56 +0000
- Why it matters: Worth tracking as Microsoft Research pushes on systems efficiency via an implementation framework.
- Summary: Analyses of congestion, transmission expansion, demand growth, and system resilience all depend on network models with realistic Page title: Building realistic electric transmission grid dataset at scale: a pipeline from open dataset - Microsoft Research…. Title: Building realistic electric transmission grid dataset at scale: a pipeline from open dataset Base summary: Microsoft Research is excited to release an open dataset of approximate transmission topology of the U.S. power grid derived from publicly…. pipeline open dataset is best read as an implementation framework in systems efficiency.
- Link: https://www.microsoft.com/en-us/research/blog/building-realistic-electric-transmission-grid-dataset-at-scale-a-pipeline-from-open-dataset/
4. ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
- Source: arXiv
- Published: 2026-05-07T17:59:58Z
- Why it matters: Adds a stronger benchmark in 3D and visual generation. Stands out for credible evaluation pressure.
- Summary: We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. ActCam is best read as a stronger benchmark in 3D and visual generation.
- Link: https://arxiv.org/abs/2605.06667v1
- PDF: https://arxiv.org/pdf/2605.06667v1
5. BAMI: Training-Free Bias Mitigation in GUI Grounding
- Source: arXiv
- Published: 2026-05-07T17:59:31Z
- Why it matters: Adds a stronger benchmark in developer tooling. Stands out for credible evaluation pressure.
- Summary: However, in complex scenarios like the ScreenSpot-Pro benchmark, existing models often suffer from suboptimal performance. For instance, applying our method to the TianXi-Action-7B model boosts its accuracy on the ScreenSpot-Pro benchmark from 51.9\% to 57.8\%. BAMI is best read as a stronger benchmark in developer tooling.
- Link: https://arxiv.org/abs/2605.06664v1
- PDF: https://arxiv.org/pdf/2605.06664v1
Coverage notes
- Candidates considered: 67
- Sources included: arXiv topic queries plus selected research/lab/blog feeds.
- Selection policy: novelty, likely downstream importance, technical substance, and recent coverage avoidance.