Research Digest Archive
Plain static pages generated from the workspace digest markdown files. This directory is self-contained and can be published as-is.
Lumen Research Digest — 2026-10-09
2026-10-09 · 7 items
- 1. FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs
- 2. GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images
- 3. SuperNav: An Agentic Navigation System for Any Task in Any Scene
- 4. SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
- 5. 4-Tensor Attention Model for Semantic Physical Reality
- 6. VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding
- 7. Rendering-Free Lookahead for Question-Guided Active Vision
Lumen Research Digest — 2026-10-08
2026-10-08 · 7 items
- 1. TileSkipper: Region-Adaptive Tile Pruning for 3D Gaussian Splatting
- 2. SPLATIFY: Reproduce, Discover, Innovate! From Papers and Ideas to Trainable 3DGS Code
- 3. DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting
- 4. vLLM-Omni Technical Report: A Unified Serving Runtime for Omni-Modality Generation
- 5. Context-aware Attention-based Gaussian Mixture Models for Vehicular Trajectory Prediction
- 6. SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models
- 7. Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
Lumen Research Digest — 2026-10-07
2026-10-07 · 7 items
- 1. MoonGS: High-quality Representation of the Lunar Surface via Gaussian Splatting Using Robust Depth Features from Image Pairs
- 2. OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception
- 3. DepthWorld: 3D World Model for Robot Manipulation
- 4. WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
- 5. Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering
- 6. DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given
- 7. Robotizing Human Videos with Physically Consistent Interactions
Lumen Research Digest — 2026-10-06
2026-10-06 · 7 items
- 1. Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting
- 2. VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding
- 3. Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants
- 4. Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos
- 5. BrainTRACE: Tracing Longitudinal, Multimodal, and Volumetric Evidence in Brain MRI Clinical Reasoning
- 6. Code2Games: Enabling Coding Agents for Gaming World Generation
- 7. ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image
Lumen Research Digest — 2026-10-05
2026-10-05 · 7 items
- 1. ManifoldSplat: Language-Guided Semantic Shape Editing of 3D Gaussian Head Avatars
- 2. EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures
- 3. HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
- 4. CORNAV: Construction-Aware Reasoning for Robot Navigation on Active Worksites
- 5. Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning
- 6. OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection
- 7. WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites
Lumen Research Digest — 2026-10-04
2026-10-04 · 7 items
- 1. SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
- 2. LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction
- 3. World Observer: Joint Actor-Observer Generation for Persistent World Modeling
- 4. Code Owns the Simulation, Jev Owns the Evaluation
- 5. Continual Learning for 6-DoF Grasp Synthesis via Experience and Demonstrations
- 6. Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
- 7. PAGER: Partial-to-global Alignment via Geometric and Relational Distillation
Lumen Research Digest — 2026-10-03
2026-10-03 · 7 items
- 1. Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
- 2. CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding
- 3. TouchTherm: Building Multimodal Digital Twins of Objects for Tactile and Thermal Rendering
- 4. GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking
- 5. MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation
- 6. Lens Flare Removal and Reconstruction
- 7. FutureWorlds: Learning Robotic World Models from Alternative Futures
Lumen Research Digest — 2026-10-02
2026-10-02 · 7 items
- 1. DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
- 2. Are Frontier VLM Agents Ready to Be Robot Generalists? An Empirical Study with the Embodied Agent Arena
- 3. MemFit: Efficient Long-Term Agentic Memory
- 4. CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight
- 5. Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents
- 6. Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
- 7. VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
Lumen Research Digest — 2026-10-01
2026-10-01 · 7 items
- 1. WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents
- 2. MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
- 3. LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
- 4. EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
- 5. STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
- 6. RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance
- 7. Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
Lumen Research Digest — 2026-09-30
2026-09-30 · 7 items
- 1. Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
- 2. Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding
- 3. Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
- 4. DispFlow-GS: Displacement Flow Supervision with Motion Disentangling for Monocular Deformable 3D Gaussian Splatting
- 5. Distilling Privileged Control Barrier Functions into RGB-Only Safety Filters for Dynamic Visual Navigation
- 6. OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
- 7. Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection
Lumen Research Digest — 2026-09-29
2026-09-29 · 7 items
- 1. CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation
- 2. Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos
- 3. Recent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent Intelligence
- 4. Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors
- 5. Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
- 6. JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
- 7. GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
Lumen Research Digest — 2026-09-28
2026-09-28 · 7 items
- 1. DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation
- 2. InternW0-$\Delta$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
- 3. SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery
- 4. Gauss What You Need: Compact Gaussian Splatting Across Scene Scales
- 5. MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes
- 6. Structured Reasoning Agentic Framework for Interpretable Critical View of Safety Assessment
- 7. RECAST: From Log Replay to Closed-Loop Driving Simulation with View-Complete Actors
Lumen Research Digest — 2026-09-27
2026-09-27 · 7 items
- 1. OceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning
- 2. Jev-Mobile: Jev as an Executor for Mobile GUI Agents
- 3. Industrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space
- 4. AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution
- 5. When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression
- 6. PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation
- 7. SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection
Lumen Research Digest — 2026-09-26
2026-09-26 · 7 items
- 1. RAPID: Robot Agentic Programming from Demonstrations
- 2. Multimodal Thinking with Renderable Programs
- 3. RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning
- 4. Markerless Multi-Modal Autonomous Robotic Inspection of Large Space Structures
- 5. SEE Challenge 2026: Event-Guided Brightness Adjustment Across a Broad Illumination Range
- 6. Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents
- 7. Pistis Technical Report
Lumen Research Digest — 2026-09-25
2026-09-25 · 7 items
- 1. World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
- 2. Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation
- 3. SWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering
- 4. From Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments
- 5. BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines
- 6. HarnessPAI: An Evolving Harness for Physical AI
- 7. Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning
Lumen Research Digest — 2026-09-24
2026-09-24 · 7 items
- 1. EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory
- 2. GaussPDE: Graph-Based Partial Differential Equation-Driven Rendering for 3D Gaussian Splatting
- 3. ACTS: A multi-tier benchmark evaluating LLM cipher identification under controlled blind conditions
- 4. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
- 5. HINT-Blimp: Human INTent Inference from Multimodal Cues for Robotic Blimps
- 6. CoRelNav: Collaborative Relational Navigation for Multi-Robot Spatially Constrained Semantic Navigation
- 7. Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations
Lumen Research Digest — 2026-09-23
2026-09-23 · 4 items
Lumen Research Digest — 2026-09-22
2026-09-22 · 4 items
Lumen Research Digest — 2026-09-21
2026-09-21 · 4 items
Lumen Research Digest — 2026-09-20
2026-09-20 · 4 items
Lumen Research Digest — 2026-09-19
2026-09-19 · 4 items
Lumen Research Digest — 2026-09-18
2026-09-18 · 4 items
Lumen Research Digest — 2026-09-17
2026-09-17 · 4 items
Lumen Research Digest — 2026-09-16
2026-09-16 · 5 items
- 1. PanoGS-SLAM: Panoramic 3D Gaussian Splatting SLAM
- 2. How Fyxer built an AI executive assistant people trust
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. RobResilience: Implementing and Evaluating a Resilience Framework for Cyber-Physical Embodied Systems
- 5. You Shall Not Pass into Ring-0! A User Privacy-Friendly Anti-Cheat Architecture for Personal Computers
Lumen Research Digest — 2026-09-15
2026-09-15 · 4 items
Lumen Research Digest — 2026-09-14
2026-09-14 · 4 items
Lumen Research Digest — 2026-09-13
2026-09-13 · 4 items
- 1. How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules
- 2. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
- 3. Cognition helps Devin test its own work with GPT‑6 Astra
- 4. Echoverse: Deep, evolving environments for computer-use agents
Lumen Research Digest — 2026-09-12
2026-09-12 · 6 items
- 1. Learning Agent-based Model Predictive Control for Holistic Vehicle Performance
- 2. Rapidly scaling online storage to serve over 1 billion ChatGPT users
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. BlueSTAR: Tiered Agentic Architecture for Autonomous Cyber Defense
- 5. Artificial Id: Drive and Persistent Alignment in Agentic AI
- 6. Perplexity trusts GPT-6 Astra with end-to-end systems
Lumen Research Digest — 2026-09-11
2026-09-11 · 6 items
- 1. SenseNova-U1.5: Towards Native Unified Visual Intelligence
- 2. Introducing the Agents API
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. MindTopo: Can Foundation Models Reason in Topological Space?
- 5. Domain-Specific Hallucination Detection in Large Language Models
- 6. Now everyone can put data to work
Lumen Research Digest — 2026-09-10
2026-09-10 · 6 items
- 1. Programmable World Model
- 2. Paul Christiano joins OpenAI Foundation Board
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
- 5. Towards Tackling Application Logic Flaws through Autonomous Formal-Logic Modeling and Automated Reasoning
- 6. GPT-6 Astra: The next generation in intelligence for work
Lumen Research Digest — 2026-09-09
2026-09-09 · 6 items
- 1. ReCite: Agentic Reasoning for Faithful Citation
- 2. 1Password increases engineering productivity 21% with Codex
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
- 5. A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes
- 6. How GPT-5.6 Sol helps run quantum computing experiments
Lumen Research Digest — 2026-09-08
2026-09-08 · 5 items
- 1. WorldSculpt: Generating Compositional Worlds from Grounded Videos
- 2. Legora reviewed 41 documents in minutes with GPT-6 Astra
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation
- 5. CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls
Lumen Research Digest — 2026-09-07
2026-09-07 · 5 items
- 1. RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
- 2. Research acceleration: The view inside OpenAI
- 3. Broadening access to Skala creates a faster path to predictive DFT
- 4. Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation
- 5. Compact Neural Appearance Models for Efficient Gaussian Splatting
Lumen Research Digest — 2026-09-06
2026-09-06 · 5 items
- 1. PatchBench: Evaluating AI Agents for Vulnerability Patching
- 2. Daybreak for Frontline Defenders: $1B to protect essential services
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation
- 5. Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations
Lumen Research Digest — 2026-09-05
2026-09-05 · 5 items
- 1. Formation Matrix and Energy-based Control of Multi-Agent Systems
- 2. GPT-6 Astra: A new generation of intelligence
- 3. Orchard: An open framework for scalable agentic AI
- 4. Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
- 5. Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
Lumen Research Digest — 2026-09-04
2026-09-04 · 6 items
- 1. Principia: Relational Physics Tests for Video Models
- 2. Safety overview: GPT-6 Astra
- 3. EvoLib: Turning experience into evolving knowledge
- 4. SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center
- 5. Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
- 6. Playco cut manual fixes 50% prototyping games with GPT-6 Astra
Lumen Research Digest — 2026-09-03
2026-09-03 · 6 items
- 1. RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation
- 2. How law firm Gilbert + Tobin governs and scales AI with OpenAI
- 3. Verifying Rust cryptography in SymCrypt, from standards to code
- 4. Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework
- 5. Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis
- 6. ATV Big Air Tour turned 3 days of work into 3 hours with ChatGPT
Lumen Research Digest — 2026-09-02
2026-09-02 · 6 items
- 1. CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
- 2. How AI-native companies turn workflows into operating capability
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting
- 5. TempCloze: Can Video-LLMs Identify the Missing Middle?
- 6. Path to Astra: critical capabilities and frontier safeguards
Lumen Research Digest — 2026-09-01
2026-09-01 · 5 items
- 1. Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
- 2. Polimill builds Japan's next-generation public AI infrastructure
- 3. GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
- 4. BRF-GS: Hyperspectral Bidirectional Reflectance Factor Modeling and Image Generation Based on 3D Gaussian Splatting
- 5. Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
Lumen Research Digest — 2026-08-31
2026-08-31 · 5 items
- 1. ChainSplat: A Physics-Inspired Screw-Theoretic Model for Learning Deformable Linear Object Dynamics from Multi-View RGB Videos
- 2. Expanding OpenAI’s presence in Brazil
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction
- 5. LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment
Lumen Research Digest — 2026-08-30
2026-08-30 · 5 items
- 1. Comparative Evaluation of 3D Reconstruction Methods for Immersive Visualization of Laboratory Objects
- 2. Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
- 5. Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
Lumen Research Digest — 2026-08-29
2026-08-29 · 5 items
- 1. UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
- 2. Learning never stops: How AI makes learning continuous
- 3. MindTopo reveals VLMs’ spatial reasoning abilities
- 4. Embodied Scene Rearrangement Planning
- 5. R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models
Lumen Research Digest — 2026-08-28
2026-08-28 · 4 items
- 1. How loveholidays is making everyone a builder with Codex
- 2. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
- 3. The full stack behind abundant intelligence
- 4. Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Lumen Research Digest — 2026-08-27
2026-08-27 · 6 items
- 1. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
- 2. Bringing ChatGPT for Teachers to more U.S. school districts
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
- 5. Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence
- 6. The Hugging Face incident and the road ahead
Lumen Research Digest — 2026-08-26
2026-08-26 · 6 items
- 1. BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes
- 2. Jalapeño’s first results show industry-leading speed and efficiency in AI inference
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
- 5. From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
- 6. Introducing the Admin plugin for ChatGPT Work and Codex
Lumen Research Digest — 2026-08-25
2026-08-25 · 5 items
- 1. FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors
- 2. Advancing price-performance for developers with GPT‑5.6 in Kiro
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
- 5. SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
Lumen Research Digest — 2026-08-24
2026-08-24 · 5 items
- 1. ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations
- 2. Introducing AI Futures
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. AI with Authority, from Application to Silicon
- 5. Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis
Lumen Research Digest — 2026-08-23
2026-08-23 · 5 items
- 1. Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
- 2. ChatGPT Ads expands across Europe
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models
- 5. The Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents
Lumen Research Digest — 2026-08-22
2026-08-22 · 5 items
- 1. AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- 2. Offering Zero Data Retention for frontier models
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
- 5. MidTool: Mid-training Data Synthesis for Agentic Tool Use
Lumen Research Digest — 2026-08-21
2026-08-21 · 6 items
- 1. DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery
- 2. Broadening access to Skala creates a faster path to predictive DFT
- 3. Stampli cuts launch hours by 68% using ChatGPT Work
- 4. 4DAnyone: Create Anyone in 4D from a Casual Monocular Video
- 5. Video2DoorTraversal: Push Door Traversal via Simulated Door Twins
- 6. Replit expands access to software creation with GPT-5.6 Luna
Lumen Research Digest — 2026-08-20
2026-08-20 · 6 items
- 1. LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding
- 2. Pacing model development in an era of cyber-critical capabilities
- 3. Orchard: An open framework for scalable agentic AI
- 4. SPADE: Self-Play in Adaptive Synthetic Executable Environments
- 5. What is Missing from AI Post-Training AI: An Empirical Analysis
- 6. Partnering with CodeAI to prepare the first AI generation
Lumen Research Digest — 2026-08-19
2026-08-19 · 6 items
- 1. GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
- 2. Asana cleared 5 years of engineering work in 2 weeks with Codex
- 3. Flint: A visualization language for the AI era
- 4. Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation
- 5. Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
- 6. Strengthening democratic oversight in national security
Lumen Research Digest — 2026-08-18
2026-08-18 · 5 items
- 1. When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents
- 2. The Defender’s Window
- 3. Verifying Rust cryptography in SymCrypt, from standards to code
- 4. Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation
- 5. FlexWorm: Primitive-augmented Hybrid Contact-motion Planning for Suction-based Multi-segment Deformable Robots
Lumen Research Digest — 2026-08-17
2026-08-17 · 5 items
- 1. Twin: Playing an Unknown Game with a Test-Time Digital Twin
- 2. OpenAI’s letter to Governor Abbott on responsible AI infrastructure in Texas
- 3. EvoLib: Turning experience into evolving knowledge
- 4. You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model
- 5. Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Lumen Research Digest — 2026-08-16
2026-08-16 · 5 items
- 1. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
- 2. OpenAI appoints Dali Rajic as Chief Revenue Officer
- 3. Echoverse: Deep, evolving environments for computer-use agents
- 4. Vero: Can AI Agents Build Formally Verified Software Repositories?
- 5. DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Lumen Research Digest — 2026-08-15
2026-08-15 · 5 items
- 1. PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
- 2. Testing ads in ChatGPT
- 3. Flint: A visualization language for the AI era
- 4. AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
- 5. GS$^{2}$CI: Robust Gaussian Splatting For Snapshot Compressive Imaging via Large Vision Model Priors
Lumen Research Digest — 2026-08-14
2026-08-14 · 6 items
- 1. Intern-S2-Preview: Scientific Agentic Foundation Model
- 2. The builder’s guide to GPT‑5.6
- 3. Flint: A visualization language for the AI era
- 4. OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
- 5. TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval
- 6. Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
Lumen Research Digest — 2026-08-13
2026-08-13 · 6 items
- 1. AVA-Encoder: Towards Agent-Native Video Representation Learning
- 2. From assistance to execution: How enterprises put AI to work
- 3. MindTopo reveals VLMs’ spatial reasoning abilities
- 4. StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
- 5. Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
- 6. How RingCentral builds AI-native work from engineering to ops
Lumen Research Digest — 2026-08-12
2026-08-12 · 4 items
- 1. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
- 2. Daybreak models are now available on AWS
- 3. Premium seats are coming to ChatGPT Business
- 4. Aurora 1.5: Extending open foundation models for weather and Earth-system applications
Lumen Research Digest — 2026-08-11
2026-08-11 · 6 items
- 1. Defining Decentralization: An Ontological Perspective
- 2. Expanding Daybreak as the Cyber Defense Window Narrows
- 3. SkillOpt: Agent skills as trainable parameters
- 4. Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection
- 5. Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
- 6. Putting frontier cyber models in more trusted hands
Lumen Research Digest — 2026-08-10
2026-08-10 · 5 items
- 1. Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI
- 2. From asking to doing: How the world is putting ChatGPT to work
- 3. Flint: A visualization language for the AI era
- 4. I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
- 5. SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
Lumen Research Digest — 2026-08-09
2026-08-09 · 5 items
- 1. RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer
- 2. Working with the American Psychological Association on youth mental health and AI
- 3. Flint: A visualization language for the AI era
- 4. Depth-Guided Video Object Counting in Crowded Scenes
- 5. A Master-Salve Robot Manipulator for Needle-Based Teleoperation in MRI Chamber
Lumen Research Digest — 2026-08-08
2026-08-08 · 6 items
- 1. Learning When to Trust via Selective Context Preference Optimization
- 2. Responding to the next frontier of critical cyber capabilities
- 3. Flint: A visualization language for the AI era
- 4. Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
- 5. TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
- 6. How HSP GRUPPE builds AI capabilities for tax advisory
Lumen Research Digest — 2026-08-07
2026-08-07 · 5 items
- 1. GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
- 2. Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
- 3. Flint: A visualization language for the AI era
- 4. The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
- 5. MASS: Multiplayer World Models with Authoritative Shared State
Lumen Research Digest — 2026-08-06
2026-08-06 · 5 items
- 1. SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
- 2. Apple is getting this wrong
- 3. Flint: A visualization language for the AI era
- 4. AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance
- 5. Objects as Audio-Visual Modal Sound Fields
Lumen Research Digest — 2026-08-05
2026-08-05 · 6 items
- 1. When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
- 2. Third-party cyber evaluations involving OpenAI models
- 3. Flint: A visualization language for the AI era
- 4. UniWorld-Design: From Pixel Generation to Layer-Native Design
- 5. Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
- 6. New ways to learn and teach with ChatGPT Work and Codex
Lumen Research Digest — 2026-08-04
2026-08-04 · 5 items
- 1. Antares: Foundation Models for Agentic Vulnerability Localization
- 2. Orchard: An open framework for scalable agentic AI
- 3. Circles powers telco personalization with OpenAI technology
- 4. Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
- 5. A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI
Lumen Research Digest — 2026-08-03
2026-08-03 · 5 items
- 1. CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding
- 2. Ten advances in mathematics and theoretical computer science
- 3. Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
- 4. FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling
- 5. OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting
Lumen Research Digest — 2026-08-02
2026-08-02 · 5 items
- 1. Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
- 2. Advancing the price-performance frontier with GPT-5.6
- 3. Flint: A visualization language for the AI era
- 4. ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
- 5. DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation
Lumen Research Digest — 2026-08-01
2026-08-01 · 5 items
- 1. OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
- 2. Advancing responsible AI across Europe
- 3. Verifying Rust cryptography in SymCrypt, from standards to code
- 4. X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching
- 5. ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Lumen Research Digest — 2026-07-31
2026-07-31 · 6 items
- 1. FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation
- 2. Echoverse: Deep, evolving environments for computer-use agents
- 3. How avatarin built a 24/7 retail agent with GPT-Realtime
- 4. ORCA-bench: How Ready Are Language Model Agents for Oncall?
- 5. PhiZero: A World Model Built Around Physical Language
- 6. EvoLib: Turning experience into evolving knowledge
Lumen Research Digest — 2026-07-30
2026-07-30 · 6 items
- 1. Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications
- 2. How GPT-5.6 fuses frontier intelligence with frontier efficiency
- 3. Understanding the brain with AI-driven explanations and experiments
- 4. AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching
- 5. MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
- 6. How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
Lumen Research Digest — 2026-07-29
2026-07-29 · 5 items
- 1. Wonder: Video World Model Done Better
- 2. Scientific computing in the age of agentic AI
- 3. Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
- 4. Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition
- 5. Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics
Lumen Research Digest — 2026-07-28
2026-07-28 · 5 items
Lumen Research Digest — 2026-07-27
2026-07-27 · 4 items
Lumen Research Digest — 2026-07-26
2026-07-26 · 5 items
- 1. Unified Video Dense Prediction from Disjoint Data
- 2. Advancing the next era of national science
- 3. Flint: A visualization language for the AI era
- 4. MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
- 5. Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
Lumen Research Digest — 2026-07-25
2026-07-25 · 5 items
- 1. 3D-Aware VLMs with Implicit and Explicit Geometries
- 2. How news organizations are using AI to advance their vital missions
- 3. SkillOpt: Agent skills as trainable parameters
- 4. Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering
- 5. Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
Lumen Research Digest — 2026-07-24
2026-07-24 · 5 items
- 1. OpenForgeRL: Train Harness-native Agents in Any Environment
- 2. Building AI infrastructure with the Effingham County community
- 3. Flint: A visualization language for the AI era
- 4. GS-Agent: Creating 4D Physical Worlds With Generative Simulation
- 5. Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog
Lumen Research Digest — 2026-07-23
2026-07-23 · 6 items
- 1. Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
- 2. Introducing OpenAI Presence
- 3. Flint: A visualization language for the AI era
- 4. ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion
- 5. MR-Compare: A Mixed-Reality Framework for Spatially Grounded Visual Comparison of 3D Gaussian Splatting and Mesh Reconstructions with the Physical Environment
- 6. NTT DATA Group cuts incident analysis to 30 minutes with Codex
Lumen Research Digest — 2026-07-22
2026-07-22 · 5 items
- 1. They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface
- 2. OpenAI and Hugging Face partner to address security incident during model evaluation
- 3. Flint: A visualization language for the AI era
- 4. MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings
- 5. Agents in the Wild: Where Research Meets Deployment
Lumen Research Digest — 2026-07-21
2026-07-21 · 5 items
- 1. FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
- 2. Safety and alignment in an era of long-horizon models
- 3. Flint: A visualization language for the AI era
- 4. Robust Multimodal Dynamic Object Segmentation
- 5. O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
Lumen Research Digest — 2026-07-20
2026-07-20 · 5 items
- 1. Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs
- 2. Why teens deserve access to safe AI
- 3. Ire identifies another LOTUSLITE specimen
- 4. The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure
- 5. Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps
Lumen Research Digest — 2026-07-19
2026-07-19 · 5 items
- 1. RoboTTT: Context Scaling for Robot Policies
- 2. GPT-Red: Unlocking Self-Improvement for Robustness
- 3. Flint: A visualization language for the AI era
- 4. SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
- 5. MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Lumen Research Digest — 2026-07-18
2026-07-18 · 6 items
- 1. Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search
- 2. How Cars24 scales conversations and builds faster with OpenAI
- 3. Flint: A visualization language for the AI era
- 4. Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA
- 5. Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents
- 6. A scorecard for the AI age
Lumen Research Digest — 2026-07-17
2026-07-17 · 4 items
- 1. Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
- 2. Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
- 3. MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization
- 4. Hierarchical Denoising For Multi-Step Visual Reasoning
Lumen Research Digest — 2026-07-16
2026-07-16 · 5 items
- 1. ProfMalPlus: Agent-Coordinated Detection of Malicious NPM Packages via Static-Dynamic Analysis Synergy
- 2. The US is advancing AI safety through state and federal action
- 3. Flint: A visualization language for the AI era
- 4. Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation
- 5. VisualRepair: Dynamic Tool Calling and Region Focusing for Visual Software Issue Repair
Lumen Research Digest — 2026-07-15
2026-07-15 · 6 items
- 1. MAMMOTH: A Multi-Modal End-to-End Policy for Off-Road Mobility Robust to Missing Modality
- 2. How to manage AI investments in the agentic era
- 3. Flint: A visualization language for the AI era
- 4. Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution
- 5. TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
- 6. How data science teams use ChatGPT Work
Lumen Research Digest — 2026-07-14
2026-07-14 · 5 items
- 1. Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
- 2. Verifying Rust cryptography in SymCrypt, from standards to code
- 3. Getting started with ChatGPT
- 4. Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems
- 5. When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
Lumen Research Digest — 2026-07-13
2026-07-13 · 5 items
- 1. Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
- 2. How Deutsche Telekom is rewiring telecommunications with AI
- 3. Understanding the brain with AI-driven explanations and experiments
- 4. 4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception
- 5. VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
Lumen Research Digest — 2026-07-12
2026-07-12 · 5 items
- 1. UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- 2. GPT-5.6: Frontier intelligence that scales with your ambition
- 3. Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
- 4. Wat3R: Underwater 3D Geometry Learning without Annotations
- 5. ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation
Lumen Research Digest — 2026-07-11
2026-07-11 · 5 items
- 1. Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction
- 2. Helping K–12 educators build practical AI skills
- 3. Flint: A visualization language for the AI era
- 4. AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding
- 5. ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
Lumen Research Digest — 2026-07-10
2026-07-10 · 4 items
Lumen Research Digest — 2026-07-09
2026-07-09 · 6 items
- 1. CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis
- 2. Flint: A visualization language for the AI era
- 3. Our approach to government and national security partnerships
- 4. Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning
- 5. Infinite Worlds with Versatile Interactions
- 6. Separating signal from noise in coding evaluations
Lumen Research Digest — 2026-07-08
2026-07-08 · 5 items
- 1. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
- 2. Australian Payments Plus moves faster with ChatGPT and Codex
- 3. SkillOpt: Agent skills as trainable parameters
- 4. CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
- 5. RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation
Lumen Research Digest — 2026-07-07
2026-07-07 · 5 items
- 1. LLM-as-a-Verifier: A General-Purpose Verification Framework
- 2. MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
- 3. How agents are transforming work
- 4. InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics
- 5. GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks
Lumen Research Digest — 2026-07-06
2026-07-06 · 5 items
- 1. Controllable Sim Agents with Behavior Latents
- 2. SkillOpt: Agent skills as trainable parameters
- 3. How agents are transforming work
- 4. GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training
- 5. ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
Lumen Research Digest — 2026-07-05
2026-07-05 · 5 items
- 1. Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study
- 2. SkillOpt: Agent skills as trainable parameters
- 3. How agents are transforming work
- 4. Steerability via constraints: a substrate for scalable oversight of coding agents
- 5. Distributed Attacks in Persistent-State AI Control
Lumen Research Digest — 2026-07-04
2026-07-04 · 5 items
- 1. Understanding Agent-Based Patching of Compiler Missed Optimizations
- 2. SkillOpt: Agent skills as trainable parameters
- 3. How agents are transforming work
- 4. Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
- 5. TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution
Lumen Research Digest — 2026-07-03
2026-07-03 · 5 items
- 1. Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas
- 2. HP Inc. launches Frontier strategic partnership with OpenAI
- 3. Ire identifies another LOTUSLITE specimen
- 4. Alignment Is All You Need For X-to-4D Generation
- 5. WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
Lumen Research Digest — 2026-07-02
2026-07-02 · 5 items
- 1. Structured 4D Latent Predictive Model for Robot Planning
- 2. Core dump epidemiology: fixing an 18-year-old bug
- 3. Extending Human Intelligence Through AI
- 4. World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video
- 5. RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue
Lumen Research Digest — 2026-07-01
2026-07-01 · 6 items
- 1. PointSplat: Compact Gaussian Splatting via Human-Centric Prediction
- 2. SkillOpt: Agent skills as trainable parameters
- 3. Introducing GeneBench-Pro
- 4. DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
- 5. ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs
- 6. How ChatGPT adoption has expanded
Lumen Research Digest — 2026-06-30
2026-06-30 · 5 items
- 1. VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
- 2. Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
- 3. Mapping Europe’s AI Workforce Opportunity
- 4. Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors
- 5. UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image
Lumen Research Digest — 2026-06-29
2026-06-29 · 5 items
- 1. StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views
- 2. How Omio is building the future of conversational travel
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. Agent-Native Immune System: Architecture, Taxonomy, and Engineering
- 5. HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
Lumen Research Digest — 2026-06-28
2026-06-28 · 5 items
- 1. Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
- 2. Helping build shared standards for advanced AI
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts
- 5. Autoregressive Boltzmann Generators
Lumen Research Digest — 2026-06-27
2026-06-27 · 5 items
- 1. Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy
- 2. Previewing GPT-5.6 Sol: a next-generation model
- 3. Vega: Zero-knowledge proofs for digital identity in the age of AI
- 4. NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems
- 5. Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
Lumen Research Digest — 2026-06-26
2026-06-26 · 5 items
- 1. PhysiFormer: Learning to Simulate Mechanics in World Space
- 2. How agents are transforming work
- 3. Understanding the brain with AI-driven explanations and experiments
- 4. OctoSense: Self-Supervised Learning for Multimodal Robot Perception
- 5. E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
Lumen Research Digest — 2026-06-25
2026-06-25 · 5 items
- 1. RoboAtlas: Contextual Active SLAM
- 2. OpenAI and Broadcom unveil LLM-optimized inference chip
- 3. Talos: Scaling rare disease diagnosis with automated, iterative genomic reanalysis
- 4. TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs
- 5. Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models
Lumen Research Digest — 2026-06-24
2026-06-24 · 6 items
- 1. Pocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAM
- 2. Patch the Planet: a Daybreak initiative to support open source maintainers
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
- 5. Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment
- 6. How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery
Lumen Research Digest — 2026-06-23
2026-06-23 · 6 items
- 1. Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse
- 2. Daybreak: Tools for securing every organization in the world
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- 5. AIR: Adaptive Interleaved Reasoning with Code in MLLMs
- 6. Codex-maxxing for long-running work
Lumen Research Digest — 2026-06-22
2026-06-22 · 5 items
- 1. Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software
- 2. Samsung Electronics brings ChatGPT and Codex to employees
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
- 5. Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
Lumen Research Digest — 2026-06-21
2026-06-21 · 5 items
- 1. Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes
- 2. GridSFM: A new, small foundation model for the electric grid
- 3. Introducing LifeSciBench
- 4. TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living
- 5. Efficient and Sound Probabilistic Verification for AI Agents
Lumen Research Digest — 2026-06-20
2026-06-20 · 5 items
- 1. Thinking in Boxes: 3D Editing in Real Images Made Easy
- 2. New usage analytics and updated spend controls for enterprises
- 3. MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
- 4. SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
- 5. LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
Lumen Research Digest — 2026-06-19
2026-06-19 · 6 items
- 1. Current World Models Lack a Persistent State Core
- 2. Improving health intelligence in ChatGPT
- 3. Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability
- 4. DeepSWIP: Quotient-WMC Counterfactuals for Neural Probabilistic Logic Programs
- 5. Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving
- 6. Using AI to help physicians diagnose rare genetic diseases affecting children
Lumen Research Digest — 2026-06-18
2026-06-18 · 6 items
- 1. OneCanvas: 3D Scene Understanding via Panoramic Reprojection
- 2. Introducing LifeSciBench
- 3. mimalloc: A new, high-performance, scalable memory allocator for the modern era
- 4. Native Active Perception as Reasoning for Omni-Modal Understanding
- 5. A Mixed-Reality Testbed for Autonomous Vehicles
- 6. A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
Lumen Research Digest — 2026-06-17
2026-06-17 · 5 items
- 1. EgoCS-400K: An Egocentric Gameplay Dataset for World Models
- 2. Predicting model behavior before release by simulating deployment
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners
- 5. Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion
Lumen Research Digest — 2026-06-16
2026-06-16 · 5 items
- 1. Geometric Action Model for Robot Policy Learning
- 2. How Preply combines AI and human tutors to personalize learning
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. R2RDreamer: 3D-aware Data Augmentation for Spatially-generalized 2D Manipulation Policies
- 5. Context-Aware RL for Agentic and Multimodal LLMs
Lumen Research Digest — 2026-06-15
2026-06-15 · 5 items
- 1. AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition
- 2. Introducing the OpenAI Partner Network
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains
- 5. RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space
Lumen Research Digest — 2026-06-14
2026-06-14 · 5 items
- 1. World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
- 2. BBVA puts AI at the core of banking with OpenAI
- 3. Extending Human Intelligence Through AI
- 4. Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks
- 5. Surflo: Consistent 3D Surface Flow Model with Global State
Lumen Research Digest — 2026-06-13
2026-06-13 · 6 items
- 1. $\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
- 2. New OpenAI Academy courses for the next era of work
- 3. Ire identifies another LOTUSLITE specimen
- 4. RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
- 5. Agents-K1: Towards Agent-native Knowledge Orchestration
- 6. PRC-linked influence operations are targeting AI debates in the US
Lumen Research Digest — 2026-06-12
2026-06-12 · 6 items
- 1. InterleaveThinker: Reinforcing Agentic Interleaved Generation
- 2. OpenAI to acquire Ona
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
- 5. Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction
- 6. Supporting Europe’s work in ensuring a trustworthy AI ecosystem
Lumen Research Digest — 2026-06-11
2026-06-11 · 6 items
- 1. DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?
- 2. Access OpenAI models and Codex through your Oracle cloud commitment
- 3. Vega: Zero-knowledge proofs for digital identity in the age of AI
- 4. A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents
- 5. Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
- 6. How an astrophysicist uses Codex to help simulate black holes
Lumen Research Digest — 2026-06-10
2026-06-10 · 6 items
- 1. WorldOlympiad: Can Your World Model Survive a Triathlon?
- 2. How engineers at Nextdoor use Codex to build without limits
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
- 5. P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning
- 6. What Codex unlocks for Notion
Lumen Research Digest — 2026-06-09
2026-06-09 · 5 items
- 1. Latent Spatial Memory for Video World Models
- 2. Built to benefit everyone: our plan
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. iMaC: Translating Actions into Motion and Contact Images for Embodied World Models
- 5. Observability for Delegated Execution in Agentic AI Systems
Lumen Research Digest — 2026-06-08
2026-06-08 · 5 items
- 1. MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism
- 2. Dreaming: Better memory for a more helpful ChatGPT
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. UniSHARP: Universal Sharp Monocular View Synthesis
- 5. Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Lumen Research Digest — 2026-06-07
2026-06-07 · 5 items
- 1. You Only Index Once: Cross-Layer Sparse Attention with Shared Routing
- 2. OpenAI public policy agenda
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. Robust Ensemble of Selectively Strengthened and Augmented Predictors
- 5. DNQ: Deep Nash Q-Network for Partially Observable n-Player Games
Lumen Research Digest — 2026-06-06
2026-06-06 · 5 items
- 1. PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding
- 2. Travelers deploys AI-powered claims countrywide with OpenAI
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation
- 5. Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators
Lumen Research Digest — 2026-06-05
2026-06-05 · 6 items
- 1. Benchmark Everything Everywhere All at Once
- 2. How Endava is redesigning software delivery around AI agents
- 3. Data Formulator 0.7: AI-powered data analytics for enterprise data
- 4. StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset
- 5. Will the Agent Recuse Itself? Measuring LLM-Agent Compliance with In-Band Access-Deny Signals
- 6. How Wasmer used Codex to build a Node.js runtime for the edge
Lumen Research Digest — 2026-06-04
2026-06-04 · 4 items
Lumen Research Digest — 2026-06-03
2026-06-03 · 4 items
Lumen Research Digest — 2026-06-02
2026-06-02 · 4 items
Lumen Research Digest — 2026-06-01
2026-06-01 · 4 items
Lumen Research Digest — 2026-05-31
2026-05-31 · 6 items
- 1. DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
- 2. Boston Children’s uses AI to unlock new diagnoses
- 3. Advancing AI for materials with MatterSim: experimental synthesis, faster simulation, and multi-task models
- 4. RoboWits: Unexpected Challenges for Robotic Creative Problem Solving
- 5. GMOS: Grounding Moving Object Segmentation in 3D Space and Time
- 6. Strengthening societal resilience with Rosalind Biodefense
Lumen Research Digest — 2026-05-30
2026-05-30 · 4 items
Lumen Research Digest — 2026-05-29
2026-05-29 · 4 items
Lumen Research Digest — 2026-05-28
2026-05-28 · 4 items
Lumen Research Digest — 2026-05-27
2026-05-27 · 4 items
- 1. OpenAI, Grupo Folha and Grupo UOL announce strategic content partnership
- 2. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 3. OpenAI named a Leader in enterprise coding agents by Gartner
- 4. MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
Lumen Research Digest — 2026-05-26
2026-05-26 · 5 items
- 1. TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
- 2. AdventHealth advances whole-person care with OpenAI
- 3. Building realistic electric transmission grid dataset at scale: a pipeline from open dataset
- 4. Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
- 5. AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond
Lumen Research Digest — 2026-05-25
2026-05-25 · 5 items
- 1. ETCHR: Editing To Clarify and Harness Reasoning
- 2. Advancing content provenance for a safer, more transparent AI ecosystem
- 3. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 4. Agentic Proving for Program Verification
- 5. SkillOpt: Executive Strategy for Self-Evolving Agent Skills
Lumen Research Digest — 2026-05-24
2026-05-24 · 5 items
- 1. Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
- 2. An OpenAI model has disproved a central conjecture in discrete geometry
- 3. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 4. LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems
- 5. MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
Lumen Research Digest — 2026-05-23
2026-05-23 · 6 items
- 1. GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
- 2. OpenAI named a Leader in enterprise coding agents by Gartner
- 3. Vega: Zero-knowledge proofs for digital identity in the age of AI
- 4. Cambrian-P: Pose-Grounded Video Understanding
- 5. AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
- 6. How Virgin Atlantic ships faster with Codex
Lumen Research Digest — 2026-05-22
2026-05-22 · 4 items
- 1. MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models
- 2. The next phase of OpenAI’s Education for Countries
- 3. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 4. OpenAI and Dell partner to bring Codex to hybrid and on-premise enterprise environments
Lumen Research Digest — 2026-05-21
2026-05-21 · 5 items
- 1. DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
- 2. How Ramp engineers accelerate code review with Codex
- 3. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 4. You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
- 5. Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
Lumen Research Digest — 2026-05-20
2026-05-20 · 5 items
- 1. ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
- 2. Introducing OpenAI for Singapore
- 3. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 4. What Do Evolutionary Coding Agents Evolve?
- 5. TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization
Lumen Research Digest — 2026-05-19
2026-05-19 · 5 items
- 1. Code as Agent Harness
- 2. OpenAI and Dell partner to bring Codex to hybrid and on-premise enterprise environments
- 3. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 4. Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory
- 5. ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
Lumen Research Digest — 2026-05-18
2026-05-18 · 6 items
- 1. Argus: Evidence Assembly for Scalable Deep Research Agents
- 2. OpenAI and Malta partner to bring ChatGPT Plus to all citizens
- 3. Red-teaming a network of agents: Understanding what breaks when AI agents interact at scale
- 4. Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
- 5. IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation
- 6. How data science teams use Codex
Lumen Research Digest — 2026-05-17
2026-05-17 · 4 items
Lumen Research Digest — 2026-05-16
2026-05-16 · 6 items
- 1. Pelican-Unified 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action
- 2. Databricks brings GPT-5.5 to enterprise agent workflows
- 3. Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability
- 4. APWA: A Distributed Architecture for Parallelizable Agentic Workflows
- 5. Articraft: An Agentic System for Scalable Articulated 3D Asset Generation
- 6. Work with Codex from anywhere
Lumen Research Digest — 2026-05-15
2026-05-15 · 4 items
Lumen Research Digest — 2026-05-14
2026-05-14 · 7 items
- 1. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
- 2. Building a safe, effective sandbox to enable Codex on Windows
- 3. mimalloc: A new, high-performance, scalable memory allocator for the modern era
- 4. Harnessing Agentic Evolution
- 5. Good Agentic Friends Do Not Just Give Verbal Advice: They Can Update Your Weights
- 6. Our response to the TanStack npm supply chain attack
- 7. GridSFM: A new, small foundation model for the electric grid
Lumen Research Digest — 2026-05-13
2026-05-13 · 6 items
- 1. LychSim: A Controllable and Interactive Simulation Framework for Vision Research
- 2. What Parameter Golf taught us about AI-assisted research
- 3. Advancing AI for materials with MatterSim: experimental synthesis, faster simulation, and multi-task models
- 4. SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- 5. MEME: Multi-entity & Evolving Memory Evaluation
- 6. How NVIDIA engineers and researchers build with Codex
Lumen Research Digest — 2026-05-12
2026-05-12 · 5 items
- 1. BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD
- 2. SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
- 3. Advancing voice intelligence with new models in the API
- 4. CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
- 5. From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
Lumen Research Digest — 2026-05-11
2026-05-11 · 5 items
- 1. Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
- 2. Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber
- 3. Microsoft at NSDI 2026: Advances in large-scale networked systems
- 4. LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
- 5. Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
Lumen Research Digest — 2026-05-10
2026-05-10 · 5 items
- 1. Relit-LiVE: Relight Video by Jointly Learning Environment Video
- 2. How ChatGPT learns about the world while protecting privacy
- 3. AutoAdapt: Automated domain adaptation for large language models
- 4. SkillOS: Learning Skill Curation for Self-Evolving Agents
- 5. AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents
Lumen Research Digest — 2026-05-09
2026-05-09 · 5 items
- 1. MedHorizon: Towards Long-context Medical Video Understanding in the Wild
- 2. Running Codex safely at OpenAI
- 3. Building realistic electric transmission grid dataset at scale: a pipeline from open dataset
- 4. ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
- 5. BAMI: Training-Free Bias Mitigation in GUI Grounding
Lumen Research Digest — 2026-05-08
2026-05-08 · 4 items
Lumen Research Digest — 2026-05-07
2026-05-07 · 6 items
- 1. LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents
- 2. How frontier enterprises are building an AI advantage
- 3. Can we AI our way to a more sustainable world?
- 4. Executable World Models for ARC-AGI-3 in the Era of Coding Agents
- 5. Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours
- 6. Singular Bank helps bankers move fast with ChatGPT and Codex
Lumen Research Digest — 2026-05-06
2026-05-06 · 5 items
- 1. An Agent-Oriented Pluggable Experience-RAG Skill for Experience-Driven Retrieval Strategy Orchestration
- 2. Microsoft at NSDI 2026: Advances in large-scale networked systems
- 3. New ways to buy ChatGPT ads
- 4. Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
- 5. Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
Lumen Research Digest — 2026-05-05
2026-05-05 · 5 items
- 1. Standing on the Shoulders of Giants: Stabilized Knowledge Distillation for Cross--Language Code Clone Detection
- 2. OpenAI and PwC collaborate to reimagine the office of the CFO
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents
- 5. When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
Lumen Research Digest — 2026-05-04
2026-05-04 · 5 items
- 1. Self-Adaptive Multi-Agent LLM-Based Security Pattern Selection for IoT Systems
- 2. Our commitment to community safety
- 3. GroundedPlanBench: Spatially grounded long-horizon task planning for robot manipulation
- 4. Position: agentic AI orchestration should be Bayes-consistent
- 5. Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game
Lumen Research Digest — 2026-05-03
2026-05-03 · 5 items
- 1. AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
- 2. Building the compute infrastructure for the Intelligence Age
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
- 5. FlexiTac: A Low-Cost, Open-Source, Scalable Tactile Sensing Solution for Robotic Systems
Lumen Research Digest — 2026-05-02
2026-05-02 · 5 items
- 1. Exploration Hacking: Can LLMs Learn to Resist RL Training?
- 2. Where the goblins came from
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction
- 5. PRISM: Pre-alignment via Black-box On-policy Distillation for Multimodal Reinforcement Learning
Lumen Research Digest — 2026-05-01
2026-05-01 · 5 items
- 1. Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
- 2. Red-teaming a network of agents: Understanding what breaks when AI agents interact at scale
- 3. Introducing Advanced Account Security
- 4. HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
- 5. Generalizable Sparse-View 3D Reconstruction from Unconstrained Images
Lumen Research Digest — 2026-04-30
2026-04-30 · 5 items
- 1. Color-Encoded Illumination for High-Speed Volumetric Scene Reconstruction
- 2. Cybersecurity in the Intelligence Age
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. Bian Que: An Agentic Framework with Flexible Skill Arrangement for Online System Operations
- 5. World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
Lumen Research Digest — 2026-04-29
2026-04-29 · 5 items
- 1. Recursive Multi-Agent Systems
- 2. OpenAI models, Codex, and Managed Agents come to AWS
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. Toward Multimodal Conversational AI for Age-Related Macular Degeneration
- 5. Variational Neural Belief Parameterizations for Robust Dexterous Grasping under Multimodal Uncertainty
Lumen Research Digest — 2026-04-28
2026-04-28 · 6 items
- 1. Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation
- 2. An open-source spec for orchestration: Symphony
- 3. Ideas: Steering AI toward the work future we want
- 4. AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
- 5. World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
- 6. Choco automates food distribution with AI agents
Lumen Research Digest — 2026-04-27
2026-04-27 · 5 items
- 1. QuantClaw: Precision Where It Matters for OpenClaw
- 2. How to get started with Codex
- 3. New Future of Work: AI is driving rapid change, uneven benefits
- 4. EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
- 5. ATRS: Adaptive Trajectory Re-splitting via a Shared Neural Policy for Parallel Optimization
Lumen Research Digest — 2026-04-26
2026-04-26 · 6 items
- 1. Vista4D: Video Reshooting with 4D Point Clouds
- 2. Codex settings
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis
- 5. Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study
- 6. Introducing GPT-5.5
Lumen Research Digest — 2026-04-25
2026-04-25 · 6 items
Lumen Research Digest — 2026-04-24
2026-04-24 · 6 items
- 1. Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
- 2. Automations
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. Context Unrolling in Omni Models
- 5. Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows
- 6. Top 10 uses for Codex at work
Lumen Research Digest — 2026-04-23
2026-04-23 · 6 items
- 1. DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
- 2. Introducing workspace agents in ChatGPT
- 3. AutoAdapt: Automated domain adaptation for large language models
- 4. Automatic Ontology Construction Using LLMs as an External Layer of Memory, Verification, and Planning for Hybrid Intelligent Systems
- 5. Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
- 6. Speeding up agentic workflows with WebSockets in the Responses API
Lumen Research Digest — 2026-04-22
2026-04-22 · 5 items
- 1. A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding
- 2. Scaling Codex to enterprises worldwide
- 3. Will machines ever be intelligent?
- 4. CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
- 5. SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
Lumen Research Digest — 2026-04-21
2026-04-21 · 5 items
- 1. MultiWorld: Scalable Multi-Agent Multi-View Video World Models
- 2. Can we AI our way to a more sustainable world?
- 3. OpenAI helps Hyatt advance AI among colleagues
- 4. OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
- 5. Using large language models for embodied planning introduces systematic safety risks
Lumen Research Digest — 2026-04-20
2026-04-20 · 5 items
- 1. FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
- 2. ChatGPT for research
- 3. ADeLe: Predicting and explaining AI performance across tasks
- 4. DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs
- 5. Semantic Area Graph Reasoning for Multi-Robot Language-Guided Search
Lumen Research Digest — 2026-04-19
2026-04-19 · 5 items
- 1. Blue Data Intelligence Layer: Streaming Data and Agents for Multi-source Multi-modal Data-Centric Applications
- 2. Creating images with ChatGPT
- 3. PlugMem: Transforming raw agent interactions into reusable knowledge
- 4. RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
- 5. Prism: Symbolic Superoptimization of Tensor Programs
Lumen Research Digest — 2026-04-18
2026-04-18 · 5 items
- 1. MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
- 2. Codex for (almost) everything
- 3. GroundedPlanBench: Spatially grounded long-horizon task planning for robot manipulation
- 4. TokenLight: Precise Lighting Control in Images using Attribute Tokens
- 5. Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
Lumen Research Digest — 2026-04-17
2026-04-17 · 5 items
- 1. GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
- 2. Introducing GPT-Rosalind for life sciences research
- 3. Think in Latent Thoughts: A New Paradigm for Gloss-Free Sign Language Translation
- 4. CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
- 5. Accelerating the cyber defense ecosystem that protects us all
Lumen Research Digest — 2026-04-16
2026-04-16 · 4 items
Lumen Research Digest — 2026-04-15
2026-04-15 · 5 items
Lumen Research Digest — 2026-04-14
2026-04-14 · 5 items
- 1. LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
- 2. Enterprises power agentic workflows in Cloudflare Agent Cloud with OpenAI
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
- 5. LottieGPT: Tokenizing Vector Animation for Autoregressive Generation
Lumen Research Digest — 2026-04-13
2026-04-13 · 3 items
Lumen Research Digest — 2026-04-12
2026-04-12 · 5 items
- 1. RewardFlow: Generate Images by Optimizing What You Reward
- 2. Our response to the Axios developer tool compromise
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
- 5. E-3DPSM: A State Machine for Event-Based Egocentric 3D Human Pose Estimation
Lumen Research Digest — 2026-04-11
2026-04-11 · 5 items
- 1. Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
- 2. Applications of AI at OpenAI
- 3. AsgardBench: A benchmark for visually grounded interactive planning
- 4. OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
- 5. PSI: Shared State as the Missing Layer for Coherent AI-Generated Instruments in Personal AI Agents
Lumen Research Digest — 2026-04-10
2026-04-10 · 7 items
- 1. Visually-grounded Humanoid Agents
- 2. CyberAgent moves faster with ChatGPT Enterprise and Codex
- 3. New Future of Work: AI is driving rapid change, uneven benefits
- 4. Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- 5. AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
- 6. Ideas: Steering AI toward the work future we want
- 7. OpenAI Full Fan Mode Contest: Terms & Conditions
Lumen Research Digest — 2026-04-09
2026-04-09 · 4 items
Lumen Research Digest — 2026-04-08
2026-04-08 · 5 items
- 1. CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments
- 2. Announcing the OpenAI Safety Fellowship
- 3. Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model
- 4. CoStream: Codec-Guided Resource-Efficient System for Video Streaming Analytics
- 5. MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control
Lumen Research Digest — 2026-04-07
2026-04-07 · 5 items
- 1. FileGram: Grounding Agent Personalization in File-System Behavioral Traces
- 2. Industrial policy for the Intelligence Age
- 3. Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model
- 4. Analyzing Symbolic Properties for DRL Agents in Systems and Networking
- 5. QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
Lumen Research Digest — 2026-04-06
2026-04-06 · 5 items
- 1. Coupled Control, Structured Memory, and Verifiable Action in Agentic AI (SCRAT -- Stochastic Control with Retrieval and Auditable Trajectories): A Comparative Perspective from Squirrel Locomotion and Scatter-Hoarding
- 2. Helping disaster response teams turn AI into action across Asia
- 3. Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model
- 4. FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation
- 5. A Systematic Security Evaluation of OpenClaw and Its Variants
Lumen Research Digest — 2026-04-05
2026-04-05 · 5 items
- 1. Modulate-and-Map: Crossmodal Feature Mapping with Cross-View Modulation for 3D Anomaly Detection
- 2. STADLER reshapes knowledge work at a 230-year-old company
- 3. Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model
- 4. Batched Contextual Reinforcement: A Task-Scaling Law for Efficient Reasoning
- 5. Model-Based Reinforcement Learning for Control under Time-Varying Dynamics
Lumen Research Digest — 2026-04-04
2026-04-04 · 5 items
- 1. SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
- 2. OpenAI acquires TBPN
- 3. Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model
- 4. Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining
- 5. Generative World Renderer
Lumen Research Digest — 2026-04-03
2026-04-03 · 5 items
Lumen Research Digest — 2026-04-02
2026-04-02 · 4 items
Lumen Research Digest — 2026-04-01
2026-04-01 · 6 items
- 1. SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
- 2. PlugMem: Transforming raw agent interactions into reusable knowledge
- 3. Accelerating the next phase of AI
- 4. EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
- 5. HapCompass: A Rotational Haptic Device for Contact-Rich Robotic Teleoperation
- 6. CORPGEN advances AI agents for real work
Lumen Research Digest — 2026-03-31
2026-03-31 · 4 items
Lumen Research Digest — 2026-03-30
2026-03-30 · 6 items
- 1. Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
- 2. Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model
- 3. Creating with Sora Safely
- 4. GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
- 5. Drive-Through 3D Vehicle Exterior Reconstruction via Dynamic-Scene SfM and Distortion-Aware Gaussian Splatting
- 6. AsgardBench: A benchmark for visually grounded interactive planning
Lumen Research Digest — 2026-03-29
2026-03-29 · 7 items
- 1. Introducing the OpenAI Safety Bug Bounty program
- 2. Powering product discovery in ChatGPT
- 3. A New Framework for Evaluating Voice Agents (EVA)
- 4. Holotron-12B - High Throughput Computer Use Agent
- 5. How we monitor internal coding agents for misalignment
- 6. Inside our approach to the Model Spec
- 7. Helping developers build safer AI experiences for teens
Lumen Research Digest — 2026-03-28
2026-03-28 · 7 items
- 1. Vega: Learning to Drive with Natural Language Instructions
- 2. Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
- 3. Back to Basics: Revisiting ASR in the Age of Voice Agents
- 4. The Kitchen Loop: User-Spec-Driven Development for a Self-Evolving Codebase
- 5. Natural-Language Agent Harnesses
- 6. Agent Factories for High Level Synthesis: How Far Can General-Purpose Coding Agents Go in Hardware Optimization?
- 7. Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
Lumen Research Digest — 2026-03-27
2026-03-27 · 7 items
- 1. Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
- 2. Vega: Learning to Drive with Natural Language Instructions
- 3. Is Mathematical Problem-Solving Expertise in Large Language Models Associated with Assessment Performance?
- 4. LanteRn: Latent Visual Structured Reasoning
- 5. R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- 6. Missing-Aware Multimodal Fusion for Unified Microservice Incident Management
- 7. RefAlign: Representation Alignment for Reference-to-Video Generation