Lumen Research Digest — 2026-03-27
A selective scan of cutting-edge work across AI, automation, graphics, and computer science. This is ranked for novelty and likely significance rather than simply recency.
Big picture
- Agentic and reasoning-heavy systems continue to dominate the high-signal end of AI work.
- Graphics and generative visual research is pushing toward real-time, high-fidelity interactive pipelines.
- Systems work remains tightly coupled to model usefulness through inference, scale, and tooling efficiency.
Selected items
1. Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
- Source: arXiv
- Published: 2026-03-26T17:59:59Z
- Why it matters: arXiv: Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting. Existing feed-forward 3D Gaussian Splatting methods predict pixel-aligned primitives, leading to a quadratic growth in primitive count as resolution increases. This fundamentally limits their scalability, making high-resolution synthesis such as 4K intractable. We introduce LGTM (Less Gaussians, Texture…
- Take: Existing feed-forward 3D Gaussian Splatting methods predict pixel-aligned primitives, leading to a quadratic growth in primitive count as resolution increases. This fundamentally limits their scalability, making high-resolution synthesis such as 4K intractable. We introduce LGTM (Less Gaussians, Texture More), a feed-forward framework that overcomes this resolution scaling barrier. By predicting compact Gaussian primitives coupled with per-primitive textures, LGTM decouples geometric complexity from rendering resolution. This approach enables high-fidelity 4K novel view synthesis without per-scene optimization, a capability previously out of reach for feed-forward methods, all while using significantly fewer Gaussian primitives. Project page: https://yxlao.github.io/lgtm/
- Link: https://arxiv.org/abs/2603.25745v1
- PDF: https://arxiv.org/pdf/2603.25745v1
2. Vega: Learning to Drive with Natural Language Instructions
- Source: arXiv
- Published: 2026-03-26T17:59:56Z
- Why it matters: arXiv: Vega: Learning to Drive with Natural Language Instructions. Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process. However, most existing pipelines only utilize the language modality for scene descriptions or reasoning and lack the flexibility to follow diverse user instructions for personalized driving. To…
- Take: Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process. However, most existing pipelines only utilize the language modality for scene descriptions or reasoning and lack the flexibility to follow diverse user instructions for personalized driving. To address this, we first construct a large-scale driving dataset (InstructScene) containing around 100,000 scenes annotated with diverse driving instructions with the corresponding trajectories. We then propose a unified Vision-Language-World-Action model, Vega, for instruction-based generation and planning. We employ the autoregressive paradigm to process visual inputs (vision) and language instructions (language) and the diffusion paradigm to generate future predictions (world modeling) and trajectories (action). We perform joint attention to enable interactions between the…
- Link: https://arxiv.org/abs/2603.25741v1
- PDF: https://arxiv.org/pdf/2603.25741v1
3. Is Mathematical Problem-Solving Expertise in Large Language Models Associated with Assessment Performance?
- Source: arXiv
- Published: 2026-03-26T16:43:54Z
- Why it matters: arXiv: Is Mathematical Problem-Solving Expertise in Large Language Models Associated with Assessment Performance?. Large Language Models (LLMs) are increasingly used in math education not only as problem solvers but also as assessors of learners' reasoning. However, it remains unclear whether stronger math problem-solving ability is associated with stronger step-level assessment…
- Take: Large Language Models (LLMs) are increasingly used in math education not only as problem solvers but also as assessors of learners' reasoning. However, it remains unclear whether stronger math problem-solving ability is associated with stronger step-level assessment performance. This study examines that relationship using the GSM8K and MATH subsets of PROCESSBENCH, a human-annotated benchmark for identifying the earliest erroneous step in mathematical reasoning. We evaluate two LLM-based math tutor agent settings, instantiated with GPT-4 and GPT-5, in two independent tasks on the same math problems: solving the original problem and assessing a benchmark-provided solution by predicting the earliest erroneous step. Results show a consistent within-model pattern: assessment accuracy is substantially higher on math problem items the same model solved correctly than on items it solved…
- Link: https://arxiv.org/abs/2603.25633v1
- PDF: https://arxiv.org/pdf/2603.25633v1
4. LanteRn: Latent Visual Structured Reasoning
- Source: arXiv
- Published: 2026-03-26T16:41:59Z
- Why it matters: arXiv: LanteRn: Latent Visual Structured Reasoning. While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, most LMMs default to verbalizing perceptual content into text, a strong limitation for tasks requiring fine-grained spatial and visual understanding. While recent approaches take steps…
- Take: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, most LMMs default to verbalizing perceptual content into text, a strong limitation for tasks requiring fine-grained spatial and visual understanding. While recent approaches take steps toward thinking with images by invoking tools or generating intermediate images, they either rely on external modules, or incur unnecessary computation by reasoning directly in pixel space. In this paper, we introduce LanteRn, a framework that enables LMMs to interleave language with compact latent visual representations, allowing visual reasoning to occur directly in latent space. LanteRn augments a vision-language transformer with the ability to generate and attend to continuous visual thought embeddings during inference. We train the model in two stages:…
- Link: https://arxiv.org/abs/2603.25629v1
- PDF: https://arxiv.org/pdf/2603.25629v1
5. R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- Source: arXiv
- Published: 2026-03-26T17:58:04Z
- Why it matters: arXiv: R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning. Robust perception and reasoning require consistency across sensory modalities. Yet current multimodal models often violate this principle, yielding contradictory predictions for visual and textual representations of the same concept. Rather than masking these failures with standard voting…
- Take: Robust perception and reasoning require consistency across sensory modalities. Yet current multimodal models often violate this principle, yielding contradictory predictions for visual and textual representations of the same concept. Rather than masking these failures with standard voting mechanisms, which can amplify systematic biases, we show that cross-modal inconsistency provides a rich and natural signal for learning. We introduce RC2, a reinforcement learning framework that resolves internal conflicts by enforcing cross-modal cycle consistency. By requiring a model to perform backward inference, switch modalities, and reliably reconstruct the answer through forward inference, we obtain a dense, label-free reward. This cyclic constraint encourages the model to align its internal representations autonomously. Optimizing for this structure mitigates modality-specific errors and…
- Link: https://arxiv.org/abs/2603.25720v1
- PDF: https://arxiv.org/pdf/2603.25720v1
6. Missing-Aware Multimodal Fusion for Unified Microservice Incident Management
- Source: arXiv
- Published: 2026-03-26T15:14:57Z
- Why it matters: arXiv: Missing-Aware Multimodal Fusion for Unified Microservice Incident Management. Automated incident management is critical for microservice reliability. While recent unified frameworks leverage multimodal data for joint optimization, they unrealistically assume perfect data completeness. In practice, network fluctuations and agent failures frequently cause missing modalities.…
- Take: Automated incident management is critical for microservice reliability. While recent unified frameworks leverage multimodal data for joint optimization, they unrealistically assume perfect data completeness. In practice, network fluctuations and agent failures frequently cause missing modalities. Existing approaches relying on static placeholders introduce imputation noise that masks anomalies and degrades performance. To address this, we propose ARMOR, a robust self-supervised framework designed for missing modality scenarios. ARMOR features: (i) a modality-specific asymmetric encoder that isolates distribution disparities among metrics, logs, and traces; and (ii) a missing-aware gated fusion mechanism utilizing learnable placeholders and dynamic bias compensation to prevent cross-modal interference from incomplete inputs. By employing self-supervised auto-regression with…
- Link: https://arxiv.org/abs/2603.25538v1
- PDF: https://arxiv.org/pdf/2603.25538v1
7. RefAlign: Representation Alignment for Reference-to-Video Generation
- Source: arXiv
- Published: 2026-03-26T17:59:57Z
- Why it matters: arXiv: RefAlign: Representation Alignment for Reference-to-Video Generation. Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on. In practice, existing R2V methods typically introduce additional…
- Take: Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on. In practice, existing R2V methods typically introduce additional high-level semantic or cross-modal features alongside the VAE latent representation of the reference image and jointly feed them into the diffusion Transformer (DiT). These auxiliary representations provide semantic guidance and act as implicit alignment signals, which can partially alleviate pixel-level information leakage in the VAE latent space. However, they may still struggle to address copy--paste artifacts and multi-subject confusion caused by modality mismatch across heterogeneous encoder features. In this paper, we propose RefAlign, a representation alignment framework that…
- Link: https://arxiv.org/abs/2603.25743v1
- PDF: https://arxiv.org/pdf/2603.25743v1
Coverage notes
- Candidates considered: 56
- Sources included: arXiv plus selected research/lab/blog feeds.
- Selection policy: novelty, likely downstream importance, and technical substance.