Big ideas.
A little less reading.
The substance of every paper, without the sprawl.
Concise text. Complete visuals. Context one click away.
GaussVLA
Geometry-Aware Spatial Reasoning for Vision-Language-Action Model
A geometry-aware robot policy that turns visual features into 3D Gaussian tokens and compact spatial reasoning.
LGSR
Improving Policy Learning via Language-Guided State Representation in World Models
Let task language select the visual information retained by a compact world-model state.
DSEI
Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models
Compress sentences into adaptive semantic segments, then predict those segments in latent space.
DELE-w0.5
Inferring Action from Future Latent State for Robotic Manipulation
Train with future-state prediction; execute actions without generating future video.
HypoEvolve
Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
A genetic search coordinates scientific agents, with biological evidence reserved for evaluation.
CoMind
Understanding Collaborative Human Activity from Multiple Minds and Views
Synchronized views, gaze, dialogue, and 3D scans for studying collaborative intent.
GeomVLA
Unifying Scene, Motion, and Action in 3D
Keep scene features, predicted motion, and action generation in a shared 3D frame.
No papers match your search.