PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
25 Aug 2026
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models generally inherit pretrained representations without using this contextual capacity…
30.1/h