Large Language Models (LLMs) have recently shown strong reasoning capabilities, motivating their use in complex decision-making environments. StarCraft II (SC2), with its massive state-action space and partial observability, is a challenging testbed. However, existing LLM-based SC2 agents primarily focus on improving the policy itself, leaving the integration of a learnable, action-conditioned dynamics model into the decision loop largely unexplored. In this work, we propose StarWM and StarWM-Agent, and conduct the first systematic study of the learnability and decision utility of player-view, action-conditioned textual world models for SC2. StarWM predicts short-horizon future observations under partial observability. StarWM-Agent integrates StarWM into a lightweight Generate-Simulate-Refine loop for foresight-driven policy refinement. Extensive experiments show that StarWM substantially outperforms zero-shot baselines across multiple dimensions, while StarWM-Agent achieves consistent win-rate gains of 30%, 15%, and 30% against the SC2 built-in AI at Hard (LV5), Harder (LV6), and VeryHard (LV7), respectively. Additional analyses show that our method is complementary to existing history-summarization SC2-agent approaches, with their combination achieving the strongest online performance in our evaluation.
World Models for Policy Refinement in StarCraft II
Large Language Models (LLMs) have recently shown strong reasoning capabilities, motivating their use in complex decision-making environments. StarCraft II (SC2), with its massive state-action space and partial observability, is a challenging testbed.
- Preview

- Year
- 2026
- Venue
- arXiv 2026
- Stars
- 10
- Authors
- 9
- Hosting
- Abstract onlyARXIV-DEFAULT
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2602.14857ARXIV-DEFAULT
- TL;DR
- Semantic Scholar