Research
Learning Transferable Skills in Action RPGs via Directed Skill Graphs and Selective Adaptation
Overview Research area: Reinforcement learning for real-time game control, specifically hierarchical and modular skill decomposition for lifelong / continual learning. Technical level: Intermediate. T

- arXiv
- 2601.17923
- Published
- 2026-01-25
- Authors
- Ali Najar
AI summary
Overview
Research area: Reinforcement learning for real-time game control, specifically hierarchical and modular skill decomposition for lifelong / continual learning.
Technical level: Intermediate. The core idea is intuitive (split a hard control problem into small, reusable skills), but the paper assumes familiarity with RL basics such as DQN, curricula, transfer, and fine-tuning. No deep mathematical background is required.
Scope: The paper studies whether representing combat in Dark Souls III as a directed graph of five small skills, trained in a hierarchical curriculum, enables sample-efficient learning and lets a Phase 1 agent adapt to Phase 2 by fine-tuning only two of those skills.
What This Paper Is About
Training a single "monolithic" reinforcement learning policy to play a complex real-time action game is sample-inefficient and brittle: the same parameters must represent camera control, targeting, movement, dodging, and attack/heal decisions all at once. The paper asks whether instead decomposing combat into a directed graph of five narrowly scoped skills (camera, lock-on, movement, dodging, heal–attack decisions), trained one after another in a dependency-ordered curriculum, makes learning more efficient and makes later adaptation cheaper. The test case is Dark Souls III, where the agent fights the first boss, Iudex Gundyr, and then must transfer from the boss's Phase 1 to Phase 2.
Key Contributions
- The paper formulates Dark Souls III combat as a directed skill graph and instantiates a modular agent with five reusable skills: camera control (C), lock-on (L), movement (M), dodging (D), and a heal–attack decision policy (H).
- It proposes a hierarchical training protocol in which skills are trained sequentially along the dependency chain C → L → M → D → H, with upstream policies frozen while downstream skills are trained, improving sample efficiency by isolating narrow competencies.
- It demonstrates selective post-training under a Phase 1 to Phase 2 domain shift: keeping camera, lock-on, and movement fixed and fine-tuning only the phase-sensitive dodge and heal–attack policies.
- It provides ablation evidence showing which skills are necessary for the composed agent's performance and which remain useful across domains.
Main Findings
- Sample efficiency from the skill graph: A competitive Phase 1 policy was obtained with an overall interaction budget of approximately 230k steps. The paper reports that the atomic end-to-end baseline "gets no where close to learning a reliable combat behavior even after plenty of steps."
- The end-to-end baseline failed: The single monolithic DQN policy, observing the same 25-dimensional state and using a 16-action space, plateaued early (already by roughly 250k steps) and was stopped before the planned 500k steps due to wall-clock cost. Its learned behavior collapsed into a poor survival heuristic: locking on and repeatedly dodging backward without a reliable dodge policy or effective attack strategy. It achieved a 0.0% win rate in the Phase 1 table.
- Upstream skills learn quickly: In Figures 2, camera, lock-on, and movement reach near-maximal return quickly under the curriculum.
- Downstream skills are the hard part: Dodging requires precise timing and was the hardest component; the paper notes the maximum achievable return for the dodge policy under its shaping is approximately 10, with returns near 0 already corresponding to surviving a non-trivial duration. The heal–attack policy reaches a reasonable attack strategy quickly but reliable healing is hindered by structural data sparsity (healing opportunities are capped by the environment); its maximum return under the reward scale is approximately 15, and the agent won in fewer than half of trials.
- Ablations confirm specialization: Over 25 episodes per setting in Phase 1, randomizing both dodge and heal–attack gives a 0.0% win rate; randomizing only heal–attack (with dodge trained) gives 4.0%; randomizing only dodge (with heal–attack trained) gives 16.0%; with both trained, 44.0%. The paper attributes the higher win rate under a randomized dodge policy to a shift toward aggression, since ending the fight quickly is the only viable path when defense is unreliable.
- Zero-shot transfer to Phase 2 is partially successful: Without additional training, 33.3% win rate from mid-range starts and 12.5% from long-range starts. Transfer was probed at two engagement distances because the Phase 2 boss has higher health and damage but is less aggressive at close range.
- Selective fine-tuning recovers performance cheaply: Fine-tuning only the dodge and heal–attack policies (mid-range starts) raised the Phase 2 win rate to 52%, showing adaptation can be localized to a small subset of policies under a limited interaction budget.
- Dodge progress is visible in episode length: The dodge agent's mean episode length increased from roughly 150 to 300 steps over training, meaning it survived about twice as long as a random-dodge baseline even without access to healing.
Methodology in Plain English
- Interface: Rather than learning from pixels (explicitly not evaluated due to compute and engineering constraints), the agent reads a compact state directly from the game's process memory. Cheat Engine was used offline to locate variables such as position, pose, resources, lock status, and animation signals; during training these are read with the Python interface pyMeow. Cheat Engine is not used in the training loop. This yields a global state of dimension 25.
- Skill decomposition: Each skill sees only the slice of the state relevant to its responsibility, constructed by lightweight feature engineering. Camera sees the unit direction to the enemy, the camera direction vector, and the camera–target angle (7 dims). Lock-on sees the camera–target angle and lock status (2 dims). Movement sees player and enemy positions (6 dims). Dodge sees enemy animation ID and progress, both orientations, stamina, player HP, and distance (7 dims). Heal–attack sees enemy and player animation IDs and progress, stamina, player and enemy HP, both orientations, remaining Estus flasks, and distance (11 dims).
- Actions: Each skill has a small discrete action set built from game keybindings: camera 5 actions, lock-on 2, movement 9, dodge 2, heal–attack 3. Skills run concurrently (multi-threaded) and their outputs are merged into a single control signal applied at a fixed rate, approximating synchronous composition.
- Curriculum: Train C, freeze it, train L, freeze it, and so on to H. Freezing upstream skills constrains the reachable state distribution for later skills, reducing their exploration burden. The paper frames this as encouraging "cooperative specialization," where each skill optimizes its own objective while minimizing interference with established upstream competencies.
- Rewards: Each skill has its own hand-designed reward reflecting its narrow responsibility and using only generic combat variables, not boss-specific scripts. Camera is penalized by the camera–target angle (with a small bonus when alignment is within 0.6). Lock-on gets +1 for a valid lock and −1 otherwise. Movement is penalized proportionally to distance (weight 1/10). Dodge rewards being alive, penalizes HP loss, applies a terminal death penalty, and penalizes low stamina (weights 0.02, 5, 5, 0.05). Heal–attack trades HP change, enemy HP change, death, and success (weights 5, 15, 5, 5).
- Learning algorithm: A deliberately simple, widely used value-based baseline, DQN, was used for all skills, implemented with Stable-Baselines3 using default hyperparameters, a single environment instance, observation normalization via VecNormalize, learning rate 3×10⁻⁴, and batch size 256. The paper argues this conservative choice makes the test of the skill-graph idea stronger, since vanilla DQN is known to be brittle under non-stationarity and provides no explicit mechanism to prevent forgetting.
- Environment handling: The evaluation environment is the first boss encounter, Iudex Gundyr, with a total health pool of 1037 HP, split into Phase 1 (approximately 415 HP) and Phase 2 (approximately 500 HP). To keep the boss loaded and avoid expensive restarts, episodes are terminated early: Phase 1 episodes end once boss health falls below 622, and Phase 2 episodes end once boss HP falls below 60. Player HP and boss HP were normalized to [0,1] to keep reward and value scales comparable across phases. In-game death handling was disabled; the agent is treated as dead when normalized player HP falls below 0.05. A default Knight class character was used for all runs.
- Episode horizons and timing: Skill-specific maximum episode lengths were 128 steps for camera and movement, 64 for lock-on, 512 for dodge, 1024 for heal–attack, and 2048 for the end-to-end baseline. Actions are executed through key-press macros with fixed real-time delays: 0.1 s for camera, lock-on, and movement actions; 0.5 s for a dodge; 0.2 s for a light attack; 0.3 s for drinking Estus; 0.5 s for attempting to drink with zero flasks; and 0.1 s for idle. Replay buffer capacity for modular skills was set to approximately one-third of the corresponding training interaction budget; the end-to-end baseline used 100,000 transitions.
- Ablation protocol: A selected policy is replaced with a uniform random policy over its action space while all other skills are held fixed, providing a direct test of whether a skill is necessary and whether upstream skills remain useful when downstream skills are removed.
- Evaluation: Figures 2 and 3 average over five evaluation episodes at each 1k-step checkpoint, with 95% confidence intervals shaded. Win rates in Table 1 are measured over 25 episodes per setting.
Why This Matters
Impact on research. The paper argues that structuring agents around skill dependencies is a practical pathway toward evolving, continually learning agents in complex real-time environments. Its distinctive claim is not peak game performance but that a small, hand-specified skill graph plus a staged curriculum can deliver transferable behavior with a deliberately simple learner (vanilla DQN) and a modest interaction budget, and that adaptation under domain shift can be confined to the few skills that are actually phase-sensitive. This connects skill-graph representations, modular policies, curriculum learning, and continual RL under limited interaction budgets.
Real-world applications (potential, as motivated by the paper's framing):
- Real-time control systems where subproblems are heterogeneous and must run concurrently, such as robot locomotion combined with perception and decision-making, where decomposing control into narrow modules reduces interference.
- Continual learning in deployed systems that face non-stationary environments, where localized fine-tuning of a few components could avoid retraining an entire policy.
- Simulated training environments for games and agents where interaction is expensive and early termination thresholds can substitute for full episode resets.
- Benchmarking of transfer and adaptation methods, using phase transitions within an encounter as a controlled domain shift with different initialization distances.
Industry relevance. Game AI and simulation teams face exactly the constraints the paper targets: tight reaction loops, partial observability, long-horizon credit assignment, and coupled subproblems. The paper's result that targeted fine-tuning of just two skills recovers Phase 2 performance under a limited interaction budget is directly relevant to anyone maintaining AI agents across game patches, difficulty changes, or new content. The use of a process-memory state interface rather than pixels also reflects a common practical shortcut in game RL research, though the paper is explicit that pixel-based perception was not evaluated.
Future Directions
- Replace the memory-readout interface with pixel-based perception. The paper explicitly avoids pixel perception due to compute and engineering constraints, so whether the skill graph still helps with learned visual input is an open question.
- Generalize beyond one boss and one game. All experiments use a single encounter (Iudex Gundyr) in a single title (Dark Souls III). Whether the five-skill decomposition and the C → L → M → D → H dependency chain transfer to other bosses or games is not tested.
- Make healing learnable. Reliable healing was hindered by structural data sparsity, since the number of healing opportunities is capped by the environment (one flask in Phase 1 and two in Phase 2), making credit assignment hard for DQN. The paper leaves unresolved how to fix this.
- Scale to stronger learners and harder shifts. Because vanilla DQN was chosen deliberately as a conservative baseline with no mechanism to prevent forgetting, it is unclear how much better the approach would perform with methods designed for non-stationarity, or how it would handle shifts larger than a within-fight phase change.
- The paper does not report the number of training seeds, compute hardware, or detailed per-skill training budgets beyond the approximately 230k-step overall Phase 1 figure, so reproducibility at that level of detail remains unreported.
Target Audience
Researchers and practitioners in reinforcement learning, particularly those working on hierarchical RL, modular policies, curricula, transfer learning, and continual/lifelong learning. It is also relevant to game AI engineers who need agents that adapt to content or difficulty changes without full retraining. Readers should have basic familiarity with RL terminology such as policies, rewards, DQN, fine-tuning, and zero-shot transfer; no advanced mathematics is needed to follow the argument.
Authors’ abstract
Lifelong agents should expand their competence over time without retraining from scratch or overwriting previously learned behaviors. We investigate this in a challenging real-time control setting (Dark Souls III) by representing combat as a directed skill graph and training its components in a hierarchical curriculum. The resulting agent decomposes control into five reusable skills: camera control, target lock-on, movement, dodging, and a heal-attack decision policy, each optimized for a narrow responsibility. This factorization improves sample efficiency by reducing the burden on any single policy and supports selective post-training: when the environment shifts from Phase 1 to Phase 2, only a subset of skills must be adapted, while upstream skills remain transferable. Empirically, we find that targeted fine-tuning of just two skills rapidly recovers performance under a limited interaction budget, suggesting that skill-graph curricula together with selective fine-tuning offer a practical pathway toward evolving, continually learning agents in complex real-time environments.