Tencent Kaiwu · Reinforcement Learning
A multi-objective hierarchical PPO agent for a partially observable survival game, balancing survival, collection, and exploration.
- My role
- Team lead, reinforcement-learning design, and training engineering
- Methods & tools
- PPO · Actor-Critic · Hierarchical policy · LSTM / Transformer Policy
01
Problem & challenge
The agent must evade pursuit and collect resources in a complex map, bringing partial observability, conflicting objectives, sparse rewards, and long-horizon planning into one task.
02
Key contributions
Designed hierarchical actors and multi-head critics to separate movement, skill use, and objective-specific value estimates.
Fused local observations, target relations, and state history with an LSTM / Transformer policy for short- and long-horizon context.
Designed multi-objective rewards for survival, resource collection, and exploration, with staged reward shaping for sparse feedback and conflicting goals.
Iterated sampling, evaluation, and training configurations, using failure replays to diagnose policy collapse and insufficient exploration.
03
Outcome & disclosure boundary
The solution placed near the top of the regional qualifier and received a national finals third prize; the public repository documents the network, rewards, and training design.
View on GitHub