Policy monitor
While a policy trains or runs, what should be on the screen: the envs, the reward terms, or what the policy looks at?
Loading engine
Grid + curve
- Cart
- Pole (alive)
- Pole (terminated)
- Mean return
Built with Parallel envs (64)
64 cart-pole environments stepped in lockstep; a linear policy improves by random-search hill climbing on the mean return. The grid view is what Isaac Lab / Viser show, at toy scale.