XVIS / Embodied AI

Policy monitor

While a policy trains or runs, what should be on the screen: the envs, the reward terms, or what the policy looks at?

Input
64 vectorized cart-pole envs, shaped reward terms, patch attention
Output
Env grid + training curve, reward decomposition, attention heatmap
Demos
3 · all run in the browser
Loading engine
Grid + curve
  • Cart
  • Pole (alive)
  • Pole (terminated)
  • Mean return

Built with Parallel envs (64)

64 cart-pole environments stepped in lockstep; a linear policy improves by random-search hill climbing on the mean return. The grid view is what Isaac Lab / Viser show, at toy scale.

  • Canvas 2DRasterized plots and overlays
  • Isaac LabReference for parallel RL training
  • ViserReference for multi-env web viewers