Deferred from the OnnxRLAdapter session: a training-side env wrapper as a separate shinro.envs module — deliberately out of the deploy/inference path.
Would let RL stacks (sb3, CleanRL, ...) train against shinro plants without polluting the deploy package.
Priority: P4 — deferred.
Deferred from the OnnxRLAdapter session: a training-side env wrapper as a separate shinro.envs module — deliberately out of the deploy/inference path.
Would let RL stacks (sb3, CleanRL, ...) train against shinro plants without polluting the deploy package.
Priority: P4 — deferred.