Xuyang Chen, Keyu Yan, Guojian Wang, Lin Zhao
Official PyTorch implementation
This repository contains the official implementation of VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning.
VIPO is a model-based offline reinforcement learning method that improves dynamics model learning by introducing a value function inconsistency penalty. Instead of relying only on heuristic uncertainty estimation, VIPO uses the discrepancy between value estimates from offline data and model-generated rollouts as an additional self-supervised signal for model training.
We are still in the process of updating and organizing this repository. The current version includes the core implementation, while some configuration files, scripts, and usage instructions are still being added.
[2026-05]Initial repository release.
Install D4RL and pytorch.
VIPO is integrated into OfflineRL-Kit and LEQ, see guidance.
Our implementation is developed with reference to several excellent open-source repositories, including:
We sincerely thank the authors and maintainers of these repositories for their valuable contributions to the offline reinforcement learning community.
If you find this work useful, please consider citing our paper:
@misc{chen2026vipovaluefunctioninconsistency,
title={VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning},
author={Xuyang Chen and Keyu Yan and Guojian Wang and Lin Zhao},
year={2026},
eprint={2504.11944},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2504.11944},
}
