CartPole-v1
CartPole-v1
CartPole-v1 - Imitation Learning Datasets
This is a dataset created by Imitation Learning Datasets project.
It was created by using Stable Baselines weights from a PPO policy from HuggingFace.
Description
The dataset consists of 1,000 episodes with an average episodic reward of 500.
Each entry consists of:
obs (list): observation with length 4.
action (int): action (0 or 1).
reward (float): reward point for that timestep.
episode_returns (bool): if that state was the… See the full description on the dataset page: https://huggingface.co/datasets/NathanGavenski/CartPole-v1.ppo-CartPole-v1
Dataset Card for "ppo-CartPole-v1"
More Information needed
