NathanGavenski/MountainCar-v0
MountainCar-v0 - Imitation Learning Datasets This is a dataset created by Imitation Learning Datasets project. It was created by using Stable Baselines weights from a DQN policy from HuggingFace. Description The dataset consists of 1,000 episodes with an average episodic reward of -98.817. Each entry consists of: obs (list): observation with length 2. action (int): action (0 or 1). reward (float): reward point for that timestep. episode_returns (bool): if that… See the full description on the dataset page: https://huggingface.co/datasets/NathanGavenski/MountainCar-v0.
139
Update README.md
Upload teacher.jsonl
Update README.md
initial commit
