datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
atari_dqnKimi-K2-5000x-Code
Kimi K2 Instruct 0905 - Coding-focused distillation (5000 samples)
I used the Kimi K2 Instruct 0905 model to distill 5000 samples of only code-related prompts.
API from NVIDIA NIM and Fireworks AI
All samples are code related only!
Mix of all languages.
Free to use for anyone with apache 2.0 license.
atari_dqn_eps_20_4_NOVdqnagent-v0.1-dataset3Pong_DQN_1Image size width: 64 and height: 64
Game specifications:
CPU speed: 0.5
Player speed: 0.5
Ball speed: 0.75
Reward function: Basic (1, -1, 0, 0, 0)
Hyperparameters:
LR: 0.0001,
Anneal length: 1000000
Evaluation Results:
Games played : 100
Agent won : 83
Agent lost : 17
dqncode-datasetPong_DQN_humanatari_dqn_epsdiscord_scam_detectiondqnagent-v0.1-dataset2Pong_DQN_5Image size width: 64 and height: 64
Game specifications:
CPU speed: 0.5
Player speed: 0.5
Ball speed: 0.75
Reward function: Basic (1, -1, 0, 0, 0)
Hyperparameters:
LR: 0.0001,
Anneal length: 1000000
atari_dqn_eps_50dqnagent-v0.1-datasetbreakout_dqndqngpt-identityPong_DQN_2Image size width: 84 and height: 84
Game specifications:
CPU speed: 0.25
Player speed: 0.5
Ball speed: 0.75
Reward function: Basic (1, -1, 0, 0, 0)
Hyperparameters:
LR: 0.0001,
Anneal length: 1000000
Evaluation:
Agent Won: 44
Agent Lost: 56
Pong_DQN_4Image size width: 64 and height: 48
Game specifications:
CPU speed: 0.5
Player speed: 0.5
Ball speed: 0.75
Reward function: Basic (1, -1, 0, 0, 0)
Hyperparameters:
LR: 0.0001
Anneal length: 1000000
Evaluation:
Agent Won: 0
Agent Lost: 100
atari_dqn_eps_20_videosSLM-RL-DQNPong_DQN_3Image size width: 84 and height: 84
Game specifications:
CPU speed: 0.5
Player speed: 0.5
Ball speed: 0.75
Reward function: Basic (1, -1, 0, 0, 0)
Hyperparameters:
LR: 0.0001
Anneal length: 1000000
Evaluation:
Agent won: 33
Agent lost: 67
atari_dqn_eps_20dqnSpaceIhighway_dqn_attention_train012nva-toudou_shion
