datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
autotrain-data-1w6s-u4vt-i7yo
Dataset Card for "autotrain-data-1w6s-u4vt-i7yo"
More Information needed
autotrain-wider-faceautotrainer-v0
autotrainer-v0
AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown
Dataset Info
Rows: 1
Columns: 1
Columns
Column
Type
Description
state_json
Value('string')
Full autotrainer state as JSON string
Generation Parameters
{
"script_name": "run_round.py",
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"description": "AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown",
"experiment_id": "autotrainer-v0"… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/autotrainer-v0.autotrain-data-gtzs-bj3r-fz0k
Dataset Card for "autotrain-data-gtzs-bj3r-fz0k"
More Information needed
autotrain-data-Medical_Terminology_Zephyr
Dataset Card for "autotrain-data-Medical_Terminology_Zephyr"
More Information needed
autotrain-data-Medical_Terminology_Zephyr_2
Dataset Card for "autotrain-data-Medical_Terminology_Zephyr_2"
More Information needed
nyxmed-autotrain-v2autotrain-data-zesty
Dataset Card for "autotrain-data-zesty"
More Information needed
corporate-speak-dataset-autotrainautotrainer-v1
AutoTrainer v1
Agentic training harness where Claude decides every training step's method, data, and hyperparameters.
Configs
steps: Per-step decisions, metrics, agent traces, and costs
eval_traces: Per-question evaluation results with difficulty breakdown (n_args=2-10)
autotrainer-v1-run-v2-11stepscode-review-python-autotrain
Python Code Review Dataset
Filtered and formatted version of ronantakizawa/github-codereview for fine-tuning code review models.
Dataset Summary
This dataset contains Python code snippets with corresponding review comments, formatted as conversations for instruction tuning.
Splits
Split
Samples
train
~40,000
validation
~800
test
~800
Format
Each sample contains a messages column with conversation format:
{
"messages": [… See the full description on the dataset page: https://huggingface.co/datasets/PrathamKotian26/code-review-python-autotrain.autotrain-data-gabba_v2autotrain-data-test
Dataset Card for "autotrain-data-test"
More Information needed
autotrain-data-autotrain-9owpg-gq2mv
Dataset Card for "autotrain-data-autotrain-9owpg-gq2mv"
More Information needed
autotrain-data-dfun-lk90-yhtx
Dataset Card for "autotrain-data-dfun-lk90-yhtx"
More Information needed
autotrain-data-2288348c0731
Dataset Card for "autotrain-data-2288348c0731"
More Information needed
autotrain-data-8a00-9wrj-ig4m
Dataset Card for "autotrain-data-8a00-9wrj-ig4m"
More Information needed
autotrain-data-Nuclear_Fusion_Falcon
Dataset Card for "autotrain-data-Nuclear_Fusion_Falcon"
More Information needed
qwen-sft-huanhuan-autotrainordis-autotrain-pure-skill
Ordis Pure Skill Features for AutoTrain
Dataset for boat racing prediction using only skill-based features (no odds).
Files
train.parquet: Training data (21,684 samples)
valid.parquet: Validation data (28,316 samples)
Features
221 pure skill features
Binary target: label (1=top-3 finish, 0=other)
Usage with AutoTrain
Go to https://huggingface.co/autotrain
Create new project → Tabular → Binary Classification
Use dataset:… See the full description on the dataset page: https://huggingface.co/datasets/sugiken/ordis-autotrain-pure-skill.autotrainer-v0-eval-tracesautotrain-data-autotrain-gakg2-0flbd
Dataset Card for "autotrain-data-autotrain-gakg2-0flbd"
More Information needed
autotrain-data-test-data
Dataset Card for "autotrain-data-test-data"
More Information needed
autotrain-data-DistilBert-500-500
Dataset Card for "autotrain-data-DistilBert-500-500"
More Information needed
autotrain-data-roblox-usernames
Dataset Card for "autotrain-data-roblox-usernames"
More Information needed
autotrainer-v0-roundsautotrain-dataset-ed1f9e0eautotrain-data-autotrain-0e0vl-kem7f
Dataset Card for "autotrain-data-autotrain-0e0vl-kem7f"
More Information needed
autotrain-data-autotrain-7u119-vc77x
