datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
autotrain-data-qa-team-car-review-project
AutoTrain Dataset for project: qa-team-car-review-project
Dataset Descritpion
This dataset has been automatically processed by AutoTrain for project qa-team-car-review-project.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": " ",
"target": 1
},
{
"text": " Mazda truck costs less than the sister… See the full description on the dataset page: https://huggingface.co/datasets/Gflorent/autotrain-data-qa-team-car-review-project.autotrain-data-mabama-dloraautotrain-data-javascript-traing-1
AutoTrain Dataset for project: javascript-traing-1
Dataset Description
This dataset has been automatically processed by AutoTrain for project javascript-traing-1.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"target": "test/NavbarSpec.js",
"feat_repo_name": "aabenoja/react-bootstrap",
"text": "import React from 'react';\nimport… See the full description on the dataset page: https://huggingface.co/datasets/ars-1/autotrain-data-javascript-traing-1.autotrain-data-quality-customer-reviews
AutoTrain Dataset for project: quality-customer-reviews
Dataset Descritpion
This dataset has been automatically processed by AutoTrain for project quality-customer-reviews.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": " Love this truck, I think it is light years better than the competition. I have driven or… See the full description on the dataset page: https://huggingface.co/datasets/Gflorent/autotrain-data-quality-customer-reviews.autotrain-sdxl-vintage-face-style-loraautotrain-data-clasificacion_pisicinas
AutoTrain Dataset for project: clasificacion_pisicinas
Dataset Description
This dataset has been automatically processed by AutoTrain for project clasificacion_pisicinas.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<11x10 RGB PIL image>",
"target": 1
},
{
"image": "<12x15 RGB PIL image>",
"target": 1
}]… See the full description on the dataset page: https://huggingface.co/datasets/Eredim/autotrain-data-clasificacion_pisicinas.autotrain-myanmar-kathein-festival-cartoonautotrain-data-nl-to-sqlautotrain-data-1w6s-u4vt-i7yo
Dataset Card for "autotrain-data-1w6s-u4vt-i7yo"
More Information needed
autotrain-north-atlantic-right-whale-lora-mk-1Munyarwanda-AI-AutoTrain
Munyarwanda AI - AutoTrain dataset
AutoTrain-ready version of arcange9/Munyarwanda-AI-Dataset v0.2.
Every example is pre-formatted in Qwen chat template as a single text column (5,542 train / 22 validation rows).
Intended recipe (Hugging Face AutoTrain, base model Qwen/Qwen3-0.6B):
LLM task, causal LM
text column: text
LoRA/PEFT + int4 quantization to fit free-tier GPUs
autotrain-custom-mistral-trainingautotrain-myanmar-ancient-maleautotrain-data-numai2autotrainer-v0
autotrainer-v0
AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown
Dataset Info
Rows: 1
Columns: 1
Columns
Column
Type
Description
state_json
Value('string')
Full autotrainer state as JSON string
Generation Parameters
{
"script_name": "run_round.py",
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"description": "AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown",
"experiment_id": "autotrainer-v0"… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/autotrainer-v0.autotrain-drone-humpback-whale-lora-1OpenHermes-2.5-Autotrain-SFT
This is the converted OpenHermes 2.5 dataset, available here: teknium/OpenHermes-2.5
All credit goes to teknium for creating the original dataset.
This version has been specifically formatted for training large language models (LLMs) using HuggingFace AutoTrain.
The dataset now contains a single text column, optimized for the LLM SFT training method. You can find other versions of the dataset in my repository as well.
I have filtered the dataset in various ways.
For example, if you're not… See the full description on the dataset page: https://huggingface.co/datasets/MugenYume/OpenHermes-2.5-Autotrain-SFT.autotrain-rosemarryautotrain-flx-fash-testautotrain-data-gtzs-bj3r-fz0k
Dataset Card for "autotrain-data-gtzs-bj3r-fz0k"
More Information needed
autotrain-data-numaiautotrain-data-Medical_Terminology_Zephyr
Dataset Card for "autotrain-data-Medical_Terminology_Zephyr"
More Information needed
autotrain-data-Medical_Terminology_Zephyr_2
Dataset Card for "autotrain-data-Medical_Terminology_Zephyr_2"
More Information needed
trthminh1112__autotrain-llama32-1b-finetune-details
Dataset Card for Evaluation run of trthminh1112/autotrain-llama32-1b-finetune
Dataset automatically created during the evaluation run of model trthminh1112/autotrain-llama32-1b-finetune
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/trthminh1112__autotrain-llama32-1b-finetune-details.autotrain-meme-tokennyxmed-autotrain-v2autotrain-data-zesty
Dataset Card for "autotrain-data-zesty"
More Information needed
corporate-speak-dataset-autotrainautotrainer-v1
AutoTrainer v1
Agentic training harness where Claude decides every training step's method, data, and hyperparameters.
Configs
steps: Per-step decisions, metrics, agent traces, and costs
eval_traces: Per-question evaluation results with difficulty breakdown (n_args=2-10)
autotrainer-v1-run-v2-11steps
