datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
autotrain-data-numai2autotrainer-v0
autotrainer-v0
AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown
Dataset Info
Rows: 1
Columns: 1
Columns
Column
Type
Description
state_json
Value('string')
Full autotrainer state as JSON string
Generation Parameters
{
"script_name": "run_round.py",
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"description": "AutoTrainer-v0: LLM-agent-controlled GRPO training on Countdown",
"experiment_id": "autotrainer-v0"… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/autotrainer-v0.autotrain-data-numaitrthminh1112__autotrain-llama32-1b-finetune-details
Dataset Card for Evaluation run of trthminh1112/autotrain-llama32-1b-finetune
Dataset automatically created during the evaluation run of model trthminh1112/autotrain-llama32-1b-finetune
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/trthminh1112__autotrain-llama32-1b-finetune-details.autotrainer-v1
AutoTrainer v1
Agentic training harness where Claude decides every training step's method, data, and hyperparameters.
Configs
steps: Per-step decisions, metrics, agent traces, and costs
eval_traces: Per-question evaluation results with difficulty breakdown (n_args=2-10)
autotrainer-v1-run-v2-11stepsautotrain-data-8a00-9wrj-ig4m
Dataset Card for "autotrain-data-8a00-9wrj-ig4m"
More Information needed
autotrain-data-Nuclear_Fusion_Falcon
Dataset Card for "autotrain-data-Nuclear_Fusion_Falcon"
More Information needed
ordis-autotrain-pure-skill
Ordis Pure Skill Features for AutoTrain
Dataset for boat racing prediction using only skill-based features (no odds).
Files
train.parquet: Training data (21,684 samples)
valid.parquet: Validation data (28,316 samples)
Features
221 pure skill features
Binary target: label (1=top-3 finish, 0=other)
Usage with AutoTrain
Go to https://huggingface.co/autotrain
Create new project → Tabular → Binary Classification
Use dataset:… See the full description on the dataset page: https://huggingface.co/datasets/sugiken/ordis-autotrain-pure-skill.autotrainer-v0-eval-tracesabhishek__autotrain-llama3-70b-orpo-v2-details
Dataset Card for Evaluation run of abhishek/autotrain-llama3-70b-orpo-v2
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-70b-orpo-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-llama3-70b-orpo-v2-details.autotrain-data-ai-explains-codeautotrainer-v1-v2autotrain-data-imagineabhishek__autotrain-llama3-orpo-v2-details
Dataset Card for Evaluation run of abhishek/autotrain-llama3-orpo-v2
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-orpo-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-llama3-orpo-v2-details.abhishek__autotrain-llama3-70b-orpo-v1-details
Dataset Card for Evaluation run of abhishek/autotrain-llama3-70b-orpo-v1
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-70b-orpo-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-llama3-70b-orpo-v1-details.abhishek__autotrain-vr4a1-e5mms-details
Dataset Card for Evaluation run of abhishek/autotrain-vr4a1-e5mms
Dataset automatically created during the evaluation run of model abhishek/autotrain-vr4a1-e5mms
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-vr4a1-e5mms-details.abhishek__autotrain-llama3-70b-orpo-v1autotrain-data-amdal-mining-llama2-7b-cleanabhishek__autotrain-0tmgq-5tpbg-details
Dataset Card for Evaluation run of abhishek/autotrain-0tmgq-5tpbg
Dataset automatically created during the evaluation run of model abhishek/autotrain-0tmgq-5tpbg
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-0tmgq-5tpbg-details.mida-autotrain2
RuTurboAlpaca
Dataset of ChatGPT-generated instructions in Russian.
Code: rulm/self_instruct
Code is based on Stanford Alpaca and self-instruct.
29822 examples
Preliminary evaluation by an expert based on 400 samples:
83% of samples contain correct instructions
63% of samples have correct instructions and outputs
Crowdsouring-based evaluation on 3500 samples:
90% of samples contain correct instructions
68% of samples have correct instructions and outputs
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/t4t455/mida-autotrain2.autotrain-data-ai-explains-codedetails_abhishek__autotrain-llama3-70b-orpo-v1
Dataset Card for Evaluation run of abhishek/autotrain-llama3-70b-orpo-v1
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-70b-orpo-v1.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_abhishek__autotrain-llama3-70b-orpo-v1.abhishek__autotrain-llama3-orpo-v2autotrain-data-qaaaaautotrain-data-address-parsingautotrain-data-isensor-on-xdr-playbook-for-threat-huntimport pandas as pd
# Load the dataset
df = pd.read_csv('cyber_security_breaches.csv')
# Print the first 14 rows of the dataset
print(df.head())
# Get the number of rows and columns in the dataset
print(df.shape)
# Get the summary statistics of the dataset
print(df.describe())
# Get the unique values of a column
print(df['Year'].unique())
# Filter the dataset based on a condition
print(df[df['Year'] == 2023])
language: English
"isensor XDR for Threat Hunt in Computational… See the full description on the dataset page: https://huggingface.co/datasets/Nyameri/autotrain-data-isensor-on-xdr-playbook-for-threat-hunt.details_abhishek__autotrain-llama3-70b-orpo-v2
Dataset Card for Evaluation run of abhishek/autotrain-llama3-70b-orpo-v2
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-70b-orpo-v2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_abhishek__autotrain-llama3-70b-orpo-v2.autotrain-data-inflation_pred_rusauto-train-test-cr-ad
