datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenHermes-2.5-Autotrain-SFT
This is the converted OpenHermes 2.5 dataset, available here: teknium/OpenHermes-2.5
All credit goes to teknium for creating the original dataset.
This version has been specifically formatted for training large language models (LLMs) using HuggingFace AutoTrain.
The dataset now contains a single text column, optimized for the LLM SFT training method. You can find other versions of the dataset in my repository as well.
I have filtered the dataset in various ways.
For example, if you're not… See the full description on the dataset page: https://huggingface.co/datasets/MugenYume/OpenHermes-2.5-Autotrain-SFT.trthminh1112__autotrain-llama32-1b-finetune-details
Dataset Card for Evaluation run of trthminh1112/autotrain-llama32-1b-finetune
Dataset automatically created during the evaluation run of model trthminh1112/autotrain-llama32-1b-finetune
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/trthminh1112__autotrain-llama32-1b-finetune-details.autotrain_qa_neuroautotrain-data-test1autotrain-data-llama-autotrainauto-train-1-1776291068490abhishek__autotrain-llama3-70b-orpo-v2-details
Dataset Card for Evaluation run of abhishek/autotrain-llama3-70b-orpo-v2
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-70b-orpo-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-llama3-70b-orpo-v2-details.autotrain-data-test-v2autotrain-poumpourasautotrain-data-dialect-translationautotrain-data-rekomendasi-mata-kuliahauto-train-1-1776353045609autotrain-data-ruzrtabhishek__autotrain-llama3-orpo-v2-details
Dataset Card for Evaluation run of abhishek/autotrain-llama3-orpo-v2
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-orpo-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-llama3-orpo-v2-details.abhishek__autotrain-llama3-70b-orpo-v1-details
Dataset Card for Evaluation run of abhishek/autotrain-llama3-70b-orpo-v1
Dataset automatically created during the evaluation run of model abhishek/autotrain-llama3-70b-orpo-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-llama3-70b-orpo-v1-details.abhishek__autotrain-vr4a1-e5mms-details
Dataset Card for Evaluation run of abhishek/autotrain-vr4a1-e5mms
Dataset automatically created during the evaluation run of model abhishek/autotrain-vr4a1-e5mms
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-vr4a1-e5mms-details.autotrain-data-amdal-mining-llama2-7b-cleanautotrain-dataset-2autotrain-1o6jf-k6pnmabhishek__autotrain-0tmgq-5tpbg-details
Dataset Card for Evaluation run of abhishek/autotrain-0tmgq-5tpbg
Dataset automatically created during the evaluation run of model abhishek/autotrain-0tmgq-5tpbg
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abhishek__autotrain-0tmgq-5tpbg-details.auto-train-1-1774182339942auto-train-1-1774220619885mida-autotrain2
RuTurboAlpaca
Dataset of ChatGPT-generated instructions in Russian.
Code: rulm/self_instruct
Code is based on Stanford Alpaca and self-instruct.
29822 examples
Preliminary evaluation by an expert based on 400 samples:
83% of samples contain correct instructions
63% of samples have correct instructions and outputs
Crowdsouring-based evaluation on 3500 samples:
90% of samples contain correct instructions
68% of samples have correct instructions and outputs
Prompt template:… See the full description on the dataset page: https://huggingface.co/datasets/t4t455/mida-autotrain2.autotrain-data-writingautotrain-data-chinese-nerautotrain-data-52y0-943s-i8cyauto-train-1-1774285994483mida-autotrain
SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set)
Dataset Summary
SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/t4t455/mida-autotrain.autotrain-data-uo-codeautotrain-data-address-parsing
