datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously.
GPT-4-Self-Instruct-TurkishAs per the community's request, here we share a Turkish dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Turkish. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Turkish.GPT-4-Self-Instruct-JapaneseHere we share a Japanese dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Japanese. This dataset will be updated continuously.
GPT-4-Self-Instruct-GreekAs per the community's request, here we share a Greek dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied.
We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Greek. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Greek.CausalLM__14B-details
Dataset Card for Evaluation run of CausalLM/14B
Dataset automatically created during the evaluation run of model CausalLM/14B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__14B-details.CausalLM__34b-beta-details
Dataset Card for Evaluation run of CausalLM/34b-beta
Dataset automatically created during the evaluation run of model CausalLM/34b-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__34b-beta-details.CausalLM__preview-1-hf-details
Dataset Card for Evaluation run of CausalLM/preview-1-hf
Dataset automatically created during the evaluation run of model CausalLM/preview-1-hf
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__preview-1-hf-details.Refined-Anime-Text
Refined Anime Text for Continual Pre-training of Language Models
This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Refined-Anime-Text.Retrieval-SFT-Chat
Retrieval-Based Multi-Turn Chat SFT Synthetic Data
A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture.
In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.
