CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CausalLM /GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously. texttext-generation10K<n<100K30 likes94 downloads2y agoHugging Face02CausalLM /GPT-4-Self-Instruct-TurkishAs per the community's request, here we share a Turkish dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Turkish. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Turkish.text1K<n<10K24 likes72 downloads2y agoHugging Face03CausalLM /GPT-4-Self-Instruct-JapaneseHere we share a Japanese dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Japanese. This dataset will be updated continuously. text1K<n<10K18 likes67 downloads2y agoHugging Face04CausalLM /GPT-4-Self-Instruct-GreekAs per the community's request, here we share a Greek dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Greek. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Greek.text1K<n<10K10 likes33 downloads2y agoHugging Face05open-llm-leaderboard /CausalLM__14B-detailsgated Dataset Card for Evaluation run of CausalLM/14B Dataset automatically created during the evaluation run of model CausalLM/14B The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__14B-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face06open-llm-leaderboard /CausalLM__34b-beta-detailsgated Dataset Card for Evaluation run of CausalLM/34b-beta Dataset automatically created during the evaluation run of model CausalLM/34b-beta The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__34b-beta-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face07open-llm-leaderboard /CausalLM__preview-1-hf-detailsgated Dataset Card for Evaluation run of CausalLM/preview-1-hf Dataset automatically created during the evaluation run of model CausalLM/preview-1-hf The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__preview-1-hf-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face08CausalLM /Refined-Anime-Textgated Refined Anime Text for Continual Pre-training of Language Models This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Refined-Anime-Text.texttext-generation1M<n<10M274 likes22 downloads2y agoHugging Face09CausalLM /Retrieval-SFT-Chatgated Retrieval-Based Multi-Turn Chat SFT Synthetic Data A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture. In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.textquestion-answering100K<n<1M61 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.