CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1k downloads1y agoHugging Face02hugosousa /TimeQA TimeQA Check out the original GitHub repo to learn more about the dataset. textquestion-answering10K<n<100K5 likes484 downloads3y agoHugging Face03hugosousa /professor_heideltime_en Professor HeidelTime Professor HeidelTime is a project to create a multilingual corpus weakly labeled with HeidelTime, a temporal tagger. Corpus Details The weak labeling was performed in six languages. Here are the specifics of the corpus for each language: Dataset Language Documents From To Tokens Timexs All the News 2.0 EN 24,642 2016-01-01 2020-04-0218,755,616 254,803 Italian Crime News IT 9,619 2011-01-01 2021-12-31 3,296,898 58,823 German News… See the full description on the dataset page: https://huggingface.co/datasets/hugosousa/professor_heideltime_en.texttoken-classification100K<n<1M0 likes131 downloads3y agoHugging Face04hugosousa /Publico Público This dataset was build by translating a set of 34,157 news from Público, an European Portuguese news paper. The news have been translated using Google Translator. To now more about the data visit the Github repos used to scrape and translate the news. text100K<n<1M1 likes44 downloads3y agoHugging Face05hugomautner /n-gramstext1M<n<10M0 likes34 downloads14d agoHugging Face06hugosousa /ProfessorHeidelTime Professor HeidelTime Paper GitHub Professor HeidelTime is a project to create a multilingual corpus weakly labeled with HeidelTime, a temporal tagger. Corpus Details The weak labeling was performed in six languages. Here are the specifics of the corpus for each language: Dataset Language Documents From To Tokens Timexs All the News 2.0 EN 24,642 2016-01-01 2020-04-02 18,755,616 254,803 Italian Crime News IT 9,619 2011-01-01 2021-12-31 3,296,898 58,823… See the full description on the dataset page: https://huggingface.co/datasets/hugosousa/ProfessorHeidelTime.texttoken-classification100K<n<1M2 likes31 downloads3y agoHugging Face07hugo74130 /sm64-tas-dataset SM64 Speedrun / TAS Reasoning Dataset Question → <think> reasoning → answer pairs about Super Mario 64 speedrunning and Tool-Assisted Speedruns (TAS), in ShareGPT format. Each assistant turn contains an explicit reasoning trace inside <think>...</think> followed by the final answer, matching the native thinking format of Qwen3-style models. Files File Rows Use dataset_v11.jsonl 2721 full dataset dataset_v11_train.jsonl 2585 training split (95%)… See the full description on the dataset page: https://huggingface.co/datasets/hugo74130/sm64-tas-dataset.textquestion-answering1K<n<10K1 likes30 downloads3mo agoHugging Face08hugoramallo /legal-ai-act-spanish-sft-7k⚠️ Legal and Liability Disclaimer This dataset is provided for research and educational purposes only. It does not constitute legal advice, nor does it represent an official or authoritative interpretation of Regulation (EU) 2024/1689 (EU AI Act). The content is synthetically generated and may contain errors, omissions, or hallucinations. Under no circumstances should this dataset be used as a basis for legal, compliance, or regulatory decision-making. The authors disclaim any liability for… See the full description on the dataset page: https://huggingface.co/datasets/hugoramallo/legal-ai-act-spanish-sft-7k.textquestion-answering1K<n<10K0 likes29 downloads6mo agoHugging Face09HugoZhu /PerCN PerCN Dataset Overview PerCN is a Chinese dataset for MBTI personality type prediction. Each sample contains multiple short posts from the same user, and labels are a 4-d binary vector corresponding to the four MBTI dimensions. The texts include typical Chinese social media expressions, emojis, and colloquial phrasing. Data Format The dataset is provided in JSONL format with three splits: train.jsonl, eval.jsonl, and test.jsonl. Each line is a JSON object with… See the full description on the dataset page: https://huggingface.co/datasets/HugoZhu/PerCN.texttext-classification10K<n<100K0 likes27 downloads8mo agoHugging Face10Hugomattsson /selma_ngramstext1M<n<10M0 likes22 downloads1d agoHugging Face11HugoTer /ngramstext100K<n<1M0 likes17 downloads2d agoHugging Face12hugoelec /folbar-test2text1K<n<10K0 likes8 downloads2y agoHugging Face13Hugodonotexit /filtered-awesome-chatgpt-propmts-oss-120b Filtered Awesome ChatGPT Prompts – Model Outputs Dataset Overview This dataset contains model-generated responses to prompts from the fka/awesome-chatgpt-prompts Hugging Face dataset. Each prompt was sent to the openai/gpt-oss-120b model via the OpenRouter API. The resulting dataset was then filtered to remove: Non English outputs with high language-detection confidence (fastText score < 0.7) Very short outputs (≤ 10 words) The goal of this dataset is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/filtered-awesome-chatgpt-propmts-oss-120b.texttext-generationn<1K1 likes8 downloads8mo agoHugging Face14hugosousa /WikiTimelinestext10K<n<100K0 likes7 downloads3y agoHugging Face15hugoelec /foolbar-llama3-tttext1K<n<10K0 likes7 downloads2y agoHugging Face16hugoclm /projetE3train and test for parkinson diseases textn<1K0 likes2 downloads2y agoHugging Face17hugochien /testdataset2tabularn<1K0 likes1 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.