CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Aeala /ShareGPT_Vicuna_unfiltered Dataset Card This is a reupload of this dataset that was further cleaned by gozfarb. text100K<n<1M52 likes10k downloads3y agoHugging Face02zlab-princeton /Vero-2.5M-unfiltered Vero-2.5M-unfiltered [!Note] This repository contains the full unfiltered dataset used to construct Vero-600k and Vero-1.6M, before question and answer filtering. Note that task categories are not balanced in this dataset. Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with vision-language models. This repository contains the Vero-2.5M-unfiltered dataset, a curation of 2.5M reinforcement learning samples… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/Vero-2.5M-unfiltered.image1M<n<10M1 likes2.5k downloads3mo agoHugging Face03Goekdeniz-Guelmez /Function_Calling_Unfilteredtext100K<n<1M4 likes2.2k downloads3y agoHugging Face04semran1 /synth-cc-unfilteredtext100M<n<1B1 likes1.5k downloads1y agoHugging Face05maxidl /FineNews-unfiltered FineNews WIP. Like FineWeb, but built from Common Crawl News instead of main web. For languages not listed as a split, check the data/ directory. For now, it contains the 2024-05 (May),-04 (April),-03 (March) dumps. This is the unfiltered version, with only URL filtering applied. Some initial stats Total number of documents: 35M Dump Number of docs Disk size (compressed) CC-NEWS-2024-05 11_715_084 11G CC-NEWS-2024-04 11_546_298 11G CC-NEWS-2024-03… See the full description on the dataset page: https://huggingface.co/datasets/maxidl/FineNews-unfiltered.texttext-generation10M<n<100M3 likes939 downloads2y agoHugging Face06mlfoundations-dev /bespokelabs-sky-t1-numina-amc-aime-subset-unfilteredtext1K<n<10K0 likes779 downloads2y agoHugging Face07HayatoHongoEveryonesAI /qa_verify_tir_5.9M_new_unfiltered_v1"HayatoHongoEveryonesAI/qa_verify_tir_2.9M_new_v1", "HayatoHongoEveryonesAI/qa_verify_1m_tir_3", "HayatoHongoEveryonesAI/qa_verify_1m_tir_4", "HayatoHongoEveryonesAI/qa_verify_1m_tir_5", tabular1M<n<10M0 likes333 downloads8mo agoHugging Face08HiTZ /Magpie-Llama-3.1-8B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-8B-Instruc with the MAGPIE codebase. The filtered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered System prompts used General <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n Code <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered.tabular1M<n<10M0 likes316 downloads1y agoHugging Face09HayatoHongoEveryonesAI /qa_verify_cot_new_6M_unfiltered_v7dataset_names = [ "HayatoHongoEveryonesAI/qa_verify_1m_cot_1", "HayatoHongoEveryonesAI/qa_verify_1m_cot_2", "HayatoHongoEveryonesAI/qa_verify_1m_cot_3", "HayatoHongoEveryonesAI/qa_verify_1m_cot_4", "HayatoHongoEveryonesAI/qa_verify_1m_cot_5", "HayatoHongoEveryonesAI/qa_verify_2m_cot_2", "HayatoHongoEveryonesAI/qa_verify_2m_cot_3", ] https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing tabular1M<n<10M0 likes313 downloads8mo agoHugging Face10mlfoundations-dev /PDF_and_SCP_unfiltered_organic_chemistry_questionstabular10K<n<100K0 likes305 downloads1y agoHugging Face11HiTZ /Magpie-Llama-3.1-70B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-70B-Instruc with the MAGPIE codebase. The filtered dataset can be found here: HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered System prompts used General <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n Code <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-70B-Instruct-Unfiltered.tabular1M<n<10M0 likes304 downloads1y agoHugging Face12marcov /trivia_qa_unfiltered_promptsourcetext100K<n<1M0 likes255 downloads2y agoHugging Face13QuixiAI /WizardLM_alpaca_evol_instruct_70k_unfilteredThis dataset is the WizardLM dataset victor123/evol_instruct_70k, removing instances of blatant alignment. 54974 instructions remain. inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py license: apache-2.0 language: - en pretty_name: wizardlm-unfiltered text10K<n<100K147 likes209 downloads3y agoHugging Face14learnanything /sharegpt_v3_unfiltered_cleaned_splittext10K<n<100K4 likes176 downloads3y agoHugging Face15drewparo /bigquery-swift-unfiltered GitHub Swift Repositories Dataset Description Dataset Summary This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license. Source Data Initial Data Collection and Normalization The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.tabulartext-generation100K<n<1M1 likes172 downloads3y agoHugging Face16hamishivi /tulu-3-unfiltered Tulu 3 Unfiltered This is an 'unfiltered' version of the Tulu 3 SFT mixture, created by collating the original Tulu 3 sources and avoiding downsampling. Details The dataset consists of a mix of : CoCoNot (ODC-BY-1.0) (Brahman et al., 2024) FLAN v2 (Apache 2.0) (Longpre et al., 2023) No Robots (CC-BY-NC-4.0) (Rajani et al. 2023) OpenAssistant Guanaco (Apache 2.0) (Kopf et al., 2024) Tulu 3 Persona MATH (ODC-BY-1.0) Tulu 3 Persona GSM (ODC-BY-1.0) Tulu 3 Persona Python… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/tulu-3-unfiltered.text1M<n<10M2 likes170 downloads2y agoHugging Face17QuixiAI /wizard_vicuna_70k_unfilteredThis dataset is the wizard_vicuna dataset junelee/wizard_vicuna_70k, removing conversations with alignment. 34598 conversations remain. inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered All credit to anon8231489123 I basically took his scripts and applied them to this new dataset. text10K<n<100K178 likes160 downloads3y agoHugging Face18HiTZ /Magpie-Llama-3-70B-Instruct-UnfilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase. The filtered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Filtered System prompts used General <|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n Code <|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI assistant designed to provide helpful, step-by-step guidance on coding problems. The user will ask you a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered.tabular1M<n<10M0 likes155 downloads2y agoHugging Face19KOREAson /YiSang-STEM_Code-UnfilteredYiSang-STEM_Code-Unfiltered is a collection of 1.1M long-cot reasoning traces generated via Qwen3-32B.It consists of 128,524 unique Korean prompts related to STEM or Coding topics collected from the web. This is not from the dataset used to train our KOREAson-0831 series. It's a bigger and unfiltered version, and might be used in our future iterations. Citation @article{son2025pushing, title={Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought}… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-STEM_Code-Unfiltered.text1M<n<10M10 likes147 downloads1y agoHugging Face20lsmpp /dongchedi_unfilteredimage1K<n<10K0 likes136 downloads11mo agoHugging Face21marketeam /marketing_user_prompts_unfilteredtext100K<n<1M3 likes134 downloads1y agoHugging Face22ewof /sharegpt-instruct-unfiltered-dedupedThis dataset is the ShareGPT unfiltered dataset anon8231489123/ShareGPT_Vicuna_unfiltered, removing instances of blatant alignment and removes duplicates. 33714 instructions remain. clean.py was first ran on hakurei/open-instruct-v1/subsets/sharegpt_data.json and then dedupe.py was ran on it. inspired by https://huggingface.co/datasets/ehartford/WizardLM_alpaca_evol_instruct_70k_unfiltered All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py, I then took this… See the full description on the dataset page: https://huggingface.co/datasets/ewof/sharegpt-instruct-unfiltered-deduped.text10K<n<100K7 likes107 downloads3y agoHugging Face23bigcode /self-oss-instruct-sc2-responses-unfilteredtext100K<n<1M2 likes93 downloads2y agoHugging Face24timchen0618 /browsecomp-plus-scout-runs-test300-qwen-sft-gpt-scout-unfiltered-v1textn<1K0 likes88 downloads4mo agoHugging Face25DialogueCharacter /chinese_belle_unfiltered Dataset Card for "chinese_belle_unfiltered" More Information needed text1M<n<10M1 likes87 downloads3y agoHugging Face26empero-ai /tasklist-grok4-multilingual-50000x-unfiltered TaskGen Dataset Generated with taskgen by empero-org Run Parameters Parameter Value Model grok-4-1-fast-non-reasoning Temperature 0.75 Total Tasks 50000 Concurrency 8 workers API Base https://api.x.ai/v1 Generated 2026-04-07 09:04:57 Budget Cap $15.0000 Multilingual Yes (en, de, fr, es, nl, zh, ar, ru) Language Distribution Language Code Tasks Arabic ar 6111 Chinese zh 6058 German de 6057 Spanish es 6020… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok4-multilingual-50000x-unfiltered.tabular10K<n<100K2 likes84 downloads6mo agoHugging Face27empero-ai /tasklist-grok-multilingual-100000x-unfiltered TaskGen Dataset Generated with taskgen by empero-org Run Parameters Parameter Value Model grok-4-1-fast-reasoning Temperature 0.9 Total Tasks 83052 Concurrency 30 workers API Base https://api.x.ai/v1 Generated 2026-04-07 14:31:14 Budget Cap $15.0000 Multilingual Yes (en, de, fr, es, nl, zh, ar, ru) Language Distribution Language Code Tasks Arabic ar 10446 German de 10397 Dutch nl 10353 Spanish es 10345… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok-multilingual-100000x-unfiltered.tabular100K<n<1M2 likes83 downloads6mo agoHugging Face28surrey-nlp /PLOD-unfilteredThis is the dataset repository for PLOD Dataset accepted to be published at LREC 2022. The dataset can help build sequence labelling models for the task Abbreviation Detection.texttoken-classification100K<n<1M1 likes81 downloads4y agoHugging Face29mychen76 /ShareGPT_V3_unfiltered_cleaned_small_9k Dataset Card for "ShareGPT_V3_unfiltered_cleaned_small_9k" More Information needed text1K<n<10K0 likes80 downloads3y agoHugging Face30Eurolingua /DCLM-200-100k-unfilteredtext10M<n<100M1 likes80 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.