CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01leonardoblas /us_election_2024_telegram_distilled A billion Telegram messages about the 2024 US presidential election This is a dataset of Telegram messages collected during the 2024 US presidential election. For more details, see https://dl.acm.org/doi/10.1145/3701716.3715297. ~1.03B messages, ~43K chats, ~0.8TB (distilled). ~350M English messages have toxicity- and hate-related scores from the Perspective API. For more details, see https://support.perspectiveapi.com/s/about-the-api-attributes-and-languages?language=en_US. ~350M… See the full description on the dataset page: https://huggingface.co/datasets/leonardoblas/us_election_2024_telegram_distilled.tabularn<1K1 likes15k downloads8mo agoHugging Face02ulysse1 /Distillation_RAGdocument1K<n<10K0 likes220 downloads4mo agoHugging Face03CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes153 downloads2y agoHugging Face04tussiiiii /llm-classification-distilled-v2-sharded LLM Classification Distilled v2 Sharded Overview This repository stores shard CSV files produced by the teacher-judge distillation pipeline. How to Use Run the distillation notebook once per shard: NUM_SHARDS = 4 SHARD_INDEX = 0 .. 3 After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos. Final Repositories Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.tabulartext-classification100K<n<1M0 likes118 downloads4mo agoHugging Face05Infomaniak-AI /distillation-llm-rawdocument1K<n<10K0 likes87 downloads1mo agoHugging Face06MedSalim /QA-deepseek-r1-distill-llama-70b DeepSeek-R1-LLama-70B Q&A Dataset This repository contains a curated set of 484 questions and answers generated by the DeepSeek-R1-LLama-70B model. The main goal is to evaluate the quality, coherence, and factual correctness of the model’s responses under various scenarios. Before getting excited about it, let's be realistic—large language models can produce both impressive and abysmal results. This dataset is meant to help you figure out which side of that spectrum… See the full description on the dataset page: https://huggingface.co/datasets/MedSalim/QA-deepseek-r1-distill-llama-70b.textn<1K1 likes46 downloads2y agoHugging Face07ZennyKenny /russian_llm_response_chatgpt_distill LLM Usage in Russian (Distilled Dataset) Dataset Summary LLM Usage RU Dataset is a synthetic dataset of 50,000 Russian-language human–LLM interaction logs. Each sample includes a user query, the LLM's response, timestamp, user feedback, and session metadata. The dataset was generated to explore how large language models perform in Russian — a language that tends to receive less training coverage than English. The queries and responses were distilled from GPT-4-turbo… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/russian_llm_response_chatgpt_distill.text10K<n<100K2 likes41 downloads1y agoHugging Face08tuanha1305 /DeepSeek-R1-Distilltabular100K<n<1M11 likes30 downloads2y agoHugging Face09talaviyabhavik /distillationDescription: Snapshot measurements on 27 variables from a distillation column; measured over 2.5 years. Data source: From an industrial source; variable names have been coded. e.g. Temp1 is a temperature, but we cannot disclose where it is measured on the column. Temperatures are in Fahrenheit Pressures are measured in bars FlowC1 in units of MSCFD FlowC3 and FlowC4 are in units of MBPD Temp11 = Temp3 - Temp9 = the temperature increase of the stream leaving the column and returning back, after… See the full description on the dataset page: https://huggingface.co/datasets/talaviyabhavik/distillation.tabularn<1K0 likes20 downloads3y agoHugging Face10CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-mistraltabular1K<n<10K0 likes20 downloads2y agoHugging Face11ZennyKenny /ru_virtual_assistant_chatgpt_distill 📊 Virtual Assistant Queries Dataset (Russian, Synthetic, 100K) Описание Этот датасет содержит 100,000 синтетически сгенерированных пользовательских запросов к виртуальному ассистенту на русском языке. Он предназначен для задач анализа пользовательского опыта, обработки естественного языка и предсказательного моделирования. Каждая запись представляет собой реалистичный запрос пользователя, категорию запроса, устройство, с которого он был сделан, и оценку качества… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/ru_virtual_assistant_chatgpt_distill.tabulartext-classification100K<n<1M3 likes19 downloads1y agoHugging Face12SLoonker /Grok-Code-Fast-1-Distillation-Done-By-GPT5.4 GPT 5.4 Code Distillation This dataset contains 500 randomly sampled prompt, reasoning, and output triples derived from the source dataset TeichAI/grok-code-fast-1-1000x. Columns Prompt: the user message extracted from the source conversation. Reasoning: the assistant's <think> content when present. Output: the assistant response after the <think> block. Source Attribution Prompt source and original conversation data come from TeichAI/grok-code-fast-1-1000x.… See the full description on the dataset page: https://huggingface.co/datasets/SLoonker/Grok-Code-Fast-1-Distillation-Done-By-GPT5.4.texttext-generationn<1K2 likes18 downloads5mo agoHugging Face13Roy-Shih /distill-sub-medmcqa-13Btabular10K<n<100K0 likes17 downloads3y agoHugging Face14tussiiiii /llm-classification-distilled-v1-safe-filtered LLM Classification Distilled v1 (Safe Filtered) Overview Safer filtered distilled dataset. Uses stricter agreement and consistency conditions for higher precision. Files train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_safe_filtered.csv Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form. The dataset repository is: tussiiiii/llm-classification-distilled-v1-safe-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-safe-filtered.tabulartext-classification1K<n<10K0 likes15 downloads5mo agoHugging Face15tussiiiii /llm-classification-distilled-v2-teacher-hard LLM Classification Distilled v2 (Teacher Hard) Overview Teacher-hard distilled subset merged from sharded v2 distillation outputs. Files train.csv: merged dataset file Shard Source Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded Number of shards: 4 Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form: A, B, C where C means tie. This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-teacher-hard.tabulartext-classification10K<n<100K0 likes14 downloads4mo agoHugging Face16marcuscedricridia /brew-genpt-qwen32b-distilltext10K<n<100K0 likes13 downloads2y agoHugging Face17tussiiiii /llm-classification-distilled-v1-filtered LLM Classification Distilled v1 (Filtered) Overview Filtered distilled training dataset. Built from rows where target agrees with gold and basic quality conditions are met. Files train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_filtered.csv Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form. The dataset repository is: tussiiiii/llm-classification-distilled-v1-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-filtered.tabulartext-classification1K<n<10K0 likes10 downloads5mo agoHugging Face18tussiiiii /llm-classification-distilled-v2-filtered LLM Classification Distilled v2 (Filtered) Overview Filtered distilled training dataset merged from sharded v2 distillation outputs. Files train.csv: merged dataset file Shard Source Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded Number of shards: 4 Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form: A, B, C where C means tie. This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-filtered.tabulartext-classification10K<n<100K0 likes9 downloads4mo agoHugging Face19sungjun83 /Ethics_Distilledtext10K<n<100K0 likes8 downloads3y agoHugging Face20tussiiiii /llm-classification-distilled-v2 LLM Classification Distilled v2 (Full) Overview Full distilled training dataset merged from sharded v2 distillation outputs. Files train.csv: merged dataset file Shard Source Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded Number of shards: 4 Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form: A, B, C where C means tie. This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2.tabulartext-classification10K<n<100K0 likes8 downloads4mo agoHugging Face21vinhnado /Distill-MCQ-Gentext10K<n<100K0 likes7 downloads1y agoHugging Face22tussiiiii /llm-classification-distilled-v1-teacher-hard LLM Classification Distilled v1 (Teacher Hard) Overview Hard subset where the teacher strongly preferred a label but disagreed with gold. Useful for analysis or hard-example mixing. Files train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_teacher_hard.csv Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form. The dataset repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-teacher-hard.tabulartext-classificationn<1K0 likes7 downloads5mo agoHugging Face23tussiiiii /llm-classification-distilled-v2-safe-filtered LLM Classification Distilled v2 (Safe Filtered) Overview Safe filtered distilled training dataset merged from sharded v2 distillation outputs. Files train.csv: merged dataset file Shard Source Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded Number of shards: 4 Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form: A, B, C where C means tie. This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-safe-filtered.tabulartext-classification10K<n<100K0 likes7 downloads4mo agoHugging Face24hbXNov /distill_r1_qwen1p5b_math7500_soln_32k_tokenstabular1K<n<10K1 likes6 downloads2y agoHugging Face25hbXNov /distill_r1_qwen_2.5_1.5b_32k_soln_gpt_4o_verify_remove_thinktabular1K<n<10K0 likes5 downloads2y agoHugging Face26hbXNov /distill_qwen_7b_math_train_question_solutiontext1K<n<10K0 likes4 downloads2y agoHugging Face27hbXNov /distill_qwen_7b_math_train_question_solution_gpt_4o_verifytext1K<n<10K0 likes3 downloads2y agoHugging Face28hbXNov /distill_r1_qwen_math_1.5b_128_solns_math_train_with_correctness_gpt_annotationtabular1K<n<10K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.