datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us_election_2024_telegram_distilled
A billion Telegram messages about the 2024 US presidential election
This is a dataset of Telegram messages collected during the 2024 US presidential election. For more details, see https://dl.acm.org/doi/10.1145/3701716.3715297.
~1.03B messages, ~43K chats, ~0.8TB (distilled).
~350M English messages have toxicity- and hate-related scores from the Perspective API. For more details, see https://support.perspectiveapi.com/s/about-the-api-attributes-and-languages?language=en_US.
~350M… See the full description on the dataset page: https://huggingface.co/datasets/leonardoblas/us_election_2024_telegram_distilled.Distillation_RAGDeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1llm-classification-distilled-v2-sharded
LLM Classification Distilled v2 Sharded
Overview
This repository stores shard CSV files produced by the teacher-judge distillation pipeline.
How to Use
Run the distillation notebook once per shard:
NUM_SHARDS = 4
SHARD_INDEX = 0 .. 3
After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos.
Final Repositories
Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.distillation-llm-rawQA-deepseek-r1-distill-llama-70b
DeepSeek-R1-LLama-70B Q&A Dataset
This repository contains a curated set of 484 questions and answers generated by the DeepSeek-R1-LLama-70B model. The main goal is to evaluate the quality, coherence, and factual correctness of the model’s responses under various scenarios. Before getting excited about it, let's be realistic—large language models can produce both impressive and abysmal results. This dataset is meant to help you figure out which side of that spectrum… See the full description on the dataset page: https://huggingface.co/datasets/MedSalim/QA-deepseek-r1-distill-llama-70b.russian_llm_response_chatgpt_distill
LLM Usage in Russian (Distilled Dataset)
Dataset Summary
LLM Usage RU Dataset is a synthetic dataset of 50,000 Russian-language human–LLM interaction logs. Each sample includes a user query, the LLM's response, timestamp, user feedback, and session metadata. The dataset was generated to explore how large language models perform in Russian — a language that tends to receive less training coverage than English.
The queries and responses were distilled from GPT-4-turbo… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/russian_llm_response_chatgpt_distill.DeepSeek-R1-DistilldistillationDescription: Snapshot measurements on 27 variables from a distillation column; measured over 2.5 years.
Data source: From an industrial source; variable names have been coded. e.g. Temp1 is a temperature, but we cannot disclose where it is measured on the column.
Temperatures are in Fahrenheit
Pressures are measured in bars
FlowC1 in units of MSCFD
FlowC3 and FlowC4 are in units of MBPD
Temp11 = Temp3 - Temp9 = the temperature increase of the stream leaving the column and returning back, after… See the full description on the dataset page: https://huggingface.co/datasets/talaviyabhavik/distillation.DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-mistralru_virtual_assistant_chatgpt_distill
📊 Virtual Assistant Queries Dataset (Russian, Synthetic, 100K)
Описание
Этот датасет содержит 100,000 синтетически сгенерированных пользовательских запросов к виртуальному ассистенту на русском языке. Он предназначен для задач анализа пользовательского опыта, обработки естественного языка и предсказательного моделирования.
Каждая запись представляет собой реалистичный запрос пользователя, категорию запроса, устройство, с которого он был сделан, и оценку качества… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/ru_virtual_assistant_chatgpt_distill.Grok-Code-Fast-1-Distillation-Done-By-GPT5.4
GPT 5.4 Code Distillation
This dataset contains 500 randomly sampled prompt, reasoning, and output triples derived from the source dataset TeichAI/grok-code-fast-1-1000x.
Columns
Prompt: the user message extracted from the source conversation.
Reasoning: the assistant's <think> content when present.
Output: the assistant response after the <think> block.
Source Attribution
Prompt source and original conversation data come from TeichAI/grok-code-fast-1-1000x.… See the full description on the dataset page: https://huggingface.co/datasets/SLoonker/Grok-Code-Fast-1-Distillation-Done-By-GPT5.4.distill-sub-medmcqa-13Bllm-classification-distilled-v1-safe-filtered
LLM Classification Distilled v1 (Safe Filtered)
Overview
Safer filtered distilled dataset. Uses stricter agreement and consistency conditions for higher precision.
Files
train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_safe_filtered.csv
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form.
The dataset repository is: tussiiiii/llm-classification-distilled-v1-safe-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-safe-filtered.llm-classification-distilled-v2-teacher-hard
LLM Classification Distilled v2 (Teacher Hard)
Overview
Teacher-hard distilled subset merged from sharded v2 distillation outputs.
Files
train.csv: merged dataset file
Shard Source
Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded
Number of shards: 4
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form: A, B, C where C means tie.
This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-teacher-hard.brew-genpt-qwen32b-distillllm-classification-distilled-v1-filtered
LLM Classification Distilled v1 (Filtered)
Overview
Filtered distilled training dataset. Built from rows where target agrees with gold and basic quality conditions are met.
Files
train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_filtered.csv
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form.
The dataset repository is: tussiiiii/llm-classification-distilled-v1-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-filtered.llm-classification-distilled-v2-filtered
LLM Classification Distilled v2 (Filtered)
Overview
Filtered distilled training dataset merged from sharded v2 distillation outputs.
Files
train.csv: merged dataset file
Shard Source
Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded
Number of shards: 4
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form: A, B, C where C means tie.
This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-filtered.Ethics_Distilledllm-classification-distilled-v2
LLM Classification Distilled v2 (Full)
Overview
Full distilled training dataset merged from sharded v2 distillation outputs.
Files
train.csv: merged dataset file
Shard Source
Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded
Number of shards: 4
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form: A, B, C where C means tie.
This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2.Distill-MCQ-Genllm-classification-distilled-v1-teacher-hard
LLM Classification Distilled v1 (Teacher Hard)
Overview
Hard subset where the teacher strongly preferred a label but disagreed with gold. Useful for analysis or hard-example mixing.
Files
train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_teacher_hard.csv
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form.
The dataset repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-teacher-hard.llm-classification-distilled-v2-safe-filtered
LLM Classification Distilled v2 (Safe Filtered)
Overview
Safe filtered distilled training dataset merged from sharded v2 distillation outputs.
Files
train.csv: merged dataset file
Shard Source
Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded
Number of shards: 4
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form: A, B, C where C means tie.
This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-safe-filtered.distill_r1_qwen1p5b_math7500_soln_32k_tokensdistill_r1_qwen_2.5_1.5b_32k_soln_gpt_4o_verify_remove_thinkdistill_qwen_7b_math_train_question_solutiondistill_qwen_7b_math_train_question_solution_gpt_4o_verifydistill_r1_qwen_math_1.5b_128_solns_math_train_with_correctness_gpt_annotation
