CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes983 downloads27d agoHugging Face02MasahiroKaneko /eagle Eagle 🦅: Ethical Dataset Given from Real Interactions Introduction This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs). If you use the Eagle dataset in your research, please cite the following: @inproceedings{Eagle:arxiv:2024, title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.tabulartext-generation100K<n<1M4 likes390 downloads3y agoHugging Face03thepowerfuldeez /massive-yt-edu-queue Massive YouTube Educational Video Queue Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours. Description This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.tabularautomatic-speech-recognition1M<n<10M1 likes146 downloads7mo agoHugging Face04masharma /convolearn ConvoLearn A dataset of tutor-student conversations demonstrating dialogic (knowledge-building) pedagogies. What's in here 2,134 dialogues between teachers and a simulated 7th-grade student discussing middle school Earth Science. Each conversation demonstrates one of six knowledge-building dimensions: cognitive engagement, formative assessment, accountability, cultural responsiveness, metacognition, or power dynamics. The teachers were real educators (323 credentialed… See the full description on the dataset page: https://huggingface.co/datasets/masharma/convolearn.tabulartext-generation1K<n<10K2 likes112 downloads6mo agoHugging Face05mashu-data /reddit-comments-sample Reddit Comment Trees Sample — Initial snapshot Initial sample: 43,913 posts and 59,874 comments across three communities. This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive. Overview Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction. The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.tabulartext-generation100K<n<1M0 likes68 downloads13d agoHugging Face06thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes58 downloads4mo agoHugging Face07masoudirani777 /PerShiaArep Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/masoudirani777/PerShiaArep.tabulartext-generation10M<n<100M0 likes54 downloads1mo agoHugging Face08inigomartinez /MASEU MASEU Multilingual Dataset Description MASEU is a dataset specifically constructed to enable reliable and linguistically faithful evaluation of mathematical reasoning in Basque, a low-resource language. It is based on a manually curated subset of the mawps-asdiv-a_svamp corpus, which merges three well-established benchmarks in the domain of Math Word Problems (MWPs): MAWPS, ASDiv-A, and SVAMP. These datasets were selected for their diversity in reasoning types… See the full description on the dataset page: https://huggingface.co/datasets/inigomartinez/MASEU.tabulartext-generation1K<n<10K0 likes52 downloads4mo agoHugging Face09AgentsSci /IEEE2026_BigData_MAS-4-Science-Matching SciAgentTrace An execution-layer trace resource for scientific-agent workload characterization. A protocol fixes who reasons, what each role can see, when feedback returns, and when a workflow stops. Those choices determine the sequence of model requests that produces an answer, so protocol design is also workload design. Two workflows that consume similar token totals can issue very different request sequences. SciAgentTrace records that difference. The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.tabulartext-generation1M<n<10M0 likes40 downloads2d agoHugging Face10h-alice /cooking-master-boy-subtitle Cooking Master Boy Chat Records Chinese (trditional) subtitle of anime "Cooking Master Boy" (中華一番). Introduction This is a collection of subtitles from anime "Cooking Master Boy" (中華一番). Dataset Description The dataset is in CSV format, with the following columns: episode: The episode index of subtitle belogs to. caption_index: The autoincrement ID of subtitles. time_start: The starting timecode, which subtitle supposed to appear. time_end: The ending… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/cooking-master-boy-subtitle.tabulartext-classification10K<n<100K3 likes37 downloads2y agoHugging Face11h-alice /chat-cooking-master-boy-100k Cooking Master Boy Chat Records Chat record dataset from Twitch channel "muse_tw" during the "Cooking Master Boy" (中華一番) marathon event. Introduction This is a chat dataset collected from Twitch channel "muse_tw", while the channel is hosting a marathon anime event featuring "Cooking Master Boy" (中華一番). The featured anime "Cooking Master Boy" is a Japanese manga series written and illustrated by Etsushi Ogawa. And has a big impact on meme culture, and has a cult following… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/chat-cooking-master-boy-100k.tabulartext-classification10K<n<100K1 likes25 downloads2y agoHugging Face12ShahzebKhoso /local-code-master_telemetry_arena Local Code Arena: Comprehensive Telemetry Matrix Dataset 🏆 An Empirical Dataset tracking Local Generation Throughput (TPS), Real-Time Latency, Syntactic CodeBLEU Alignments, and Functional Pass Rates across 22 Edge Architectures. 📊 Dataset Blueprint This dataset contains a consolidated, high-fidelity matrix of 11,000 unique token-generation execution loops across 22 state-of-the-art open-weights language models (ranging from 500M to 15.5B parameters). Every… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-master_telemetry_arena.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face13lamm-mit /silkome-masp Silkome MaSp lamm-mit/silkome-masp is the major ampullate spidroin (MaSp) sequence-property subset used for the SilkomeGPT study: Wei Lu, David L. Kaplan, and Markus J. Buehler, "Generative Modeling, Design, and Analysis of Spider Silk Protein Sequences for Enhanced Mechanical Properties", Advanced Functional Materials 34, 2311324 (2024). The dataset is curated from lamm-mit/silkome-full by selecting rows whose category1 is one of: MaSp, MaSp1, MaSp2, MaSp2B, MaSp3, MaSp3B… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/silkome-masp.tabulartext-generation1K<n<10K0 likes17 downloads4mo agoHugging Face14MASSJ77 /open-dialogue-dataset-5k-sampleOpen Dialogue Dataset – 5k Cleaned Instruction-Response Conversations (Free Sample) 1️.Description This dataset contains 5,000 instruction-response dialogues sampled from the full 189k cleaned dialogues dataset. Each conversation has 2 turns: User: instruction or question Assistant: response or explanation It is fully cleaned, deduplicated, and structured in JSONL and CSV formats, ready for: Fine-tuning large language models (LLMs) Instruction-tuned chatbot training NLP research and analysis… See the full description on the dataset page: https://huggingface.co/datasets/MASSJ77/open-dialogue-dataset-5k-sample.tabulartext-generation1K<n<10K1 likes12 downloads9mo agoHugging Face15robworks-software /ccisd-unified-master-2024 CCISD Unified School Master (2024) School-level records for Clear Creek Independent School District (Texas), compiled from the district's public school pages and Texas Education Agency accountability reports. Covers 39 schools with principal names, contact details, enrollment, and accountability ratings. Loading from datasets import load_dataset ds = load_dataset("robworks-software/ccisd-unified-master-2024") all_schools = ds["full"] # all 39 schools… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/ccisd-unified-master-2024.tabulartext-generationn<1K0 likes10 downloads2mo agoHugging Face16weblab-GENIAC /OpenBookQA-Japanese-maskedgated OpenBookQA-Japanese-masked 与えられた問題に対して4つの選択肢から答えを選択するデータセット allenai/openbookqaをcyberagent/calm3-22b-chatで翻訳 5,957件 train split: 4,956件(4,957件の内1件削除) validation split: 500件 test split: 499件(500件の内1件削除) ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 Format データセットの構成は以下 { "idx": ID, "id": 元ID, "question_stem_en": 英語の質問文, "choices_en": { "text": 選択肢の文章, "label": 選択肢の記号, }… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/OpenBookQA-Japanese-masked.tabulartext-generation1K<n<10K0 likes6 downloads2y agoHugging Face17h-alice /chat-cooking-master-boy-XLgated Cooking Master Boy Chat Records Chat record dataset from Twitch channel "muse_tw" during the "Cooking Master Boy" (中華一番) marathon event. Introduction This is a chat dataset collected from Twitch channel "muse_tw", while the channel is hosting a marathon anime event featuring "Cooking Master Boy" (中華一番). The featured anime "Cooking Master Boy" is a Japanese manga series written and illustrated by Etsushi Ogawa. And has a big impact on meme culture, and has a cult following… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/chat-cooking-master-boy-XL.tabulartext-classification1M<n<10M1 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.