CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openai /gsm8k Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/openai/gsm8k.texttext-generation10K<n<100K1.7k likes1.2m downloads6mo agoHugging Face02google /IFEval Dataset Card for IFEval Dataset Summary This dataset contains the prompts used in the Instruction-Following Eval (IFEval) benchmark for large language models. It contains around 500 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times" which can be verified by heuristics. To load the dataset, run: from datasets import load_dataset ifeval = load_dataset("google/IFEval") Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/google/IFEval.texttext-generationn<1K167 likes363k downloads2y agoHugging Face03Idavidrein /gpqagated Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.tabularquestion-answering1K<n<10K542 likes128k downloads1d agoHugging Face04aisa-group /PostTrainBench-Trajectories PostTrainBench Agent Traces Agent traces from PostTrainBench (GitHub), a benchmark that measures CLI agents' ability to post-train base LLMs. Task Each agent is given: A pre-trained base LLM to fine-tune An evaluation script for a specific benchmark 10 hours on an NVIDIA H100 80GB GPU The agent must autonomously improve the model's performance on the target benchmark using any post-training strategy it chooses (SFT, LoRA, RLHF, prompt engineering for data… See the full description on the dataset page: https://huggingface.co/datasets/aisa-group/PostTrainBench-Trajectories.text-generationn<1K8 likes84k downloads3d agoHugging Face05glaiveai /glaive-function-calling-v2texttext-generation100K<n<1M530 likes55k downloads3y agoHugging Face06codeparrot /github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.text-generation421 likes40k downloads4y agoHugging Face07lockon /glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en. texttext-generation1K<n<10K1 likes27k downloads2y agoHugging Face08gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes23k downloads4y agoHugging Face09Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M196 likes18k downloads25d agoHugging Face10flwrlabs /alpaca-gpt4 Dataset Card for alpaca-gpt4 This dataset originates from this repository. The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts. Dataset Details Dataset Description Each sample is comprised of four columns: instruction, input, output and text. Language(s): English License: Creative Commons NonCommercial (CC BY-NC 4.0) Dataset Sources The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.texttext-generation10K<n<100K0 likes17k downloads1y agoHugging Face11ibm-granite /ChartNet ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding 🌐 Homepage | 📖 arXiv 📝 Changelog June 3, 2026 — Release of grounded_qa subset and completed reasoning subset (both subject to Notice Regarding Data Availability) May 15, 2026 — Added link to 30K real-world charts and detailed captions dataset released by our collaborators Abaka AI/2077AI. April 29, 2026 — Release of an additional 2.5 million row subset core_permissive (subject to… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/ChartNet.imageimage-to-text1M<n<10M47 likes14k downloads4mo agoHugging Face12AdhyanshVerma /open-github-major-repos🌐 AdhyanshVerma's Open GitHub Major Repos An elite, curated collection of GitHub commit metadata from the world's most influential technology companies: Microsoft, Google, Meta, and Intel. 📖 Introduction Welcome to AdhyanshVerma's Open GitHub Major Repos dataset. This dataset focuses exclusively on high-impact, industry-standard repositories maintained by the world's leading technology giants. It utilizes the Lazy Pointer Pattern: instead of bloating your storage… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/open-github-major-repos.text-generation100K<n<1M1 likes13k downloads17d agoHugging Face13sedthh /gutenberg_english Dataset Card for Project Gutenber - English Language eBooks A collection of non-english language eBooks (48284 rows, 80%+ of all english language books available on the site) from the Project Gutenberg site with metadata removed. Originally colected for https://github.com/LAION-AI/Open-Assistant (follows the OpenAssistant training format) The METADATA column contains catalogue meta information on each book as a serialized JSON: key original column language - text_id… See the full description on the dataset page: https://huggingface.co/datasets/sedthh/gutenberg_english.texttext-generation10K<n<100K39 likes11k downloads4y agoHugging Face14PranavViswanath /jlens-gp-auditbench The AuditBench J-lens corpus: 80-layer activations and gradient-pursuit readouts Everything needed to redo J-space interpretability work on the 84 AuditBench model organisms (14 hidden behaviors x 2 instillation methods x 3 adversarial-training levels) without a GPU harvest: the raw bf16 residual stream at all 80 layers for every recorded token, and a gradient-pursuit J-lens decomposition at every (position, layer) site. The organisms are Llama-3.3-70B-Instruct with an… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/jlens-gp-auditbench.tabulartext-generation10M<n<100M0 likes9.4k downloads2mo agoHugging Face15manu /project_gutenberg Dataset Card for "Project Gutenberg" Project Gutenberg is a library of over 70,000 free eBooks, hosted at https://www.gutenberg.org/. All examples correspond to a single book, and contain a header and a footer of a few lines (delimited by a *** Start of *** and *** End of *** tags). Usage from datasets import load_dataset ds = load_dataset("manu/project_gutenberg", split="fr", streaming=True) print(next(iter(ds))) License Full license is available here:… See the full description on the dataset page: https://huggingface.co/datasets/manu/project_gutenberg.texttext-generation10K<n<100K74 likes8.8k downloads3y agoHugging Face16ONE-Lab /GUI-World GUI-World: A Dataset for GUI-Orientated Multimodal Large Language Models Dataset: GUI-World Overview GUI-World introduces a comprehensive benchmark for evaluating MLLMs in dynamic and complex GUI environments. It features extensive annotations covering six GUI scenarios and eight types of GUI-oriented questions. The dataset assesses state-of-the-art ImageLLMs and VideoLLMs, highlighting their limitations in handling dynamic and multi-step tasks. It provides… See the full description on the dataset page: https://huggingface.co/datasets/ONE-Lab/GUI-World.videoquestion-answering10K<n<100K45 likes8.1k downloads1y agoHugging Face17tmquan /anle-toaan-gov-vn Vietnamese Án lệ Corpus — anle.toaan.gov.vn 🇻🇳 Tóm tắt. Bộ dữ liệu các bản án + án lệ Việt Nam thu thập từ cổng anle.toaan.gov.vn của Tòa án nhân dân tối cao. Mỗi văn bản đi kèm markdown chuẩn hoá tiếng Việt và một lớp grounding mức câu (mỗi trích dẫn mang sentence_id + char span trỏ ngược vào markdown). Bộ dữ liệu là một phần của ViLA common-corpus và ship ba cấu hình HF theo chuẩn chung: documents (bảng chính) · embeddings (vector 4096-D Nemotron-3-Embed-8B) · reduces (toạ… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/anle-toaan-gov-vn.tabulartext-classification10K<n<100K10 likes8k downloads7d agoHugging Face18glaiveai /reasoning-v1-20m We are excited to release a synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using deepseek-ai/DeepSeek-R1-Distill-Llama-70B. While there have been multiple efforts to build open reasoning datasets for math and code tasks, we noticed a lack of large datasets containing reasoning traces for diverse non code/math topics like social and natural sciences, education, creative writing and general conversations, which is why we decided to release this… See the full description on the dataset page: https://huggingface.co/datasets/glaiveai/reasoning-v1-20m.texttext-generation10M<n<100M238 likes7.9k downloads2y agoHugging Face19cs-giung /clean-gsm8k-aug Clean GSM8K-Aug Overview The Clean GSM8K-Aug family is a revised version of whynlp/gsm8k-aug and whynlp/gsm8k-aug-nl. It retains the original question, steps, and answer schema while removing or repairing examples with incomplete or inconsistent calculation traces. Four representations of the same aligned questions and answers are available: Dataset Step representation cs-giung/clean-gsm8k-aug Infix arithmetic expressions… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/clean-gsm8k-aug.texttext-generation100K<n<1M0 likes7.7k downloads2mo agoHugging Face20AtomicChat /Qwen3.8-27B-GGUF-metrics Qwen3.8-27B GGUF, everything behind the numbers This is the working record for AtomicChat/Qwen3.8-27B-GGUF. Every figure in that model card came from a file in here, including the ones about other publishers' builds. The point of publishing it is simple. A quantization comparison is only worth reading if someone else can run it, and that needs three things nobody usually ships: the exact reference the numbers were measured against, the exact text they were measured on, and the… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/Qwen3.8-27B-GGUF-metrics.text-generation4 likes7.6k downloads1mo agoHugging Face21Shofo /shofo-tiktok-general-small Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos. Size: ~50K videos (~500GB) Modality: Video + Audio + Text (transcripts, comments, captions) Source: TikTok Schema Column Type Description file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tabularvideo-classification10K<n<100K22 likes7.4k downloads7mo agoHugging Face22Gryphe /Opus-WritingPrompts Opus Writing Prompts This is a dataset containing 3008 short stories, generated by an unrestrained Claude Opus using Reddit's Writing Prompts as a source. Each sample is generally between 4000-6000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Disclaimer: This dataset is extremely varied and includes erotica. You have been warned. Three files are included: A ShareGPT dataset, ready to be used for… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/Opus-WritingPrompts.texttext-generation1K<n<10K86 likes7.3k downloads2y agoHugging Face23theelderemo /genius-lyrics-cleaned ◎ Genius Lyrics Dataset Cleaned & Deduplicated 🤗 Hugging Face 🤗 Hugging Face DOI: 10.57967/hf/7978 DOI: 10.57967/hf/7978 revision: 9742989 revision: 9742989 A heavily cleaned, English-only, genre-filtered subset of the Genius Song… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/genius-lyrics-cleaned.texttext-generation1M<n<10M19 likes7k downloads7mo agoHugging Face24Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes6.6k downloads2y agoHugging Face25thefcraft /gentoomen-lib Gentoomen Library The Gentoomen Library is an extensive archive of technology-related resources originally shared on 4chan's /g/ board. It consists of a collection of files and directories covering various topics in computer science and technology. Here is a Basic Search Engine To Search Your Pdf Book Overview Total Size: 32.8GB Format: Extracted files and directories Topics Covered: Algorithms Scripting Technology guides Computer science… See the full description on the dataset page: https://huggingface.co/datasets/thefcraft/gentoomen-lib.text-generation2 likes6.6k downloads2y agoHugging Face26Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes6.5k downloads3y agoHugging Face27tomg-group-umd /huginn-dataset The Huginn Dataset This is a record of the dataset collection used to train the huginn-0125 model. The data is provided in a semi-prepared format. We provide 4096 parquet files for train and val each which contain the exact rows used for training and validation (on the 4096 accelerators the model was trained on). Each row is 4097 tokens long, which includes formatting tokens. The tokenizer here is the same as the model, https://huggingface.co/tomg-group-umd/huginn-0125. However… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/huginn-dataset.texttext-generation100M<n<1B9 likes6.2k downloads1y agoHugging Face28nlile /24-game Math Twenty Four (24s Game) Dataset A comprehensive dataset for the classic math twenty four game (also known as the 4 numbers game / 24s game / Game of 24). This dataset of mathematical reasoning challenges was collected from 4nums.com, featuring over 1,300 unique puzzles of the Game of 24, with difficulty metrics derived from over 6.4 million human solution attempts since 2012. In each puzzle, players must use exactly four numbers and basic arithmetic operations (+, -, ×, /) to… See the full description on the dataset page: https://huggingface.co/datasets/nlile/24-game.tabularmultiple-choice1K<n<10K14 likes6.2k downloads2y agoHugging Face29silk-road /alpaca-data-gpt4-chinesetexttext-generation10K<n<100K104 likes6k downloads3y agoHugging Face30Carlosaug47 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M4 likes6k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.