CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B99 likes727k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2023-23 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes20k downloads2y agoHugging Face03minpeter /xlam-function-calling-60k-parsed [PARSED] APIGen Function-Calling Datasets (xLAM) This dataset contains the full data from the original Salesforce/xlam-function-calling-60k Subset name multi-turn parallel multiple definition Last turn type number of dataset xlam-function-calling-60k no yes yes tool_calls 60000 This is a re-parsing formatting dataset for the xLAM official dataset. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/xlam-function-calling-60k-parsed.texttext-generation10K<n<100K3 likes20k downloads1y agoHugging Face04mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face05mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes15k downloads2y agoHugging Face06mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes8.3k downloads2y agoHugging Face07mlfoundations /MINT-1T-PDF-CC-2023-50 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes6k downloads2y agoHugging Face08Smith42 /minty-astro-ph MINT-1T ArXiv Astro-ph An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers). Overview Papers ~845k Total size ~804 GB Format WebDataset tar shards Shards 287 (astro-ph-00000.tar to astro-ph-00286.tar) Shard size ~3 GB each Source MINT-1T (Awadalla et al., 2024) Data Format Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.imagetext-generation100K<n<1M1 likes4.6k downloads5mo agoHugging Face09MindlessForMinerva /fineweb-nopotter 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We then… See the full description on the dataset page: https://huggingface.co/datasets/MindlessForMinerva/fineweb-nopotter.tabulartext-generation1B<n<10B0 likes3.6k downloads7mo agoHugging Face10JeanKaddour /minipile Dataset Card for MiniPile Dataset Description The MiniPile Challenge for Data-Efficient Language Models Dataset Summary MiniPile is a 6GB subset of the deduplicated The Pile corpus. To curate MiniPile, we perform a simple, three-step data filtering process: we (1) infer embeddings for all documents of the Pile, (2) cluster the embedding space using k-means, and (3) filter out low-quality clusters. The primary motivation for curating MiniPile is that (i) diverse… See the full description on the dataset page: https://huggingface.co/datasets/JeanKaddour/minipile.texttext-generation1M<n<10M150 likes2.6k downloads3y agoHugging Face11SkillFi /deepseek-v2-codder-minecraft-apitexttext-generationn<1K0 likes2.1k downloads1y agoHugging Face12yulan-team /YuLan-Mini-Text-Datasets News [2025.04.11] Add dataset mixture: link. [2025.03.30] Text datasets upload finished. This is text dataset. 这是文本格式的数据集。 Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here. 由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。 For more information, please refer to our datasets details and preprocess details. Contributing We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.tabulartext-generation100M<n<1B12 likes2k downloads1y agoHugging Face13MiniMaxAI /SynLogic SynLogic Dataset SynLogic is a comprehensive synthetic logical reasoning dataset designed to enhance logical reasoning capabilities in Large Language Models (LLMs) through reinforcement learning with verifiable rewards. 🐙 GitHub Repo: https://github.com/MiniMax-AI/SynLogic 📜 Paper (arXiv): https://arxiv.org/abs/2505.19641 Dataset Description SynLogic contains 35 diverse logical reasoning tasks with automatic verification capabilities, making it ideal for… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/SynLogic.texttext-generation10K<n<100K106 likes1.8k downloads1y agoHugging Face14KevinNotSmile /nuscenes-qa-mini NuScenes-QA-mini Dataset TL;DR: This dataset is used for multimodal question-answering tasks in autonomous driving scenarios. We created this dataset based on nuScenes-QA dataset for evaluation in our paper Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI. The samples are divided into day and night scenes. scene # train samples # validation samples day 2,229 2,229 night 659 659 Each sample contains… See the full description on the dataset page: https://huggingface.co/datasets/KevinNotSmile/nuscenes-qa-mini.textvisual-question-answering1K<n<10K4 likes1.3k downloads3y agoHugging Face15PursuitOfDataScience /MiniMax-M2.1-Mixture-of-Thoughts MiniMax-M2.1 Mixture of Thoughts This dataset contains responses generated by MiniMax-M2.1 for user questions from the open-r1/Mixture-of-Thoughts dataset. Dataset Description The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation. Metric Value Examples 349,317 Total Tokens 4,052,592,552 Avg Tokens/Example 11,601 Source Dataset Name:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts.tabulartext-generation100K<n<1M2 likes1.2k downloads9mo agoHugging Face16minpeter /bfcl-v1-non-live-ast-parsed [PARSED] BFCL V1 AST (non-live python) The data in this dataset is a subset of the original gorilla-llm/Berkeley-Function-Calling-Leaderboard Subset name multi-turn parallel multiple definition Last turn type number of dataset simple no no no tool_calls 400 multiple no no yes tool_calls 200 parallel no yes no tool_calls 200 parallel_multiple no yes yes tool_calls 200 This is a re-parsing formatting dataset for Python AST parts from V1 of the official dataset of… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/bfcl-v1-non-live-ast-parsed.texttext-generation1K<n<10K1 likes1.1k downloads2y agoHugging Face17minpeter /hermes-function-calling-v1-jsonl Hermes Function-Calling V1 This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models. This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/hermes-function-calling-v1-jsonl.texttext-generation10K<n<100K1 likes1k downloads2y agoHugging Face18iNeil77 /pseudo-mini-pileA small, aggressively cleaned and de-duped pre-training corpus for academic settings. It aims to recreate something akin to The Pile but prioritizes quality for the constrained token budget academic researchers live with. It has seven config subsets and an eighth all subset that combines them for a total of ~91B tokens (GPT2 Tokenizer estimate). These splits are as follows: c4_realnews: The RealNews domain subset of the C4 dataset containing news articles. openwebtext: The OpenWebText dataset… See the full description on the dataset page: https://huggingface.co/datasets/iNeil77/pseudo-mini-pile.texttext-generation100M<n<1B4 likes1k downloads3y agoHugging Face19minpeter /toolace-parsed [PARSED] ToolACE The data in this dataset is a subset of the original Team-ACE/ToolACE Subset name multi-turn parallel multiple definition Last turn type number of dataset toolace yes yes yes complex 11k This is a re-parsing formatting dataset for the ToolACE official dataset. Load the dataset from datasets import load_dataset ds = load_dataset("minpeter/toolace-parsed") print(ds) # DatasetDict({ # train: Dataset({ # features:… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/toolace-parsed.texttext-generation10K<n<100K1 likes949 downloads2y agoHugging Face20minnesotanlp /Finch-Collection Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions. 👋 Welcome to Finch Collection, the dataset proposed in Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks. It is a 156K-trajectory a large-scale dataset of 156K evolutionary search trajectories collected… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection.imagetext-generation100K<n<1M8 likes729 downloads3mo agoHugging Face21ViuAI /viu-mini-raw-pretrain ViuAI/viu-mini-raw-pretrain Raw pretraining corpus for ViuMini-MoE-242M (Hinglish-first, 40:30:30 mix target). Unified schema: text, lang, source, domain, safety_tag. text: training text (plain, or <|user|> ... <|assistant|> ... templated for QA/instruct/distillation sources) lang: hinglish | hindi | english | bilingual source: origin dataset name (e.g. indiccorp_v2, fineweb-edu, bespoke-stratos-r1, smoltalk, numina_math_cot) domain: knowledge pillar (e.g. general_hindi… See the full description on the dataset page: https://huggingface.co/datasets/ViuAI/viu-mini-raw-pretrain.texttext-generation100M<n<1B0 likes602 downloads3h agoHugging Face22armand0e /minimax-m3-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Minimax M3 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by minimax/minimax-m3. JSONL files: 31 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m3-claude-code-traces.tabulartext-generationn<1K13 likes557 downloads4mo agoHugging Face23LydiaNottingham /MindTheGap Mind the Gap Dataset Description This dataset accompanies the paper "Mind the Gap: How Elicitation Protocols Shape the Stated-Revealed Preference Gap in Language Models" and extends the original AIRiskDilemmas dataset with comprehensive model evaluation results. Authors: Pranav Mahajan, Ihor Kendiukhov, Syed Hussain, Lydia Nottingham Repository: SPAR-SvR/Mind-the-Gap Original Dataset: AIRiskDilemmas Key Contribution We systematically study how elicitation… See the full description on the dataset page: https://huggingface.co/datasets/LydiaNottingham/MindTheGap.texttext-classification1K<n<10K0 likes527 downloads8mo agoHugging Face24devngho /the-stack-mini-nonshuffledThis repo contains (up to) 30k samples of 21 languages (top 20 languages by StackOverflow survey, html/css was splited). 'javascript', 'html', 'css', 'python', 'sql', 'typescript', 'shell', 'java', 'c-sharp', 'cpp', 'c', 'php', 'powershell', 'go', 'rust', 'kotlin', 'lua', 'dart', 'assembly', 'ruby', 'swift' tabulartext-generation1M<n<10M2 likes519 downloads2y agoHugging Face25devngho /culturax-mini-nonshuffledThis repo contains 1% of each language of uonlp/CulturaX. load_dataset('devngho/culturax-mini-nonshuffled', '[lang]', split='train') # read specified language load_dataset('devngho/culturax-mini-nonshuffled', data_files="*/*", split='train') # read all language texttext-generation10M<n<100M0 likes514 downloads2y agoHugging Face26EliMC /fineweb-edu-10BT-mincols fineweb-edu: 10BT sample This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons. token counts GPT-4 tiktoken token count: token_count count 9.672101e+06 mean 1.001188e+03 std 1.834986e+03 min 3.800000e+01 25% 3.380000e+02 50% 6.090000e+02 75% 1.054000e+03 max 1.649670e+05 Total count: 9683.59 M tokens texttext-generation1M<n<10M0 likes495 downloads10mo agoHugging Face27BEE-spoke-data /fineweb-edu-10BT-mincols fineweb-edu: 10BT sample This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons. token counts GPT-4 tiktoken token count: token_count count 9.672101e+06 mean 1.001188e+03 std 1.834986e+03 min 3.800000e+01 25% 3.380000e+02 50% 6.090000e+02 75% 1.054000e+03 max 1.649670e+05 Total count: 9683.59 M tokens texttext-generation1M<n<10M1 likes485 downloads9mo agoHugging Face28MiniMaxAI /role-play-bench Role-play Benchmark A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios. Dataset Summary Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/role-play-bench.tabulartext-generation1K<n<10K151 likes455 downloads8mo agoHugging Face29barissozudogru /swe-bench-mini SWE-bench-mini 34 self-contained bug-fix tasks in the SWE-bench format — a small repository snapshot carrying a defect, a test that fails because of it, and a gold patch that fixes it (difficulty mix: 12 easy / 19 medium / 3 hard, author estimate). Built for the swe_bench_mini agent and the make demo-swe-mini evaluator in adk-agent-playground, to demonstrate the framework's range on code-modification and to exercise the CaMeL filesystem-capability gate. A second harder config… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/swe-bench-mini.texttext-generationn<1K1 likes437 downloads2mo agoHugging Face30seoulraphaellee /korean-assembly-minutes 대한민국 국회 회의록 아카이브 국회 회의록 원문(record.assembly.go.kr) PDF 에서 본문을 뽑아 모은 것이다. 본회의와 각 위원회 회의록이 모두 들어 있다. 수록 기간: 1948~1993 회의 수: 1,951건 본문 분량: 65,306,444자 구성 연도별 JSONL(gzip) 한 덩이다. from datasets import load_dataset ds = load_dataset("seoulraphaellee/korean-assembly-minutes", split="train") ds = load_dataset("seoulraphaellee/korean-assembly-minutes", data_files="data/2026.jsonl.gz", split="train") 필드 이름 설명 meeting_key 회의 식별자 (record:<id>)… See the full description on the dataset page: https://huggingface.co/datasets/seoulraphaellee/korean-assembly-minutes.tabulartext-generation10K<n<100K0 likes427 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.