CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SALT-NLP /SWE-chatgated SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code. Dataset Size… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.tabulartext-generation1M<n<10M113 likes5.6k downloads5mo agoHugging Face02coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.4k downloads8mo agoHugging Face03eth-nlped /mathdial Mathdial dataset https://arxiv.org/abs/2305.14536 MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching. Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.tabulartext-generation1K<n<10K18 likes431 downloads2y agoHugging Face04McGill-NLP /TopiOCQATopiOCQA is an information-seeking conversational dataset with challenging topic switching phenomena.tabulartext-retrieval10K<n<100K10 likes300 downloads3y agoHugging Face05Finnish-NLP /Reddit_fi_2006_2022 Dataset Card for "Reddit_fi_2006_2022" Dataset Summary Reddit_fi is a filtered and post-processed corpus consisting of comments from Reddit. Some words of caution at this stage however. Subreddits were not filtered as in ScandiReddit to filter out any specific subreddits that could have hate speech, toxicity, biased. Be careful when training language models with this data and curate you dataset properly. All Reddit comments from January 2006 up until December 2022 were… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/Reddit_fi_2006_2022.tabulartext-generation1M<n<10M2 likes262 downloads3y agoHugging Face06Finnish-NLP /HPLT_Finnish_fineweb_edu_predictedtabulartext-generation1M<n<10M0 likes229 downloads2y agoHugging Face07surrey-nlp /dialect-preferences DiaLLM — Pooled Preference Dataset (Implicit Thread) Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 45,690 preference pairs, pooling all three variety-specific sets (Australian, Northern British, Indian) without variety targeting. Used for implicit-thread DPO training, where the three varieties are pooled rather than targeted individually, preserving the variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.tabulartext-generation10K<n<100K0 likes208 downloads1mo agoHugging Face08AUEB-NLP /greek-bar-bench Dataset Card for GreekBarBench 🇬🇷🏛️⚖️ GreekBarBench is a benchmark designed to evaluate LLMs on challenging legal reasoning questions across five different legal areas from the Greek Bar exams, requiring citations to statutory articles and case facts. This repository hosts two related benchmarks: Benchmark Subsets Task GreekBarBench (GBB) greekbarbench, gbb-jme Free-text legal reasoning with citations, and LLM-judge meta-evaluation GreekBarRetrieval (GBR)… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/greek-bar-bench.tabularquestion-answering1K<n<10K5 likes198 downloads2d agoHugging Face09recogna-nlp /EduBench EduBench 📚 EduBench é um benchmark em português brasileiro para avaliação de Large Language Models (LLMs) em tarefas educacionais, composto por 3,149 questões discursivas extraídas de vestibulares de alta competitividade. GitHub Paper Dataset Description Fontes USP: Universidade de São Paulo UNICAMP: Universidade Estadual de Campinas UNESP: Universidade Estadual Paulista Período 2015-2025 (11 anos de provas) Áreas do… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/EduBench.tabularquestion-answering1K<n<10K0 likes166 downloads3mo agoHugging Face10naist-nlp /XQ-MEval Dataset Card for XQ-MEval XQ-MEval is a quality-parallel benchmark dataset for automatic evaluation metrics on cross-lingual scoring bias. Dataset Details Dataset Description XQ-MEval is a benchmark released under CC BY-S 4.0 for evaluating automatic metrics with respect to cross-lingual scoring bias. This dataset is constructed by injecting varying numbers of Multidimensional Quality Metric (MQM)-defined errors into high-quality translations… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/XQ-MEval.tabulartext-generation10K<n<100K2 likes163 downloads3mo agoHugging Face11algerian-nlp /TinyStories-Algerian-Darija TinyStories Algerian Darija Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous). The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.tabulartext-generation10K<n<100K0 likes156 downloads6d agoHugging Face12surrey-nlp /alignment-british-final DiaLLM — Northern British English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 15,449 preference pairs for Northern British English (en-UK), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-british-final.tabulartext-generation10K<n<100K0 likes138 downloads1mo agoHugging Face13surrey-nlp /alignment-indian-final DiaLLM — Indian English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 18,402 preference pairs for Indian English (en-IN), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using Multi-VALUE (Ziems… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-indian-final.tabulartext-generation10K<n<100K0 likes137 downloads1mo agoHugging Face14surrey-nlp /alignment-australian-final DiaLLM — Australian English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 11,839 preference pairs for Australian English (en-AU), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using Multi-VALUE… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-australian-final.tabulartext-generation10K<n<100K0 likes130 downloads1mo agoHugging Face15nlphuji /DOVE_Litegated 🕊️ DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation 🌐 Project Website | 📄 Read our paper Updates 📅 2025-08-13: Expansion beyond multiple-choice task: Added comprehensive PromptSuite benchmark evaluations with ~37,000 LLM outputs across 9 diverse tasks including open-ended generation, mathematical reasoning, sentiment analysis, translation, summarization, and code generation (MMLU, GSM8K, SST, WMT14, CNN/DailyMail, MuSiQue… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/DOVE_Lite.tabularmultiple-choice100M<n<1B3 likes105 downloads1y agoHugging Face16hkust-nlp /drkernel-coldstart-8k DR.Kernel Cold-Start Dataset Paper | Code This directory documents the format of hkust-nlp/drkernel-coldstart-8k. The cold-start set is used for supervised fine-tuning (SFT) before RL in DR.Kernel. As described in the paper, it is built from 5-turn multi-turn trajectories collected with KernelGYM feedback. Overview Purpose: initialize kernel-generation ability (Triton coding + iterative optimization) before TRLOO/MRS/PR/PRS RL.Data form: one row per full multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/drkernel-coldstart-8k.tabulartext-generation1K<n<10K2 likes105 downloads8mo agoHugging Face17algerian-nlp /algerian-darja-corpus Algerian Darja Corpus 11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.tabulartext-generation10K<n<100K0 likes99 downloads6d agoHugging Face18upb-nlp /RoJBMO RoJBMO: Junior Balkan Mathematical Olympiad Benchmark RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination. Sources Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.tabulartext-generationn<1K1 likes87 downloads19d agoHugging Face19nlp-waseda /VisRecall VisRecall This repository contains the VisRecall benchmark, introduced in Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs. Dataset Description Imagine a tourist finished their journey in Japan and came back to France, eager to share the places they visited with their friends. When portraying these experiences, the visual information they convey is inherently independent of language, meaning that descriptions created in… See the full description on the dataset page: https://huggingface.co/datasets/nlp-waseda/VisRecall.tabulartext-generation1K<n<10K0 likes64 downloads1y agoHugging Face20hkust-nlp /dart-math-pool-gsm8k-query-info [!NOTE] This dataset is the synthesis information of queries from the GSM8K training set, such as the numbers of raw/correct samples of each synthesis job. Usually used with dart-math-pool-gsm8k. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-gsm8k-query-info.tabulartext-generation1K<n<10K2 likes63 downloads2y agoHugging Face21hkust-nlp /dart-math-pool-math-query-info [!NOTE] This dataset is the synthesis information of queries from the MATH training set, such as the numbers of raw/correct samples of each synthesis job. Usually used with dart-math-pool-math. 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the state-of-the-art… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-pool-math-query-info.tabulartext-generation1K<n<10K0 likes58 downloads2y agoHugging Face22bavarian-nlp /bavarian-wikimedia 🥨 Bavarian Wikimedia This repo hosts Bavarian Wikipedia dumps retrieved from the Wikimedia Enterprise API. Subsets Different Wikipedia dumps can be accessed as dataset subsets - based on their snapshop creation time. Currently, the following dataset subsets exists: Dataset Subset Creation Timestamp Version Size 2026-07-01 2026-07-01T02:08:28.508597145Z 1e39fd12592e52e4f41fd477f5d3e5dc 153 MB Usage from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/bavarian-nlp/bavarian-wikimedia.tabulartext-generation10K<n<100K0 likes58 downloads21d agoHugging Face23Nischithhh /NLP r/IPMATtards Reddit Dataset Dataset Description This dataset contains scraped posts and comments from the r/IPMATtards subreddit, a community dedicated to aspirants of the Integrated Programme in Management Aptitude Test (IPMAT) in India. The data is structured into two main components: Posts: Top-level submissions including titles, body text, scores, and metadata. Comments: Threaded replies associated with the posts, including recursion depth and parent-child… See the full description on the dataset page: https://huggingface.co/datasets/Nischithhh/NLP.tabulartext-generation1K<n<10K0 likes56 downloads14d agoHugging Face24guanvireak /khmer-nlp-technical-corpus khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus Dataset Summary This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers. Dataset Statistics Total Documents: 3 Train Documents: 3 Total Words: 8,002 Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.tabulartext-generationn<1K0 likes40 downloads9d agoHugging Face25nlp-with-deeplearning /ko.SHP 🚢 Korean Stanford Human Preferences Dataset (Ko.SHP) 이 데이터셋은 자체 구축한 번역기를 활용하여 stanfordnlp/SHP 데이터셋을 번역한 것입니다. 아래의 내용은 해당 번역기로 README 파일을 번역한 것입니다. 참고 부탁드립니다. If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022). Summary SHP는 요리에서 법률 조언에 이르기까지 18가지 다른 주제 영역의 질문/지침에 대한 응답에 대한 385K 집단 인간 선호도 데이터 세트이다. 기본 설정은 다른 응답에 대 한 한 응답의 유용성을 반영 하기 위한 것이며 RLHF 보상 모델 및 NLG 평가 모델 (예: SteamSHP)을 훈련 하는 데… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.tabulartext-generation100K<n<1M1 likes34 downloads3y agoHugging Face26SUSTech-NLP /MixReward MixReward: A Large-Scale Multilingual Preference Dataset Overview MixReward is a large-scale, high-quality multilingual preference dataset comprising 64,528 examples across 6 domains and 103 languages. It is designed to train unified reasoning reward models that support multiple evaluation paradigms (pairwise, listwise, and pointwise). This dataset is introduced in the following paper, accepted at ICML 2026 (the 43rd International Conference on Machine Learning):… See the full description on the dataset page: https://huggingface.co/datasets/SUSTech-NLP/MixReward.tabulartext-generation10K<n<100K0 likes34 downloads3mo agoHugging Face27disi-unibo-nlp /PRISM PRISM: Impact of Decoding Strategies for Abstractive Document Summarization at Test Time Dataset Description PRISM is a comprehensive evaluation dataset for studying the impact of different decoding strategies on abstractive document summarization performance. The dataset contains results from 9 decoding strategies applied to 8 models across 6 datasets, providing a systematic comparison of generation approaches. Dataset Summary This dataset contains evaluation… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/PRISM.tabularsummarization1K<n<10K0 likes32 downloads1y agoHugging Face28NLPForUA /dumy-zno-ukrainian-math-history-geo-r1-o1 DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers) DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks. The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian: Думи мої, думи мої, Лихо мені з вами! Нащо стали на папері Сумними рядами?.. Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.tabulartext-generation1K<n<10K2 likes32 downloads1y agoHugging Face29Finnish-NLP /belebele-fi-filtered-sft Dataset Card for Finnish-NLP/benebele Creation process Finnish subset loaded from facebook/belebele tabulartext-generationn<1K0 likes29 downloads3y agoHugging Face30orai-nlp /ultrafeedback_eu Ultrafeedback Binarized machine translated preference dataset for Basque Dataset Creation Source Data Machine translated to Basque from the Ultrafeedback Binarized dataset. Annotations Annotation process Machine translated to Basque from the Ultrafeedback Binarized dataset. Citation [optional] If you use this dataset please cite the following reference: @misc{Llama-eus, title = {Llama-eus-8B, a foundational sub-10 billion… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/ultrafeedback_eu.tabulartext-generation10K<n<100K0 likes22 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.