CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MU-NLPC /Calc-asdiv_a Dataset Card for Calc-asdiv_a Summary The dataset is a collection of simple math word problems focused on arithmetics. It is derived from the arithmetic subset of ASDiv (original repo). The main addition in this dataset variant is the chain column. It was created by converting the solution to a simple html-like language that can be easily parsed (e.g. by BeautifulSoup). The data contains 3 types of tags: gadget: A tag whose content is intended to be evaluated by calling… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-asdiv_a.tabular1K<n<10K1 likes16k downloads3y agoHugging Face02SALT-NLP /SWE-chatgated SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code. Dataset Size… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.tabulartext-generation1M<n<10M112 likes5.6k downloads5mo agoHugging Face03hitachi-nlp /proofwriter_processed_OWAtabular10K<n<100K2 likes4k downloads2y agoHugging Face04princeton-nlp /QuRatedPajama-260B QuRatedPajama Paper: QuRating: Selecting High-Quality Data for Training Language Models A 260B token subset of cerebras/SlimPajama-627B, annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria: Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers Facts & Trivia - how much factual and trivia knowledge the text contains, where specific facts and obscure trivia are preferred over more… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-260B.tabular100M<n<1B7 likes3.9k downloads2y agoHugging Face05coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.3k downloads8mo agoHugging Face06s-nlp /Mintaka_Graph_Features_T5-xl-ssm Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm" More Information needed tabular100K<n<1M0 likes1.9k downloads2y agoHugging Face07nilc-nlp /assin2 Dataset Card for ASSIN 2 Dataset Summary The ASSIN 2 corpus is composed of rather simple sentences. Following the procedures of SemEval 2014 Task 1. The training and validation data are composed, respectively, of 6,500 and 500 sentence pairs in Brazilian Portuguese, annotated for entailment and semantic similarity. Semantic similarity values range from 1 to 5, and text entailment classes are either entailment or none. The test data are composed of approximately 3,000… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin2.tabulartext-classification1K<n<10K15 likes1.8k downloads3y agoHugging Face08Finnish-NLP /mc4_3.1.0_fi_cleaned Dataset Card for "mc4_3.1.0_fi_cleaned" More Information needed tabular10M<n<100M0 likes1.4k downloads3y agoHugging Face09OALL /details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.tabular100K<n<1M0 likes1.1k downloads2y agoHugging Face10UBC-NLP /EgyMMLU Dataset Card for EgyMMLU Dataset Description Dataset Summary Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Dataset Summary EgyMMLU is a benchmark created to test the… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/EgyMMLU.tabular10K<n<100K0 likes1.1k downloads11mo agoHugging Face11causal-nlp /corr2cause Dataset card for corr2cause TODO tabular100K<n<1M31 likes1.1k downloads3y agoHugging Face12princeton-nlp /LitSearch LitSearch: A Retrieval Benchmark for Scientific Literature Search This dataset contains the query set and retrieval corpus for our paper LitSearch: A Retrieval Benchmark for Scientific Literature Search. We introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers. LitSearch is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/LitSearch.tabular100K<n<1M23 likes1k downloads2y agoHugging Face13yale-nlp /FOLIOgatedtabular1K<n<10K73 likes924 downloads3y agoHugging Face14Finnish-NLP /oscar_2301_fi_cleaned Dataset Card for "oscar_2301_fi_cleaned" More Information needed tabular1M<n<10M0 likes763 downloads3y agoHugging Face15nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes697 downloads2y agoHugging Face16SALT-NLP /hle-context-baseline-deeptabular10K<n<100K0 likes678 downloads2mo agoHugging Face17surrey-nlp /Low-resource-QE-DA-dataset Low-resource QE-DA Dataset Direct Assessment (DA) quality estimation data for English→Indic (Gujarati, Hindi, Marathi, Tamil, Telugu) and related Estonian/Nepali/Sinhala pairs, released with the ALOPE work on LLM-based QE. Paper: Sindhujan, A., Qian, S., Matthew, C.C.C., Orasan, C., and Kanojia, D. (2024). ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models. In Second Conference on Language Modeling. (arXiv) Task: Sentence-level quality… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/Low-resource-QE-DA-dataset.tabularother100K<n<1M0 likes601 downloads10mo agoHugging Face18s-nlp /KGQASubgraphsRanking 📰 News [12/2023] Publishing of the original paper "Large Language Models Meets Knowledge Graph to Answer Factoid Questions". This paper first introduces the novelty of the extracted subgraphs; which provide valuable information for different methods of ranking. The paper leveraged T5-like models, and achieve SOTA results with Graph2Text ranking. Dataset Summary KGQASubgraphsRanking is the total-packaged dataset for both publications mentioned in the News section.… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/KGQASubgraphsRanking.tabular100K<n<1M1 likes585 downloads2y agoHugging Face19NLP-FBK /multilingual-medical-reasoning-tracesThis datasets containes the traces generated to answer multiple-choice medical questions in Italian, Englihs, and Spanish. The dataset is structured in 3 parts, one per language. Each part is composed by 2 splits, one containing the examples generated from medqa, one from medmcqa. The columns are: id, representing an unique identifier full_question, representing the medical question options, a dictionary of options to answer the question and their identifiers list_of_options, a list of the… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/multilingual-medical-reasoning-traces.tabular100K<n<1M1 likes535 downloads7mo agoHugging Face20eth-nlped /mathdial Mathdial dataset https://arxiv.org/abs/2305.14536 MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems. MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching. Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.tabulartext-generation1K<n<10K18 likes440 downloads2y agoHugging Face21sixuexing /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.tabular1M<n<10M1 likes437 downloads1y agoHugging Face22hitachi-nlp /FLD.v2 Dataset Card for "FLD.v2" For the schema of the dataset, see here. For the whole of the project, see our project page. More Information needed tabular10K<n<100K15 likes420 downloads3y agoHugging Face23nyu-dice-lab /lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.tabular100K<n<1M0 likes416 downloads2y agoHugging Face24nilc-nlp /assin Dataset Card for ASSIN Dataset Summary The ASSIN (Avaliação de Similaridade Semântica e INferência textual) corpus is a corpus annotated with pairs of sentences written in Portuguese that is suitable for the exploration of textual entailment and paraphrasing classifiers. The corpus contains pairs of sentences extracted from news articles written in European Portuguese (EP) and Brazilian Portuguese (BP), obtained from Google News Portugal and Brazil, respectively. To… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin.tabulartext-classification10K<n<100K10 likes405 downloads3y agoHugging Face25nlpatunt /D_persuade_2 Persuade_2 The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements) contains over 25,000 argumentative essays written by 6th–12th grade students in the United States, covering 15 distinct prompts across two writing tasks: independent and source-based writing. The corpus also provides detailed individual and demographic information for each writer. This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.tabular10K<n<100K0 likes364 downloads6mo agoHugging Face26s-nlp /EnokiQA EnokiQA EnokiQA is an annotated dataset for fine-grained hallucination detection in long-form question answering. Each example contains a factual question, a no-context LLM answer, the full Wikipedia article used as verification evidence, sentence-grouped factual triples, and per-triple NLI and hallucination probabilities. The dataset is dual-granularity: every hallucination label is attached to a claim (an extracted triple) and projected to a character span of the answer. The… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/EnokiQA.tabularquestion-answering1K<n<10K3 likes337 downloads19d agoHugging Face27nlpatunt /D_ASAP-AES D_ASAP-AES This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation. For the original dataset with labels, see below. Original Dataset 🔗 ASAP-AES on Kaggle Citation If you use this dataset, please cite the original: @misc{asap_aes, title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.tabular10K<n<100K0 likes312 downloads6mo agoHugging Face28NLP-POL /instagram-political-communication-it Instagram Political Communication (Italy) — NLP-POL Dataset Summary This dataset is part of NLP-POL (NLP for Political Communication), a research project focused on the analysis of political communication strategies through Natural Language Processing. The dataset contains Instagram posts and comments collected from more than 300 Italian political figures, primarily members of the Italian Parliament (with a strong focus on Deputies). It includes both content published by… See the full description on the dataset page: https://huggingface.co/datasets/NLP-POL/instagram-political-communication-it.tabulartext-classification1M<n<10M3 likes302 downloads9mo agoHugging Face29McGill-NLP /TopiOCQATopiOCQA is an information-seeking conversational dataset with challenging topic switching phenomena.tabulartext-retrieval10K<n<100K10 likes291 downloads3y agoHugging Face30SALT-NLP /hle-context-baseline-gpt41tabular10K<n<100K0 likes287 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.