CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nlile /NuminaMath-1.5-RL-Verifiable Dataset Card for NuminaMath-1.5-RL-Verifiable Dataset Summary NuminaMath-1.5-RL-Verifiable is a curated subset of the NuminaMath-1.5 dataset, specifically filtered to support reinforcement learning applications requiring verifiable outcomes. This collection consists of 131,063 math word problems from the original dataset that meet strict filtering criteria: all problems have definitive numerical answers, validated problem statements and solutions, and come from… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-RL-Verifiable.texttext-generation100K<n<1M10 likes8.7k downloads2y agoHugging Face02tasksource /FOL-nli Dataset Card for "FOL-nli" https://github.com/sileod/unigram/ https://arxiv.org/abs/2406.11035 Citation: @article{sileo2024scaling, title={Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars}, author={Sileo, Damien}, journal={arXiv preprint arXiv:2406.11035}, year={2024} } texttext-classification100K<n<1M3 likes426 downloads9mo agoHugging Face03nlile /math_benchmark_test_saturation LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024) This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems. Original source data: Math Word Problem Solving on MATH (Papers with Code) About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.tabularquestion-answeringn<1K0 likes183 downloads2y agoHugging Face04ruanchaves /faquad-nli Dataset Card for FaQuAD-NLI Dataset Summary FaQuAD is a Portuguese reading comprehension dataset that follows the format of the Stanford Question Answering Dataset (SQuAD). It is a pioneer Portuguese reading comprehension dataset using the challenging format of SQuAD. The dataset aims to address the problem of abundant questions sent by academics whose answers are found in available institutional documents in the Brazilian higher education system. It consists of 900… See the full description on the dataset page: https://huggingface.co/datasets/ruanchaves/faquad-nli.tabularquestion-answering1K<n<10K4 likes135 downloads1y agoHugging Face05araag2 /SemEval_NLI4CT NLI4CT: Multi-Evidence Natural Language Inference for Clinical Trial Reports and SemEval-2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials Dataset Description Links Homepage: sites.google Repository: Github2024 Paper: arXiv2023 / arXiv2024 Leaderboard: Codalab2023 Contact (Original Authors):Maël Jullien (mael.jullien@postgrad.manchester.ac.uk) Contact (Curator): Artur Guimarães (artur.guimas@gmail.com) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/araag2/SemEval_NLI4CT.texttext-classification10K<n<100K0 likes89 downloads1y agoHugging Face06nicholasKluge /faquad-nli-parquet Dataset Card for FaQuAD-NLI THIS IS A TEMPORARY COPY OF THE ORIGINAL ruanchaves/faquad-nli. WHY? As of datasets==4.0, loading scripts and trust_remote_code are no longer supported. This breaks things, like the lm-evaluation-harness-pt, which people who work with Portuguese LLMs need for running evals. As soon as ruanchaves updates his version, I'll delete this copy. Dataset Summary FaQuAD is a Portuguese reading comprehension dataset that follows the format of the… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/faquad-nli-parquet.tabularquestion-answering1K<n<10K1 likes32 downloads1y agoHugging Face07NLie2 /rewrite-questions-real-words-sciency real_words_sciency.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file real_words_sciency.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.tabulartext-generationn<1K0 likes32 downloads1y agoHugging Face08NLie2 /rewrite-questions-nonsensical-biology nonsensical_biology.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file nonsensical_biology.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-nonsensical-biology.tabulartext-generationn<1K0 likes30 downloads1y agoHugging Face09NLie2 /rewrite-questions-gibberish gibberish.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file gibberish.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the correct answer… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-gibberish.tabulartext-generationn<1K0 likes28 downloads1y agoHugging Face10Singhchandann /all_nli_marathigated All Nli Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The All Nli Marathi dataset is a meticulously curated collection of 570810 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine learning models… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/all_nli_marathi.texttext-classification100K<n<1M0 likes23 downloads2y agoHugging Face11Singhchandann /all-nli-pair-score_marathigated All-Nli-Pair-Score Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The All-Nli-Pair-Score Marathi dataset is a meticulously curated collection of 981336 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/all-nli-pair-score_marathi.texttext-classification100K<n<1M0 likes12 downloads1y agoHugging Face12Singhchandann /nli-for-simcse_marathigated Nli-For-Simcse Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The Nli-For-Simcse Marathi dataset is a meticulously curated collection of 274951 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/nli-for-simcse_marathi.texttext-classification100K<n<1M0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.