CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nlile /NuminaMath-1.5-RL-Verifiable Dataset Card for NuminaMath-1.5-RL-Verifiable Dataset Summary NuminaMath-1.5-RL-Verifiable is a curated subset of the NuminaMath-1.5 dataset, specifically filtered to support reinforcement learning applications requiring verifiable outcomes. This collection consists of 131,063 math word problems from the original dataset that meet strict filtering criteria: all problems have definitive numerical answers, validated problem statements and solutions, and come from… See the full description on the dataset page: https://huggingface.co/datasets/nlile/NuminaMath-1.5-RL-Verifiable.texttext-generation100K<n<1M10 likes8.6k downloads2y agoHugging Face02nlile /24-game Math Twenty Four (24s Game) Dataset A comprehensive dataset for the classic math twenty four game (also known as the 4 numbers game / 24s game / Game of 24). This dataset of mathematical reasoning challenges was collected from 4nums.com, featuring over 1,300 unique puzzles of the Game of 24, with difficulty metrics derived from over 6.4 million human solution attempts since 2012. In each puzzle, players must use exactly four numbers and basic arithmetic operations (+, -, ×, /) to… See the full description on the dataset page: https://huggingface.co/datasets/nlile/24-game.tabularmultiple-choice1K<n<10K14 likes6.1k downloads2y agoHugging Face03aisingapore /NLR-NLIgated SEA Abstractive Summarization SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from IndoNLI for Indonesian, IndicXNLI for Tamil, and XNLI for Thai and Vietnamese. Supported Tasks and Leaderboards SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLR-NLI.texttext-generation1K<n<10K0 likes2.1k downloads9mo agoHugging Face04tasksource /FOL-nli Dataset Card for "FOL-nli" https://github.com/sileod/unigram/ https://arxiv.org/abs/2406.11035 Citation: @article{sileo2024scaling, title={Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars}, author={Sileo, Damien}, journal={arXiv preprint arXiv:2406.11035}, year={2024} } texttext-classification100K<n<1M3 likes443 downloads9mo agoHugging Face05nlile /math_benchmark_test_saturation LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024) This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems. Original source data: Math Word Problem Solving on MATH (Papers with Code) About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.tabularquestion-answeringn<1K0 likes182 downloads2y agoHugging Face06ctu-aic /nli_it_collection Dataset Card for Natural Language Inference Instruction Tuning Collection This dataset is a collection of various NLI datasets in Czech and English, transformed into an instruction tuning format based on the FLAN approach. Dataset Details Dataset Description This dataset is a collection of English and Czech NLI datasets. Its primary purpose is instruction tuning (supervised fine tuning) of decoder LLMs. The used datasets were converted using a FLAN-like… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/nli_it_collection.texttext-generation1M<n<10M0 likes107 downloads1y agoHugging Face07Lots-of-LoRAs /task936_defeasible_nli_snli_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task936_defeasible_nli_snli_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task936_defeasible_nli_snli_classification.texttext-generation1K<n<10K0 likes90 downloads2y agoHugging Face08nlip /DIWALI DIWALI - Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context Paper | Code | Project page We present a novel Culture Specific Items (CSIs) dataset for Indian culture covering 17 facets. Please refer to our project pagefor quick details. Facets considered: food, dance, festivals, names, jewellery, places, traditions, languages, clothing, games, rituals, architectures, drinks, arts… See the full description on the dataset page: https://huggingface.co/datasets/nlip/DIWALI.texttext-generation1K<n<10K7 likes59 downloads5mo agoHugging Face09nlile /peerreview-bench PeerReview Bench Expert-annotated review items from scientific papers, organized for three complementary evaluation tasks. All data in this dataset is intended for evaluation, not training. All configs reference a shared, deduplicated file store (submitted_papers) via SHA256 content hashes. Every config exposes a single eval split. Configs reviewer For evaluating AI reviewers (models that generate reviews from a paper). One row per paper. Minimal fields:… See the full description on the dataset page: https://huggingface.co/datasets/nlile/peerreview-bench.tabulartext-classification10K<n<100K0 likes40 downloads6mo agoHugging Face10NLie2 /rewrite-questions-nonsensical-biology nonsensical_biology.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file nonsensical_biology.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-nonsensical-biology.tabulartext-generationn<1K0 likes33 downloads1y agoHugging Face11NLie2 /rewrite-questions-real-words-sciency real_words_sciency.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file real_words_sciency.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.tabulartext-generationn<1K0 likes32 downloads1y agoHugging Face12Lots-of-LoRAs /task937_defeasible_nli_social_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task937_defeasible_nli_social_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task937_defeasible_nli_social_classification.texttext-generation1K<n<10K0 likes28 downloads2y agoHugging Face13NLie2 /rewrite-questions-gibberish gibberish.csv - Question Rewriting Dataset This dataset contains question rewriting outputs from the file gibberish.csv. Dataset Structure The dataset contains the following columns: custom_id: Unique identifier for each question style: Rewriting style applied (e.g., "gibberish") index: Numerical index original: Original question text rewritten: Rewritten version of the question options: Multiple choice options (list format) correct: Index of the correct answer… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-gibberish.tabulartext-generationn<1K0 likes28 downloads1y agoHugging Face14Singhchandann /all_nli_marathigated All Nli Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The All Nli Marathi dataset is a meticulously curated collection of 570810 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine learning models… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/all_nli_marathi.texttext-classification100K<n<1M0 likes23 downloads2y agoHugging Face15Lots-of-LoRAs /task935_defeasible_nli_atomic_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task935_defeasible_nli_atomic_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task935_defeasible_nli_atomic_classification.texttext-generation1K<n<10K0 likes20 downloads2y agoHugging Face16Singhchandann /all-nli-pair-score_marathigated All-Nli-Pair-Score Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The All-Nli-Pair-Score Marathi dataset is a meticulously curated collection of 981336 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/all-nli-pair-score_marathi.texttext-classification100K<n<1M0 likes12 downloads1y agoHugging Face17Singhchandann /nli-for-simcse_marathigated Nli-For-Simcse Marathi Dataset: High-Quality Marathi NLP Corpus 📌 Overview The Nli-For-Simcse Marathi dataset is a meticulously curated collection of 274951 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness. This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/nli-for-simcse_marathi.texttext-classification100K<n<1M0 likes11 downloads1y agoHugging Face18Ponimash /nli_dataset NLI Task Specification Dataset (MASL) Авторы: Рудаков И. С., Понимаш З. А Описание датасета Синтетический датасет для задачи преобразования команд на естественном языке в структурированное описание задач с использованием формата MASL (Multi-agent system language). Датасет содержит 9 999 примеров пар: Входные данные: текстовая команда пользователя на русском языке Выходные данные: JSON структура с входными и выходными коннекторами Назначение Датасет… See the full description on the dataset page: https://huggingface.co/datasets/Ponimash/nli_dataset.texttranslation1K<n<10K1 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.