CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
011-800-SHARED-TASKS /COLING-2025-CHIPSAL Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-CHIPSAL.tabular100K<n<1M2 likes92 downloads2y agoHugging Face02CAMeL-Lab /BAREC-Shared-Task-2025-sent BAREC Shared Task 2025 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-sent.tabulartext-classification10K<n<100K2 likes68 downloads1y agoHugging Face031-800-SHARED-TASKS /COLING-2025-GENAI-3 🚨 RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors 🚨 🌐 Website, 🖥️ Github, 📝 Paper RAID is the largest & most comprehensive dataset for evaluating AI-generated text detectors. It contains over 10 million documents spanning 11 LLMs, 11 genres, 4 decoding strategies, and 12 adversarial attacks. It is designed to be the go-to location for trustworthy third-party evaluation of both open-source and closed-source generated text detectors. Load… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-GENAI-3.texttext-classification1M<n<10M0 likes50 downloads2y agoHugging Face04CAMeL-Lab /BAREC-Shared-Task-2025-doc BAREC Shared Task 2025 Dataset Summary BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes. The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.tabulartext-classification1K<n<10K2 likes38 downloads1y agoHugging Face051-800-SHARED-TASKS /COLING-2025-FINNLP-FMD Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-FINNLP-FMD.text1K<n<10K1 likes28 downloads2y agoHugging Face06darrow-ai /LegalLensNER-SharedTasktextn<1K0 likes24 downloads2y agoHugging Face07darrow-ai /LegalLensNLI-SharedTasktextn<1K3 likes21 downloads2y agoHugging Face081-800-SHARED-TASKS /telugu_news_classificationtexttext-classification10K<n<100K0 likes14 downloads2y agoHugging Face09kcrl /Shared_Task_Fake_News_binarytext1K<n<10K0 likes13 downloads2y agoHugging Face101-800-SHARED-TASKS /telugu-summarization-generation Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/telugu-summarization-generation.texttext-generation100K<n<1M0 likes12 downloads2y agoHugging Face11Amalq /shared_TaskAtext1K<n<10K0 likes8 downloads4y agoHugging Face12s-nlp /TextGraphs17-shared-task-datasetWe present a dataset for graph-based question answering. The dataset consists of <question; candidate answer> pairs. For each candidate, we present a graph that is obtained by finding the shortest path between named entities mentioned in a question and a candidate answer. As a knowledge graph, we adopted Wikidata. Our dataset has the following fields: sample_id - an identifier for <question, candidate answer>; question - question text; questionEntity - comma-separated list of names (textual… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/TextGraphs17-shared-task-dataset.textquestion-answering10K<n<100K0 likes8 downloads3y agoHugging Face131-800-SHARED-TASKS /COLING-2025-GENAI-MULTItext100K<n<1M1 likes7 downloads2y agoHugging Face141-800-SHARED-TASKS /telugu_teknium_GPTeacher_general_instruct_filtered_romanizedtext10K<n<100K0 likes6 downloads2y agoHugging Face151-800-SHARED-TASKS /COLING-2025-GENAI-MONOtext100K<n<1M1 likes5 downloads2y agoHugging Face16Care4langGW /shared_taskBtextn<1K0 likes3 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.