CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01naver-clova-ix /cord-v2image1K<n<10K126 likes8k downloads4y agoHugging Face02medalpaca /medical_meadow_cord19 CORD 19 Dataset Summary In response to the COVID-19 pandemic, the White House and a coalition of leading research groups have prepared the COVID-19 Open Research Dataset (CORD-19). CORD-19 is a resource of over 1,000,000 scholarly articles, including over 400,000 with full text, about COVID-19, SARS-CoV-2, and related coronaviruses. This freely available dataset is provided to the global research community to apply recent advances in natural language processing and other… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_cord19.textsummarization100K<n<1M10 likes1.3k downloads3y agoHugging Face03naver-clova-ix /cord-v1image1K<n<10K20 likes519 downloads4y agoHugging Face04sileod /cordis-bench CordisBench CordisBench tests whether language models can reason about the consequences of component lifecycle changes in dynamic agent harnesses. Each record contains an exact, programmatically generated oracle. Set-valued tasks use Jaccard similarity, sequence prediction uses per-observable accuracy, and executable reconfiguration is checked by running the proposed lifecycle operations. This repository packages the frozen V2.0.1 release from sileod/cordis-bench.… See the full description on the dataset page: https://huggingface.co/datasets/sileod/cordis-bench.textquestion-answering1K<n<10K2 likes174 downloads1mo agoHugging Face05buthaya /cordDescriptionThe CORD (Consolidated Receipt Dataset) dataset contains receipts annotated for key information extraction. It was released for the 2019 ICDAR competition on scanned receipts. Content 1,000 receipts (800 train/ 100 val/ 100 test) Entities include menu items, totals, store information, and dates OCR text + layout information available More fine-grained annotations than in SROIE (e.g. line items in receipts) Useful for benchmarking models on dense receipt parsing Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/buthaya/cord.tabularn<1K0 likes113 downloads1y agoHugging Face06pritamdeka /cord-19-fulltext Dataset Card for [pritamdeka/cord-19-fulltext] Dataset Description Dataset Summary This is a modified cord19 dataset which contains only the fulltext field. This can be used directly for language modelling tasks. Languages English Citation Information @article{Wang2020CORD19TC, title={CORD-19: The Covid-19 Open Research Dataset}, author={Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin… See the full description on the dataset page: https://huggingface.co/datasets/pritamdeka/cord-19-fulltext.text100K<n<1M2 likes106 downloads5y agoHugging Face07CordwainerSmith /GolemGuard GolemGuard: Hebrew Privacy Information Detection Corpus GolemGuard is a comprehensive Hebrew language dataset specifically designed for training and evaluating models for Personal Identifiable Information (PII) detection and masking. The dataset contains ~600MB of synthetic text data representing various document types and communication formats commonly found in Israeli professional and administrative contexts. Source Data Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/GolemGuard.texttoken-classification100K<n<1M0 likes105 downloads2y agoHugging Face08mychen76 /receipt_cord_ocr_v2 Dataset Card for "receipt_cord_ocr_v2" More Information needed image1K<n<10K1 likes89 downloads3y agoHugging Face09katanaml /cordhttps://huggingface.co/datasets/katanaml/cordimage1K<n<10K3 likes83 downloads5y agoHugging Face10mystic-leung /medical_cord19 Description This dataset contains large amounts of biomedical abstracts and corresponding summaries. textsummarization100K<n<1M5 likes65 downloads3y agoHugging Face11marianbasti /cordeba CordeBA: Corpus de Buenos Aires Descripción CordeBA es una colección de registros orales de conversaciones espontáneas informales entre hablantes de la provincia de Buenos Aires, Argentina. El corpus está compuesto por documentos de audio de discurso dialogal no dirigido y sus respectivas transcripciones. Esta primera versión del corpus incluye 24 registros orales, con edades media y mediana de los participantes de 29.74 y 24 años respectivamente. El objetivo principal es… See the full description on the dataset page: https://huggingface.co/datasets/marianbasti/cordeba.audion<1K1 likes61 downloads1y agoHugging Face12mp-02 /cordimagetoken-classification1K<n<10K0 likes60 downloads2y agoHugging Face13SinaAhmadi /CORDI CORDI — Corpus of Dialogues in Central Kurdish ➡️ See the repository on GitHub This repository provides resources for language and speech technology for Central Kurdish varieties discussed in our LREC-COLING 2024 paper, particularly the first annotated corpus of spoken Central Kurdish varieties — CORDI. Given the financial burden of traditional ways of documenting languages and varieties as in fieldwork, we follow a rather novel alternative where movies and series are… See the full description on the dataset page: https://huggingface.co/datasets/SinaAhmadi/CORDI.audio100K<n<1M2 likes60 downloads2y agoHugging Face14SotiriosKastanas /cord100imagen<1K0 likes59 downloads4y agoHugging Face15matg41 /cord_demo_gerimagen<1K0 likes57 downloads3y agoHugging Face16jfargus /w9_cord_completeimage1K<n<10K0 likes56 downloads1y agoHugging Face17zechao /cord-v1image1K<n<10K0 likes53 downloads6mo agoHugging Face18tejas2102 /cord-v2image1K<n<10K0 likes52 downloads3mo agoHugging Face19Nyaaneet /cord-v2-custom Dataset Card for "cord-v2-custom" More Information needed image1K<n<10K1 likes50 downloads4y agoHugging Face20mychen76 /cord-ocr-text-in-image-v2 Dataset Card for "cord-ocr-text-in-image-v2" More Information needed imagen<1K5 likes50 downloads3y agoHugging Face21IBoH /layoutlmv3_cord Dataset Card for "layoutlmv3_cord" Original Dataset is "naver-clova-ix/cord-v2" This dataset is modified for learning. More Information needed image1K<n<10K1 likes49 downloads3y agoHugging Face22CordwainerSmith /CustomerPersonas Synthetic Customer Experience Persona Overview The Synthetic Customer Experience Persona Dataset is a large-scale synthetic corpus of customer service personas, designed to aid in the development and evaluation of AI models for customer service applications. Inspired by Tencent AI Labs' Persona Hub, this dataset provides a diverse array of customer profiles across multiple industries. Dataset Statistics Total Personas: 250,000 Industries Covered: 6 (Retail… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/CustomerPersonas.texttext-generation100K<n<1M1 likes49 downloads2y agoHugging Face23PranavHarshan /sharegpt_formatted_cord19_fulltexttext100K<n<1M0 likes48 downloads2y agoHugging Face24NeuralMetrics /cord-v2 Neural Metrics · The receipt benchmark everyone quotes. CORD is the standard consolidated receipt dataset, with detailed line-item and field-level annotations. If a document-understanding paper reports receipt numbers, they are usually CORD numbers. We use it for: comparable, publishable receipt extraction scores - line-item table parsing under messy real-world layouts. Attribution This is an unmodified fork of naver-clova-ix/cord-v2, created by the Qwen… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/cord-v2.image1K<n<10K0 likes48 downloads2mo agoHugging Face25CShorten /CORD19-init-160ktext100K<n<1M0 likes40 downloads4y agoHugging Face26epuertas94 /CorDiCas CorDiCas CorDiCas es un prototipo de corpus diacrónico cuyos documentos proceden de una colección de más de 120 documentos inéditos de carácter semiprivado, cuya temática gira en torno a la sedentarización e inserción forzosas de la población gitana durante el siglo XVIII. En la siguiente tabla se ofrece la información estructurada sobre los periodos que se abordan en la colección: Signatura Periodo N.º textos AMH_01430 1745 - 1746 4 textos 1748 14 textos 1749 Más… See the full description on the dataset page: https://huggingface.co/datasets/epuertas94/CorDiCas.imagetext-classificationn<1K1 likes40 downloads2y agoHugging Face27David19930 /CORDtextn<1K0 likes39 downloads2y agoHugging Face28RCaz /eu-funding-cordis-qatext100K<n<1M0 likes38 downloads5mo agoHugging Face29CShorten /1000-CORD19-Papers-Texttext10K<n<100K2 likes36 downloads4y agoHugging Face30tasiam /cord-v2image1K<n<10K0 likes36 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.