CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01naver-clova-ix /cord-v2image1K<n<10K126 likes8k downloads4y agoHugging Face02allenai /cord19The Covid-19 Open Research Dataset (CORD-19) is a growing resource of scientific papers on Covid-19 and related historical coronavirus research. CORD-19 is designed to facilitate the development of text mining and information retrieval systems over its rich collection of metadata and structured full text papers. Since its release, CORD-19 has been downloaded over 75K times and has served as the basis of many Covid-19 text mining and discovery systems. The dataset itself isn't defining a specific task, but there is a Kaggle challenge that define 17 open research questions to be solved with the dataset: https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge/tasksother100K<n<1M9 likes2.7k downloads4y agoHugging Face03Cordelia /WheelArm_WoZ_Pilot_Dataset WheelArm Synchronized Dataset A multimodal dataset of wheelchair-mounted robot arm demonstrations for assistive daily-living tasks. Each episode captures a single task performed by a human operator and includes synchronized RGB video, depth, robot kinematics, audio, and natural-language dialogue with ambiguity annotations. Dataset Summary WheelArm is a real-robot dataset collected from a Kinova Gen3 6-DOF manipulator arm mounted on a powered wheelchair.… See the full description on the dataset page: https://huggingface.co/datasets/Cordelia/WheelArm_WoZ_Pilot_Dataset.imageroboticsn<1K1 likes2k downloads4mo agoHugging Face04ml-infra-toloka /cord-receipt-imagesimage1K<n<10K0 likes1.4k downloads1mo agoHugging Face05medalpaca /medical_meadow_cord19 CORD 19 Dataset Summary In response to the COVID-19 pandemic, the White House and a coalition of leading research groups have prepared the COVID-19 Open Research Dataset (CORD-19). CORD-19 is a resource of over 1,000,000 scholarly articles, including over 400,000 with full text, about COVID-19, SARS-CoV-2, and related coronaviruses. This freely available dataset is provided to the global research community to apply recent advances in natural language processing and other… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_cord19.textsummarization100K<n<1M10 likes1.3k downloads3y agoHugging Face06FengQiuxuan /CordViP0 likes811 downloads1y agoHugging Face07naver-clova-ix /cord-v1image1K<n<10K20 likes519 downloads4y agoHugging Face08Auroraky /CordViP_flipcup1 likes477 downloads11mo agoHugging Face09sileod /cordis-bench CordisBench CordisBench tests whether language models can reason about the consequences of component lifecycle changes in dynamic agent harnesses. Each record contains an exact, programmatically generated oracle. Set-valued tasks use Jaccard similarity, sequence prediction uses per-observable accuracy, and executable reconfiguration is checked by running the proposed lifecycle operations. This repository packages the frozen V2.0.1 release from sileod/cordis-bench.… See the full description on the dataset page: https://huggingface.co/datasets/sileod/cordis-bench.textquestion-answering1K<n<10K2 likes174 downloads1mo agoHugging Face10Heliosoph /CORD CORD (receipts) — reshaped mirror A reshaped mirror of naver-clova-ix/cord-v1 (Park et al., NAVER Clova, 2019), packaged for one-row-per-receipt ingestion. Upstream stores the receipt as a HuggingFace Image feature (a {bytes, path} struct) plus a ground_truth JSON string; pipelines that can't decode a struct image column (or don't want a full datasets dependency) can't consume it directly. This mirror splits each receipt into an image file + a JSONL annotation line, joinable by… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/CORD.image-to-textn<1K0 likes134 downloads3mo agoHugging Face11PassionPrc /cord-grpo0 likes131 downloads3mo agoHugging Face12MarkusDressel /cordhttps://github.com/clovaai/cord0 likes127 downloads5y agoHugging Face13sarcasticcoder /cord-extraction-lora CORD structured-extraction (LoRA training set) 200 train + 20 val examples derived from CORD (naver-clova-ix/cord-v2, train split). Each example pairs a preprocessed receipt image with the target extraction JSON (receipt fields + line items in the pipeline's schema). prompt.txt - the shared instruction (schema skeleton injected) train.jsonl / val.jsonl - lines of {"id", "image": "images/..png", "target": "<json>"} images/ - the preprocessed pages (deskew / resize<=1536 /… See the full description on the dataset page: https://huggingface.co/datasets/sarcasticcoder/cord-extraction-lora.imageimage-text-to-textn<1K0 likes120 downloads3mo agoHugging Face14PerikLab /Cord_Blood0 likes115 downloads17d agoHugging Face15buthaya /cordDescriptionThe CORD (Consolidated Receipt Dataset) dataset contains receipts annotated for key information extraction. It was released for the 2019 ICDAR competition on scanned receipts. Content 1,000 receipts (800 train/ 100 val/ 100 test) Entities include menu items, totals, store information, and dates OCR text + layout information available More fine-grained annotations than in SROIE (e.g. line items in receipts) Useful for benchmarking models on dense receipt parsing Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/buthaya/cord.tabularn<1K0 likes113 downloads1y agoHugging Face16pritamdeka /cord-19-fulltext Dataset Card for [pritamdeka/cord-19-fulltext] Dataset Description Dataset Summary This is a modified cord19 dataset which contains only the fulltext field. This can be used directly for language modelling tasks. Languages English Citation Information @article{Wang2020CORD19TC, title={CORD-19: The Covid-19 Open Research Dataset}, author={Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin… See the full description on the dataset page: https://huggingface.co/datasets/pritamdeka/cord-19-fulltext.text100K<n<1M2 likes106 downloads5y agoHugging Face17CordwainerSmith /GolemGuard GolemGuard: Hebrew Privacy Information Detection Corpus GolemGuard is a comprehensive Hebrew language dataset specifically designed for training and evaluating models for Personal Identifiable Information (PII) detection and masking. The dataset contains ~600MB of synthetic text data representing various document types and communication formats commonly found in Israeli professional and administrative contexts. Source Data Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/GolemGuard.texttoken-classification100K<n<1M0 likes105 downloads2y agoHugging Face18mychen76 /receipt_cord_ocr_v2 Dataset Card for "receipt_cord_ocr_v2" More Information needed image1K<n<10K1 likes89 downloads3y agoHugging Face19katanaml /cordhttps://huggingface.co/datasets/katanaml/cordimage1K<n<10K3 likes83 downloads5y agoHugging Face20mystic-leung /medical_cord19 Description This dataset contains large amounts of biomedical abstracts and corresponding summaries. textsummarization100K<n<1M5 likes65 downloads3y agoHugging Face21christinamaria /cord-v2-ocr CORD-v2 text-only — receipt OCR text → structured JSON A text-only derivative of CORD-v2 (Consolidated Receipt Dataset; Park et al., 2019), the standard benchmark for document information extraction used by Donut and similar models. The original dataset pairs 1,000 receipt photos with a rich ground-truth schema (~30 field types across 4 groups, including menu line items). This version drops the images and pairs the OCR text of each receipt with its target parse, so text-only… See the full description on the dataset page: https://huggingface.co/datasets/christinamaria/cord-v2-ocr.text-generationn<1K0 likes62 downloads2mo agoHugging Face22nielsr /cord-layoutlmv3https://github.com/clovaai/cord/3 likes61 downloads4y agoHugging Face23marianbasti /cordeba CordeBA: Corpus de Buenos Aires Descripción CordeBA es una colección de registros orales de conversaciones espontáneas informales entre hablantes de la provincia de Buenos Aires, Argentina. El corpus está compuesto por documentos de audio de discurso dialogal no dirigido y sus respectivas transcripciones. Esta primera versión del corpus incluye 24 registros orales, con edades media y mediana de los participantes de 29.74 y 24 años respectivamente. El objetivo principal es… See the full description on the dataset page: https://huggingface.co/datasets/marianbasti/cordeba.audion<1K1 likes61 downloads1y agoHugging Face24mp-02 /cordimagetoken-classification1K<n<10K0 likes60 downloads2y agoHugging Face25SinaAhmadi /CORDI CORDI — Corpus of Dialogues in Central Kurdish ➡️ See the repository on GitHub This repository provides resources for language and speech technology for Central Kurdish varieties discussed in our LREC-COLING 2024 paper, particularly the first annotated corpus of spoken Central Kurdish varieties — CORDI. Given the financial burden of traditional ways of documenting languages and varieties as in fieldwork, we follow a rather novel alternative where movies and series are… See the full description on the dataset page: https://huggingface.co/datasets/SinaAhmadi/CORDI.audio100K<n<1M2 likes60 downloads2y agoHugging Face26SotiriosKastanas /cord100imagen<1K0 likes59 downloads4y agoHugging Face27matg41 /cord_demo_gerimagen<1K0 likes57 downloads3y agoHugging Face28jfargus /w9_cord_completeimage1K<n<10K0 likes56 downloads1y agoHugging Face29macavaney /cord19.pisa cord19.pisa Description TODO: What is the artifact? Usage # Load the artifact import pyterrier_alpha as pta artifact = pta.Artifact.from_hf('macavaney/cord19.pisa') # TODO: Show how you use the artifact Benchmarks TODO: Provide benchmarks for the artifact. Reproduction # TODO: Show how you constructed the artifact. Metadata { "type": "sparse_index", "format": "pisa", "package_hint": "pyterrier-pisa", "stemmer":… See the full description on the dataset page: https://huggingface.co/datasets/macavaney/cord19.pisa.text-retrieval0 likes54 downloads2y agoHugging Face30zechao /cord-v1image1K<n<10K0 likes53 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.