datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cord-v2cord19The Covid-19 Open Research Dataset (CORD-19) is a growing resource of scientific papers on Covid-19 and related
historical coronavirus research. CORD-19 is designed to facilitate the development of text mining and information
retrieval systems over its rich collection of metadata and structured full text papers. Since its release, CORD-19
has been downloaded over 75K times and has served as the basis of many Covid-19 text mining and discovery systems.
The dataset itself isn't defining a specific task, but there is a Kaggle challenge that define 17 open research
questions to be solved with the dataset: https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge/tasksWheelArm_WoZ_Pilot_Dataset
WheelArm Synchronized Dataset
A multimodal dataset of wheelchair-mounted robot arm demonstrations for assistive daily-living tasks.
Each episode captures a single task performed by a human operator and includes synchronized RGB video,
depth, robot kinematics, audio, and natural-language dialogue with ambiguity annotations.
Dataset Summary
WheelArm is a real-robot dataset collected from a Kinova Gen3 6-DOF manipulator arm mounted on
a powered wheelchair.… See the full description on the dataset page: https://huggingface.co/datasets/Cordelia/WheelArm_WoZ_Pilot_Dataset.cord-receipt-imagesmedical_meadow_cord19
CORD 19
Dataset Summary
In response to the COVID-19 pandemic, the White House and a coalition of leading research groups have prepared the COVID-19 Open Research Dataset (CORD-19). CORD-19 is a resource of over 1,000,000 scholarly articles, including over 400,000 with full text, about COVID-19, SARS-CoV-2, and related coronaviruses. This freely available dataset is provided to the global research community to apply recent advances in natural language processing and other… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_cord19.CordViPcord-v1CordViP_flipcupcordis-bench
CordisBench
CordisBench tests whether language models can reason about the consequences of
component lifecycle changes in dynamic agent harnesses. Each record contains an
exact, programmatically generated oracle. Set-valued tasks use Jaccard
similarity, sequence prediction uses per-observable accuracy, and executable
reconfiguration is checked by running the proposed lifecycle operations.
This repository packages the frozen V2.0.1 release from
sileod/cordis-bench.… See the full description on the dataset page: https://huggingface.co/datasets/sileod/cordis-bench.CORD
CORD (receipts) — reshaped mirror
A reshaped mirror of naver-clova-ix/cord-v1 (Park et al., NAVER Clova, 2019), packaged for one-row-per-receipt ingestion. Upstream stores the receipt as a HuggingFace Image feature (a {bytes, path} struct) plus a ground_truth JSON string; pipelines that can't decode a struct image column (or don't want a full datasets dependency) can't consume it directly. This mirror splits each receipt into an image file + a JSONL annotation line, joinable by… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/CORD.cord-grpocordhttps://github.com/clovaai/cordcord-extraction-lora
CORD structured-extraction (LoRA training set)
200 train + 20 val examples derived from CORD (naver-clova-ix/cord-v2, train
split). Each example pairs a preprocessed receipt image with the target
extraction JSON (receipt fields + line items in the pipeline's schema).
prompt.txt - the shared instruction (schema skeleton injected)
train.jsonl / val.jsonl - lines of {"id", "image": "images/..png", "target": "<json>"}
images/ - the preprocessed pages (deskew / resize<=1536 /… See the full description on the dataset page: https://huggingface.co/datasets/sarcasticcoder/cord-extraction-lora.Cord_BloodcordDescriptionThe CORD (Consolidated Receipt Dataset) dataset contains receipts annotated for key information extraction. It was released for the 2019 ICDAR competition on scanned receipts.
Content
1,000 receipts (800 train/ 100 val/ 100 test)
Entities include menu items, totals, store information, and dates
OCR text + layout information available
More fine-grained annotations than in SROIE (e.g. line items in receipts)
Useful for benchmarking models on dense receipt parsing
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/buthaya/cord.cord-19-fulltext
Dataset Card for [pritamdeka/cord-19-fulltext]
Dataset Description
Dataset Summary
This is a modified cord19 dataset which contains only the fulltext field. This can be used directly for language modelling tasks.
Languages
English
Citation Information
@article{Wang2020CORD19TC,
title={CORD-19: The Covid-19 Open Research Dataset},
author={Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin… See the full description on the dataset page: https://huggingface.co/datasets/pritamdeka/cord-19-fulltext.GolemGuard
GolemGuard: Hebrew Privacy Information Detection Corpus
GolemGuard is a comprehensive Hebrew language dataset specifically designed for training and evaluating models for Personal Identifiable Information (PII) detection and masking. The dataset contains ~600MB of synthetic text data representing various document types and communication formats commonly found in Israeli professional and administrative contexts.
Source Data
Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/GolemGuard.receipt_cord_ocr_v2
Dataset Card for "receipt_cord_ocr_v2"
More Information needed
cordhttps://huggingface.co/datasets/katanaml/cordmedical_cord19
Description
This dataset contains large amounts of biomedical abstracts and corresponding summaries.
cord-v2-ocr
CORD-v2 text-only — receipt OCR text → structured JSON
A text-only derivative of CORD-v2 (Consolidated Receipt Dataset; Park et al., 2019), the standard benchmark for document information extraction used by Donut and similar models. The original dataset pairs 1,000 receipt photos with a rich ground-truth schema (~30 field types across 4 groups, including menu line items). This version drops the images and pairs the OCR text of each receipt with its target parse, so text-only… See the full description on the dataset page: https://huggingface.co/datasets/christinamaria/cord-v2-ocr.cord-layoutlmv3https://github.com/clovaai/cord/cordeba
CordeBA: Corpus de Buenos Aires
Descripción
CordeBA es una colección de registros orales de conversaciones espontáneas informales entre hablantes de la provincia de Buenos Aires, Argentina. El corpus está compuesto por documentos de audio de discurso dialogal no dirigido y sus respectivas transcripciones.
Esta primera versión del corpus incluye 24 registros orales, con edades media y mediana de los participantes de 29.74 y 24 años respectivamente. El objetivo principal es… See the full description on the dataset page: https://huggingface.co/datasets/marianbasti/cordeba.cordCORDI
CORDI — Corpus of Dialogues in Central Kurdish
➡️ See the repository on GitHub
This repository provides resources for language and speech technology for Central Kurdish varieties discussed in our LREC-COLING 2024 paper, particularly the first annotated corpus of spoken Central Kurdish varieties — CORDI. Given the financial burden of traditional ways of documenting languages and varieties as in fieldwork, we follow a rather novel alternative where movies and series are… See the full description on the dataset page: https://huggingface.co/datasets/SinaAhmadi/CORDI.cord100cord_demo_gerw9_cord_completecord19.pisa
cord19.pisa
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('macavaney/cord19.pisa')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "sparse_index",
"format": "pisa",
"package_hint": "pyterrier-pisa",
"stemmer":… See the full description on the dataset page: https://huggingface.co/datasets/macavaney/cord19.pisa.cord-v1
