datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cord-v2medical_meadow_cord19
CORD 19
Dataset Summary
In response to the COVID-19 pandemic, the White House and a coalition of leading research groups have prepared the COVID-19 Open Research Dataset (CORD-19). CORD-19 is a resource of over 1,000,000 scholarly articles, including over 400,000 with full text, about COVID-19, SARS-CoV-2, and related coronaviruses. This freely available dataset is provided to the global research community to apply recent advances in natural language processing and other… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_cord19.cord-v1cordis-bench
CordisBench
CordisBench tests whether language models can reason about the consequences of
component lifecycle changes in dynamic agent harnesses. Each record contains an
exact, programmatically generated oracle. Set-valued tasks use Jaccard
similarity, sequence prediction uses per-observable accuracy, and executable
reconfiguration is checked by running the proposed lifecycle operations.
This repository packages the frozen V2.0.1 release from
sileod/cordis-bench.… See the full description on the dataset page: https://huggingface.co/datasets/sileod/cordis-bench.cordDescriptionThe CORD (Consolidated Receipt Dataset) dataset contains receipts annotated for key information extraction. It was released for the 2019 ICDAR competition on scanned receipts.
Content
1,000 receipts (800 train/ 100 val/ 100 test)
Entities include menu items, totals, store information, and dates
OCR text + layout information available
More fine-grained annotations than in SROIE (e.g. line items in receipts)
Useful for benchmarking models on dense receipt parsing
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/buthaya/cord.cord-19-fulltext
Dataset Card for [pritamdeka/cord-19-fulltext]
Dataset Description
Dataset Summary
This is a modified cord19 dataset which contains only the fulltext field. This can be used directly for language modelling tasks.
Languages
English
Citation Information
@article{Wang2020CORD19TC,
title={CORD-19: The Covid-19 Open Research Dataset},
author={Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin… See the full description on the dataset page: https://huggingface.co/datasets/pritamdeka/cord-19-fulltext.GolemGuard
GolemGuard: Hebrew Privacy Information Detection Corpus
GolemGuard is a comprehensive Hebrew language dataset specifically designed for training and evaluating models for Personal Identifiable Information (PII) detection and masking. The dataset contains ~600MB of synthetic text data representing various document types and communication formats commonly found in Israeli professional and administrative contexts.
Source Data
Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/GolemGuard.receipt_cord_ocr_v2
Dataset Card for "receipt_cord_ocr_v2"
More Information needed
cordhttps://huggingface.co/datasets/katanaml/cordmedical_cord19
Description
This dataset contains large amounts of biomedical abstracts and corresponding summaries.
cordeba
CordeBA: Corpus de Buenos Aires
Descripción
CordeBA es una colección de registros orales de conversaciones espontáneas informales entre hablantes de la provincia de Buenos Aires, Argentina. El corpus está compuesto por documentos de audio de discurso dialogal no dirigido y sus respectivas transcripciones.
Esta primera versión del corpus incluye 24 registros orales, con edades media y mediana de los participantes de 29.74 y 24 años respectivamente. El objetivo principal es… See the full description on the dataset page: https://huggingface.co/datasets/marianbasti/cordeba.cordCORDI
CORDI — Corpus of Dialogues in Central Kurdish
➡️ See the repository on GitHub
This repository provides resources for language and speech technology for Central Kurdish varieties discussed in our LREC-COLING 2024 paper, particularly the first annotated corpus of spoken Central Kurdish varieties — CORDI. Given the financial burden of traditional ways of documenting languages and varieties as in fieldwork, we follow a rather novel alternative where movies and series are… See the full description on the dataset page: https://huggingface.co/datasets/SinaAhmadi/CORDI.cord100cord_demo_gerw9_cord_completecord-v1cord-v2cord-v2-custom
Dataset Card for "cord-v2-custom"
More Information needed
cord-ocr-text-in-image-v2
Dataset Card for "cord-ocr-text-in-image-v2"
More Information needed
layoutlmv3_cord
Dataset Card for "layoutlmv3_cord"
Original Dataset is "naver-clova-ix/cord-v2"
This dataset is modified for learning.
More Information needed
CustomerPersonas
Synthetic Customer Experience Persona
Overview
The Synthetic Customer Experience Persona Dataset is a large-scale synthetic corpus of customer service personas, designed to aid in the development and evaluation of AI models for customer service applications. Inspired by Tencent AI Labs' Persona Hub, this dataset provides a diverse array of customer profiles across multiple industries.
Dataset Statistics
Total Personas: 250,000
Industries Covered: 6 (Retail… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/CustomerPersonas.sharegpt_formatted_cord19_fulltextcord-v2
Neural Metrics · The receipt benchmark everyone quotes.
CORD is the standard consolidated receipt dataset, with detailed line-item and field-level annotations. If a document-understanding paper reports receipt numbers, they are usually CORD numbers.
We use it for: comparable, publishable receipt extraction scores - line-item table parsing under messy real-world layouts.
Attribution
This is an unmodified fork of naver-clova-ix/cord-v2, created by the Qwen… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/cord-v2.CORD19-init-160kCorDiCas
CorDiCas
CorDiCas es un prototipo de corpus diacrónico cuyos documentos proceden de una colección de más de 120 documentos inéditos de carácter semiprivado, cuya temática gira en torno a la sedentarización e inserción forzosas de la población gitana durante el siglo XVIII.
En la siguiente tabla se ofrece la información estructurada sobre los periodos que se abordan en la colección:
Signatura
Periodo
N.º textos
AMH_01430
1745 - 1746
4 textos
1748
14 textos
1749
Más… See the full description on the dataset page: https://huggingface.co/datasets/epuertas94/CorDiCas.CORDeu-funding-cordis-qa1000-CORD19-Papers-Textcord-v2
