datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical_meadow_cord19
CORD 19
Dataset Summary
In response to the COVID-19 pandemic, the White House and a coalition of leading research groups have prepared the COVID-19 Open Research Dataset (CORD-19). CORD-19 is a resource of over 1,000,000 scholarly articles, including over 400,000 with full text, about COVID-19, SARS-CoV-2, and related coronaviruses. This freely available dataset is provided to the global research community to apply recent advances in natural language processing and other… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_cord19.cordis-bench
CordisBench
CordisBench tests whether language models can reason about the consequences of
component lifecycle changes in dynamic agent harnesses. Each record contains an
exact, programmatically generated oracle. Set-valued tasks use Jaccard
similarity, sequence prediction uses per-observable accuracy, and executable
reconfiguration is checked by running the proposed lifecycle operations.
This repository packages the frozen V2.0.1 release from
sileod/cordis-bench.… See the full description on the dataset page: https://huggingface.co/datasets/sileod/cordis-bench.cordDescriptionThe CORD (Consolidated Receipt Dataset) dataset contains receipts annotated for key information extraction. It was released for the 2019 ICDAR competition on scanned receipts.
Content
1,000 receipts (800 train/ 100 val/ 100 test)
Entities include menu items, totals, store information, and dates
OCR text + layout information available
More fine-grained annotations than in SROIE (e.g. line items in receipts)
Useful for benchmarking models on dense receipt parsing
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/buthaya/cord.GolemGuard
GolemGuard: Hebrew Privacy Information Detection Corpus
GolemGuard is a comprehensive Hebrew language dataset specifically designed for training and evaluating models for Personal Identifiable Information (PII) detection and masking. The dataset contains ~600MB of synthetic text data representing various document types and communication formats commonly found in Israeli professional and administrative contexts.
Source Data
Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/GolemGuard.CustomerPersonas
Synthetic Customer Experience Persona
Overview
The Synthetic Customer Experience Persona Dataset is a large-scale synthetic corpus of customer service personas, designed to aid in the development and evaluation of AI models for customer service applications. Inspired by Tencent AI Labs' Persona Hub, this dataset provides a diverse array of customer profiles across multiple industries.
Dataset Statistics
Total Personas: 250,000
Industries Covered: 6 (Retail… See the full description on the dataset page: https://huggingface.co/datasets/CordwainerSmith/CustomerPersonas.CORDIALreceipt-ser-cord-plus-coru-reconciled-v1adaption-cordel-factual-nordeste
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-cordel_factual_nordeste
This dataset contains pairs of prompts and completions where Brazilian 'cordel' poetry is generated in sextet stanzas based strictly on provided factual texts about Northeastern Brazilian culture, history, and geography. The source texts cover topics such as Frevo, the Cangaço (Lampião and Maria Bonita), the Caruaru Fair, and the Caatinga biome, with… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-cordel-factual-nordeste.reddit-ProgrammerHumor-testcordaidataset for cordpluß stfu dont use ik its opensource.
cord19_chunked_300_words
