datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cordDescriptionThe CORD (Consolidated Receipt Dataset) dataset contains receipts annotated for key information extraction. It was released for the 2019 ICDAR competition on scanned receipts.
Content
1,000 receipts (800 train/ 100 val/ 100 test)
Entities include menu items, totals, store information, and dates
OCR text + layout information available
More fine-grained annotations than in SROIE (e.g. line items in receipts)
Useful for benchmarking models on dense receipt parsing
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/buthaya/cord.receipt-ser-cord-plus-coru-reconciled-v1adaption-cordel-factual-nordeste
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-cordel_factual_nordeste
This dataset contains pairs of prompts and completions where Brazilian 'cordel' poetry is generated in sextet stanzas based strictly on provided factual texts about Northeastern Brazilian culture, history, and geography. The source texts cover topics such as Frevo, the Cangaço (Lampião and Maria Bonita), the Caruaru Fair, and the Caatinga biome, with… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-cordel-factual-nordeste.
