datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cord-v2WheelArm_WoZ_Pilot_Dataset
WheelArm Synchronized Dataset
A multimodal dataset of wheelchair-mounted robot arm demonstrations for assistive daily-living tasks.
Each episode captures a single task performed by a human operator and includes synchronized RGB video,
depth, robot kinematics, audio, and natural-language dialogue with ambiguity annotations.
Dataset Summary
WheelArm is a real-robot dataset collected from a Kinova Gen3 6-DOF manipulator arm mounted on
a powered wheelchair.… See the full description on the dataset page: https://huggingface.co/datasets/Cordelia/WheelArm_WoZ_Pilot_Dataset.cord-receipt-imagescord-v1cord-extraction-lora
CORD structured-extraction (LoRA training set)
200 train + 20 val examples derived from CORD (naver-clova-ix/cord-v2, train
split). Each example pairs a preprocessed receipt image with the target
extraction JSON (receipt fields + line items in the pipeline's schema).
prompt.txt - the shared instruction (schema skeleton injected)
train.jsonl / val.jsonl - lines of {"id", "image": "images/..png", "target": "<json>"}
images/ - the preprocessed pages (deskew / resize<=1536 /… See the full description on the dataset page: https://huggingface.co/datasets/sarcasticcoder/cord-extraction-lora.receipt_cord_ocr_v2
Dataset Card for "receipt_cord_ocr_v2"
More Information needed
cordhttps://huggingface.co/datasets/katanaml/cordcordcord100cord_demo_gerw9_cord_completecord-v1cord-v2cord-v2-custom
Dataset Card for "cord-v2-custom"
More Information needed
cord-ocr-text-in-image-v2
Dataset Card for "cord-ocr-text-in-image-v2"
More Information needed
layoutlmv3_cord
Dataset Card for "layoutlmv3_cord"
Original Dataset is "naver-clova-ix/cord-v2"
This dataset is modified for learning.
More Information needed
cord-v2
Neural Metrics · The receipt benchmark everyone quotes.
CORD is the standard consolidated receipt dataset, with detailed line-item and field-level annotations. If a document-understanding paper reports receipt numbers, they are usually CORD numbers.
We use it for: comparable, publishable receipt extraction scores - line-item table parsing under messy real-world layouts.
Attribution
This is an unmodified fork of naver-clova-ix/cord-v2, created by the Qwen… See the full description on the dataset page: https://huggingface.co/datasets/NeuralMetrics/cord-v2.CORD2CorDiCas
CorDiCas
CorDiCas es un prototipo de corpus diacrónico cuyos documentos proceden de una colección de más de 120 documentos inéditos de carácter semiprivado, cuya temática gira en torno a la sedentarización e inserción forzosas de la población gitana durante el siglo XVIII.
En la siguiente tabla se ofrece la información estructurada sobre los periodos que se abordan en la colección:
Signatura
Periodo
N.º textos
AMH_01430
1745 - 1746
4 textos
1748
14 textos
1749
Más… See the full description on the dataset page: https://huggingface.co/datasets/epuertas94/CorDiCas.cord-v2CORDIALearth_surface_waterw9_cord_complete_subset_aug1cord-v2cord-sroiew9_cord_ltd_fields_phila_domain_v4CORD-T1
CORD-T1
dataset_info:
features:
- name: image
dtype: image
splits:
- name: train
num_bytes: 57404876
num_examples: 15
- name: validation
num_bytes:
num_examples: 3
- name: test
num_bytes:
num_examples: 3
download_size: 57334757
dataset_size: 57404876
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
license: mit
task_categories:
- token-classification
language:
- ak
tags:
- medical… See the full description on the dataset page: https://huggingface.co/datasets/hsienchen/CORD-T1.cord_parsecord_train_cleaned
cord_train_cleaned
The cord_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
794
QA turns
2,990
answers rewritten by the cleaning pass
303
QA created by the cleaning pass (new_qa)
2,197 (73.5%)
shards
4
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but salvageable… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/cord_train_cleaned.layoutlmv3-cordv2
