datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.bag-of-documents
Bag-of-Documents: Product Search Dataset
Blog post: Distilling Retrieval Pipelines to a Single Embedding Model
Live demo: huggingface.co/spaces/dtunkelang/bag-of-documents-demo
Code: github.com/dtunkelang/bag-of-documents
Dataset Description
A large-scale bag-of-documents dataset for e-commerce product search, built on Amazon product data. Each search query is represented as a distribution of relevant products in embedding space, captured by a centroid vector and a… See the full description on the dataset page: https://huggingface.co/datasets/dtunkelang/bag-of-documents.document_dating
Dating Document Evaluation at EVALITA 2020
In the context of EVALITA 2020, we propose the task of assigning a temporal span to a document, i.e. recognising when a document was issued. The task has already been addressed in other languages, namely French, English, Polish, also in the framework of shared tasks, see for example the DÉfi Fouille de Textes (DEFT) 2010 and 2011 challenges (Grouin, 2010; Grouin, 2011), the SemEval-2015 task on Diachronic Text Evaluation (Popescu and… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/document_dating.Daemontatox__DocumentCogito-details
Dataset Card for Evaluation run of Daemontatox/DocumentCogito
Dataset automatically created during the evaluation run of model Daemontatox/DocumentCogito
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Daemontatox__DocumentCogito-details.eventx-recognition-documenttimex-recognition-document
