datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tweet_sentiment_extraction
TweetSentimentExtractionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Task category
t2c
Domains
Social, Written
Reference
https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentExtractionClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.tool-output-extraction-swebench
Tool Output Extraction Dataset
Paper | Code
Training data for squeez — a small model that prunes verbose coding agent tool output to only the evidence the agent needs next.
Task
Task-conditioned context pruning of a single tool observation for coding agents.
Given a focused extraction query and one verbose tool output, return the smallest verbatim evidence block(s) the agent should read next.
The model copies lines from the tool output — it never rewrites, summarizes, or… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/tool-output-extraction-swebench.scrapinghub-article-extraction-benchmark
Scrapinghub Article Extraction Benchmark
This dataset was originally created and distributed under MIT License by Scrapinghub on GitHub: github.com/scrapinghub/article-extraction-benchmark
It is mirrored on the HuggingFace Hub as a convenience.
price-tag-extraction
Price tag extraction dataset
This dataset contains images of price tags extracted from the Open Prices dataset, along with information extracted from this dataset.
It is intended to be used to train large visual language models (LVLMs) to extract information from price tag images, as part of the Open Prices project.
For more information about the format of this dataset, please refer to the documentation on llm-image-extraction datasets.
Dataset creation
A detailed… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/price-tag-extraction.paper_extractionraw-fact-extractionfunding-extraction-harness-benchmarkdanish-dynaword-extractionsturkish-keyword-extraction-500k
Turkish Keyword Extraction 500K v2
Yirmi alanda konu ve anahtar sözcük çıkarımı için kısa Türkçe belgeler.
Doğrulanmış boyut
Train: 490,000
Validation: 5,000
Test: 5,000
Toplam: 500,000
Ana görev sütunları: id, text, keywords, domain
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-keyword-extraction-500k.doab-metadata-extraction
DOAB Open Access Books - Metadata Extraction Dataset
Dataset Description
This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs.
The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/biglam/doab-metadata-extraction.resume-json-extraction-5k
Dataset Card for resume-json-extraction-5k
Dataset Description
This dataset contains 4,879 resume examples formatted for fine-tuning language models to extract structured JSON information from resume text.
Dataset Summary
The dataset consists of resume text paired with structured JSON outputs containing:
Job titles (current and previous)
Companies (current and previous)
Years of experience
Seniority level
Primary domain and industries
Core and secondary skills… See the full description on the dataset page: https://huggingface.co/datasets/sandeeppanem/resume-json-extraction-5k.danish-extraction-v1
danish-extraction-v1
Danish information-extraction rows over real prose, where the schema is
proposed per passage rather than fixed. Built from
danish-foundation-models/danish-dynaword
by scripts/gen_extraction_da.py.
Each source passage got its own field set: an LLM proposed 3-6 fields for that
text without seeing any values, then filled them in a separate turn. Roughly a
quarter of proposed fields come back empty, which are genuine abstention
targets rather than annotation… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-extraction-v1.json_data_extraction
Diverse Restricted JSON Data Extraction
Curated by: The paraloq analytics team.
Uses
Benchmark restricted JSON data extraction (text + JSON schema -> JSON instance)
Fine-Tune data extraction model (text + JSON schema -> JSON instance)
Fine-Tune JSON schema Retrieval model (text -> retriever -> most adequate JSON schema)
Out-of-Scope Use
Intended for research purposes only.
Dataset Structure
The data comes with the following fields:
title: The… See the full description on the dataset page: https://huggingface.co/datasets/paraloq/json_data_extraction.Skill-extraction-Tech-graded
skill-extraction-tech-graded
Graded-relevance annotations for sentences from
TechWolf/skill-extraction-tech
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
75
_id (sentence id), text (sentence)
queries
test
338
_id (sentence id), text (sentence)
corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-Tech-graded.Skill-extraction-SkillSkape-graded
skill-extraction-skillskape-graded
Graded-relevance annotations for sentences from
jjzha/skillskape
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
100
_id (sentence id), text (sentence)
queries
test
500
_id (sentence id), text (sentence)
corpus
corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-SkillSkape-graded.task181_outcome_extraction
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task181_outcome_extraction
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task181_outcome_extraction.Skill-extraction-House-graded
skill-extraction-house-graded
Graded-relevance annotations for sentences from
TechWolf/skill-extraction-house
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
61
_id (sentence id), text (sentence)
queries
test
261
_id (sentence id), text (sentence)… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-House-graded.tooth_extraction_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 200,
"total_frames": 76053,
"total_tasks": 1,
"total_videos": 400,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.extraction-wiki-ja
extraction-wiki-ja
This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This is a Japanese instruction-tuning dataset tailored for information extraction and structuring from Japanese Wikipedia text.
The dataset consists of instruction–response pairs automatically generated from Japanese Wikipedia articles. Instructions are created by prompting Qwen/Qwen2.5-32B-Instruct with passages from Wikipedia, and the… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/extraction-wiki-ja.openalex_extractionarxiv-funding-entity-extractions
arxiv-funding-entity-extractions
Funder/award entity extractions over cometadata/arxiv-funding-statements.
Extractor: funding-entity-extractor (vLLM + LoRA)
Base model: meta-llama/Llama-3.1-8B-Instruct
LoRA: cometadata/funding-extraction-llama-3.1-8b-instruct-artifact-data-mix-grpo-mixed-reward
Hardware: A100-large bf16, concurrency 256
Total rows: 1,823,650
Configs
predictions (default) — original extractions, no ROR enrichment.
predictions_with_ror — same rows… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-funding-entity-extractions.task1448_disease_entity_extraction_ncbi_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1448_disease_entity_extraction_ncbi_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1448_disease_entity_extraction_ncbi_dataset.tooth_extraction_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 32879,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_3.Skill-extraction-TechWolf-graded
skill-extraction-techwolf-graded
Graded-relevance annotations for sentences from
TechWolf/skill-extraction-techwolf
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
test
324
_id (sentence id), text (sentence)
corpus
corpus
13,891
_id (ESCO skill URI), title (English… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-TechWolf-graded.search-term-extraction
Multilingual Search-Term Extraction
Token-level labels marking, in a voice-assistant query, the search term — the
minimal topic string you would hand to a knowledge base or search engine. Given
"what is the speed of light?" the target is "speed of light"; given
"set volume to fifty" the target is nothing (there is no topic to look up).
The task is not document keyphrase extraction and not full intent/slot
NLU. It answers one question: what do I search for? — the input the OVOS… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/search-term-extraction.task1486_cell_extraction_anem_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1486_cell_extraction_anem_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1486_cell_extraction_anem_dataset.key_information_extractionposter-schedule-information-extraction
Multimodal Visual-Text Dataset for Poster Schedule Information Extraction
A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels.
Overview
What it is
127 Indonesian seminar and religious-study event posters with multimodal token-level annotations
Primary task
Schedule information extraction as token classification
Modalities
Image + text + 2D spatial layout
Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.task1510_evalution_relation_extraction
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1510_evalution_relation_extraction
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1510_evalution_relation_extraction.doab-metadata-extraction
DOAB Open Access Books - Metadata Extraction Dataset
Dataset Description
This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs.
The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/doab-metadata-extraction.
