CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /tweet_sentiment_extraction TweetSentimentExtractionClassification An MTEB dataset Massive Text Embedding Benchmark Task category t2c Domains Social, Written Reference https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["TweetSentimentExtractionClassification"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.texttext-classification10K<n<100K38 likes5.8k downloads1y agoHugging Face02KRLabsOrg /tool-output-extraction-swebench Tool Output Extraction Dataset Paper | Code Training data for squeez — a small model that prunes verbose coding agent tool output to only the evidence the agent needs next. Task Task-conditioned context pruning of a single tool observation for coding agents. Given a focused extraction query and one verbose tool output, return the smallest verbatim evidence block(s) the agent should read next. The model copies lines from the tool output — it never rewrites, summarizes, or… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/tool-output-extraction-swebench.texttext-generation10K<n<100K5 likes1.2k downloads6mo agoHugging Face03allenai /scrapinghub-article-extraction-benchmark Scrapinghub Article Extraction Benchmark This dataset was originally created and distributed under MIT License by Scrapinghub on GitHub: github.com/scrapinghub/article-extraction-benchmark It is mirrored on the HuggingFace Hub as a convenience. textn<1K0 likes921 downloads3y agoHugging Face04openfoodfacts /price-tag-extraction Price tag extraction dataset This dataset contains images of price tags extracted from the Open Prices dataset, along with information extracted from this dataset. It is intended to be used to train large visual language models (LVLMs) to extract information from price tag images, as part of the Open Prices project. For more information about the format of this dataset, please refer to the documentation on llm-image-extraction datasets. Dataset creation A detailed… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/price-tag-extraction.image10K<n<100K2 likes629 downloads8mo agoHugging Face05zaaabik /paper_extractiontabular1M<n<10M0 likes554 downloads1d agoHugging Face06obalcells /raw-fact-extractiontabular100K<n<1M0 likes465 downloads1y agoHugging Face07cometadata /funding-extraction-harness-benchmarktabular10K<n<100K0 likes275 downloads6mo agoHugging Face08syvai /danish-dynaword-extractionsgatedtext10K<n<100K0 likes232 downloads2mo agoHugging Face09GoktugD /turkish-keyword-extraction-500k Turkish Keyword Extraction 500K v2 Yirmi alanda konu ve anahtar sözcük çıkarımı için kısa Türkçe belgeler. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, text, keywords, domain Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-keyword-extraction-500k.texttoken-classification100K<n<1M0 likes231 downloads1mo agoHugging Face10biglam /doab-metadata-extraction DOAB Open Access Books - Metadata Extraction Dataset Dataset Description This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs. The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/biglam/doab-metadata-extraction.imageimage-to-text1K<n<10K14 likes221 downloads11mo agoHugging Face11sandeeppanem /resume-json-extraction-5k Dataset Card for resume-json-extraction-5k Dataset Description This dataset contains 4,879 resume examples formatted for fine-tuning language models to extract structured JSON information from resume text. Dataset Summary The dataset consists of resume text paired with structured JSON outputs containing: Job titles (current and previous) Companies (current and previous) Years of experience Seniority level Primary domain and industries Core and secondary skills… See the full description on the dataset page: https://huggingface.co/datasets/sandeeppanem/resume-json-extraction-5k.texttext-generation1K<n<10K0 likes215 downloads8mo agoHugging Face12jensjepsen /danish-extraction-v1 danish-extraction-v1 Danish information-extraction rows over real prose, where the schema is proposed per passage rather than fixed. Built from danish-foundation-models/danish-dynaword by scripts/gen_extraction_da.py. Each source passage got its own field set: an LLM proposed 3-6 fields for that text without seeing any values, then filled them in a separate turn. Roughly a quarter of proposed fields come back empty, which are genuine abstention targets rather than annotation… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-extraction-v1.texttext-generation100K<n<1M0 likes212 downloads23d agoHugging Face13paraloq /json_data_extraction Diverse Restricted JSON Data Extraction Curated by: The paraloq analytics team. Uses Benchmark restricted JSON data extraction (text + JSON schema -> JSON instance) Fine-Tune data extraction model (text + JSON schema -> JSON instance) Fine-Tune JSON schema Retrieval model (text -> retriever -> most adequate JSON schema) Out-of-Scope Use Intended for research purposes only. Dataset Structure The data comes with the following fields: title: The… See the full description on the dataset page: https://huggingface.co/datasets/paraloq/json_data_extraction.texttext-generationn<1K34 likes211 downloads3y agoHugging Face14TechWolf /Skill-extraction-Tech-graded skill-extraction-tech-graded Graded-relevance annotations for sentences from TechWolf/skill-extraction-tech against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 75 _id (sentence id), text (sentence) queries test 338 _id (sentence id), text (sentence) corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-Tech-graded.text1M<n<10M0 likes201 downloads1mo agoHugging Face15TechWolf /Skill-extraction-SkillSkape-graded skill-extraction-skillskape-graded Graded-relevance annotations for sentences from jjzha/skillskape against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 100 _id (sentence id), text (sentence) queries test 500 _id (sentence id), text (sentence) corpus corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-SkillSkape-graded.text1M<n<10M0 likes194 downloads1mo agoHugging Face16Lots-of-LoRAs /task181_outcome_extraction Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task181_outcome_extraction Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task181_outcome_extraction.texttext-generationn<1K0 likes187 downloads2y agoHugging Face17TechWolf /Skill-extraction-House-graded skill-extraction-house-graded Graded-relevance annotations for sentences from TechWolf/skill-extraction-house against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 61 _id (sentence id), text (sentence) queries test 261 _id (sentence id), text (sentence)… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-House-graded.text1M<n<10M1 likes179 downloads1mo agoHugging Face18samsam0510 /tooth_extraction_4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 200, "total_frames": 76053, "total_tasks": 1, "total_videos": 400, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:200" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.tabularrobotics10K<n<100K0 likes176 downloads2y agoHugging Face19llm-jp /extraction-wiki-ja extraction-wiki-ja This repository provides an instruction-tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This is a Japanese instruction-tuning dataset tailored for information extraction and structuring from Japanese Wikipedia text. The dataset consists of instruction–response pairs automatically generated from Japanese Wikipedia articles. Instructions are created by prompting Qwen/Qwen2.5-32B-Instruct with passages from Wikipedia, and the… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/extraction-wiki-ja.texttext-generation100K<n<1M4 likes154 downloads1y agoHugging Face20LLMDH /openalex_extractiontext10M<n<100M1 likes153 downloads2y agoHugging Face21cometadata /arxiv-funding-entity-extractions arxiv-funding-entity-extractions Funder/award entity extractions over cometadata/arxiv-funding-statements. Extractor: funding-entity-extractor (vLLM + LoRA) Base model: meta-llama/Llama-3.1-8B-Instruct LoRA: cometadata/funding-extraction-llama-3.1-8b-instruct-artifact-data-mix-grpo-mixed-reward Hardware: A100-large bf16, concurrency 256 Total rows: 1,823,650 Configs predictions (default) — original extractions, no ROR enrichment. predictions_with_ror — same rows… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-funding-entity-extractions.tabular1M<n<10M1 likes153 downloads5mo agoHugging Face22Lots-of-LoRAs /task1448_disease_entity_extraction_ncbi_dataset Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1448_disease_entity_extraction_ncbi_dataset Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1448_disease_entity_extraction_ncbi_dataset.texttext-generationn<1K0 likes146 downloads2y agoHugging Face23samsam0510 /tooth_extraction_3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 100, "total_frames": 32879, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_3.tabularrobotics10K<n<100K0 likes139 downloads2y agoHugging Face24TechWolf /Skill-extraction-TechWolf-graded skill-extraction-techwolf-graded Graded-relevance annotations for sentences from TechWolf/skill-extraction-techwolf against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries test 324 _id (sentence id), text (sentence) corpus corpus 13,891 _id (ESCO skill URI), title (English… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-TechWolf-graded.text1M<n<10M0 likes127 downloads23d agoHugging Face25TigreGotico /search-term-extraction Multilingual Search-Term Extraction Token-level labels marking, in a voice-assistant query, the search term — the minimal topic string you would hand to a knowledge base or search engine. Given "what is the speed of light?" the target is "speed of light"; given "set volume to fifty" the target is nothing (there is no topic to look up). The task is not document keyphrase extraction and not full intent/slot NLU. It answers one question: what do I search for? — the input the OVOS… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/search-term-extraction.texttoken-classification10K<n<100K0 likes123 downloads4mo agoHugging Face26Lots-of-LoRAs /task1486_cell_extraction_anem_dataset Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1486_cell_extraction_anem_dataset Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1486_cell_extraction_anem_dataset.texttext-generationn<1K0 likes121 downloads2y agoHugging Face27nanonets /key_information_extractiontextquestion-answeringn<1K6 likes116 downloads1y agoHugging Face28Ibnuck /poster-schedule-information-extraction Multimodal Visual-Text Dataset for Poster Schedule Information Extraction A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels. Overview What it is 127 Indonesian seminar and religious-study event posters with multimodal token-level annotations Primary task Schedule information extraction as token classification Modalities Image + text + 2D spatial layout Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.imagetoken-classificationn<1K1 likes115 downloads15d agoHugging Face29Lots-of-LoRAs /task1510_evalution_relation_extraction Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1510_evalution_relation_extraction Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1510_evalution_relation_extraction.texttext-generation1K<n<10K1 likes114 downloads2y agoHugging Face30davanstrien /doab-metadata-extraction DOAB Open Access Books - Metadata Extraction Dataset Dataset Description This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs. The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/doab-metadata-extraction.imageimage-to-text1K<n<10K0 likes109 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.