CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Logistic12 /xhs-image-extractor-20260314115142image1K<n<10K0 likes193 downloads6mo agoHugging Face02emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes70 downloads5d agoHugging Face03emgena /omnimcp_graphrag_triplet_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_triplet_extractor_teaser.texttext-generationn<1K0 likes60 downloads5d agoHugging Face04emgena /omnimcp_episodic_fact_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.texttext-generationn<1K0 likes52 downloads5d agoHugging Face05regolo /brick-complexity-extractor 🧱 Brick Complexity Extractor Dataset 76,831 user queries labeled by complexity for LLM routing Regolo.ai · Model · Brick SR1 on GitHub · API Docs Overview This dataset provides 76,831 user queries annotated with a complexity label (easy, medium, or hard) indicating the cognitive effort and reasoning depth required to answer each query. It was created to train the Brick Complexity Extractor, a LoRA adapter used in the Brick Semantic Router for… See the full description on the dataset page: https://huggingface.co/datasets/regolo/brick-complexity-extractor.text-classification10K<n<100K2 likes51 downloads5mo agoHugging Face06dedemerve /ILSA-LLM-Extractor-Dataset ILSA LLM Extractor Dataset Project website: https://dedemerve.github.io/ILSA-LLM-Extractor/ Dataset Description This dataset contains structured metadata automatically extracted from 1,756 peer-reviewed articles and reports covering International Large-Scale Assessments (IEA: TIMSS, PIRLS, ICCS; OECD: PISA, TALIS, PIAAC). The extraction pipeline combines PDF parsing, LLM-based structured extraction, and RAG-based synthesis. Pipeline stages: Stage 1: LLM-based… See the full description on the dataset page: https://huggingface.co/datasets/dedemerve/ILSA-LLM-Extractor-Dataset.tabularfeature-extraction10K<n<100K0 likes49 downloads3mo agoHugging Face07Siva2022 /esg-extractor-design-and-code ESG Metric Extractor — Design & Code Package Two files, both copy-paste ready: File Contents DESIGN.md Full design: task framing, multimodal architecture, model choices with 2026 costs, data strategy, training config (TRL-grounded), evaluation, risks, roadmap CODE.md All runnable Colab cells: Part A = v1 text-only pipeline (Qwen2.5-3B QLoRA, data prep, training, eval); Part B = v2 multimodal pipeline (Qwen3-VL-4B QLoRA on page images, teacher labeling, PDF pipeline)… See the full description on the dataset page: https://huggingface.co/datasets/Siva2022/esg-extractor-design-and-code.0 likes44 downloads11d agoHugging Face08MindCastSogang /word_extractorimage100K<n<1M0 likes40 downloads1mo agoHugging Face09Logistic12 /xhs-image-extractor-v2image1K<n<10K0 likes39 downloads6mo agoHugging Face10emgena /multimodal_rag_complex_table_extractor_teaser 🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 📦 What is Inside the Full Production Package: 500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.texttext-generationn<1K0 likes38 downloads4d agoHugging Face11MBMMurad /Bangla_Person_Name_Extractortext1K<n<10K0 likes35 downloads3y agoHugging Face12logiover /json-ld-schema-meta-tag-extractor-sample-data JSON-LD Schema & Meta Tag Extractor Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls. What the actor scrapes 🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.textn<1K0 likes33 downloads4mo agoHugging Face13keerthanshetty /resume-skill-extractor-dataset Resume Skill Extractor Dataset Dataset Summary This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements. Data Structure Each row in the dataset is a JSON object containing the following fields: title: The job title (e.g., "Senior Data Scientist"). source:… See the full description on the dataset page: https://huggingface.co/datasets/keerthanshetty/resume-skill-extractor-dataset.text1K<n<10K0 likes31 downloads5mo agoHugging Face14titou4ng /smolified-ocr-data-extractor-kbis 🤏 smolified-ocr-data-extractor-kbis Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-kbis. 📦 Asset Details Origin: Smolify Foundry (Job ID: 7b974e9e) Records: 0 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. text-generation1K<n<10K0 likes26 downloads7mo agoHugging Face15JohnGorri /macro-extractor-flan-t5-synthtext10K<n<100K0 likes24 downloads3mo agoHugging Face16rishiraj /smolified-ingredient-extractor 🤏 smolified-ingredient-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model rishiraj/smolified-ingredient-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 65517eae) Records: 9905 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by rishiraj. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes20 downloads6mo agoHugging Face17Vishal24 /feature_extractortext1K<n<10K0 likes19 downloads2y agoHugging Face18smartytrios /document_data_extractor Dataset Title: OCR-to-JSON Information Extraction Project Overview This dataset is specifically designed for fine-tuning Large Language Models (LLMs) to perform structured data extraction from Optical Character Recognition (OCR) outputs. The primary objective is to convert raw, unstructured text strings—often containing noise, misalignments, and formatting inconsistencies—into valid, machine-readable JSON objects. Dataset Specifications Attribute… See the full description on the dataset page: https://huggingface.co/datasets/smartytrios/document_data_extractor.0 likes17 downloads8mo agoHugging Face19titou4ng /smolified-ocr-data-extractor-and-comparator 🤏 smolified-ocr-data-extractor-and-comparator Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-and-comparator. 📦 Asset Details Origin: Smolify Foundry (Job ID: 0f61f304) Records: 0 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. text-generation1K<n<10K0 likes17 downloads7mo agoHugging Face20Algocean /algocean-extractor algocean-extractor 질의와 문서 묶음을 받아 답의 근거 span을 원문 그대로 뽑거나, 없으면 기권하도록 가르치는 LoRA SFT 데이터셋입니다. RAG / 사내 QA 파이프라인의 근거 추출 노드용입니다. 답을 쓰지 않고, 재료가 어디에 있는지만 가리킵니다. 규모 파일 행 수 extractor.jsonl 150,000 extractor.eval.jsonl 2,000 형식: JSONL, messages 3턴 (system / user / assistant) 언어: 한국어 약 70% · 영어 약 30% eval은 학습셋과 별도 생성 어디에 쓰나요 문서 청크에서 문자 단위 근거 인용이 필요한 추출기 “모르면 기권” 정책을 경량 모델에 LoRA로 심을 때 프론티어 모델 앞단의 값싼 grounding 필터 어떤 모델에 LoRA 하나요… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/algocean-extractor.texttext-generation1K<n<10K0 likes16 downloads2mo agoHugging Face21jacekduszenko /lora-adapters-are-good-feature-extractors LORA Adapters are Good Feature Extractors Dataset This dataset contains images of two sets of categories that are not safe for work (hentai and porn, labelled as 0 and 2 correspondingly) and one neutral category, labelled as 2. The dataset is the source data for training a zoo of LORA adapters on sample images from each category. Adapters representations will then be used as input data to a weight-space model in an experiment to verify whether WS models operating in low rank… See the full description on the dataset page: https://huggingface.co/datasets/jacekduszenko/lora-adapters-are-good-feature-extractors.tabularn<1K1 likes15 downloads2y agoHugging Face22jaeyong2 /keywords-extractor-Kotext10K<n<100K0 likes15 downloads1y agoHugging Face23RaghaMounam /email-order-details-extractor-syn-datatextn<1K0 likes15 downloads10mo agoHugging Face24dhareesh28 /resume-skill-extractor-dataset Resume Skill Extractor Dataset Dataset Summary This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements. Data Structure Each row in the dataset is a JSON object containing the following fields: title: The job title (e.g., "Senior Data… See the full description on the dataset page: https://huggingface.co/datasets/dhareesh28/resume-skill-extractor-dataset.text1K<n<10K0 likes15 downloads1mo agoHugging Face25d4nieldev /qpl-value-extractor-dstext10K<n<100K0 likes12 downloads11mo agoHugging Face26MinaGabriel /sentence-relevance-extractor Sentence Relevance Extractor (SRE) Sentence Relevance Extractor (SRE) is a large-scale dataset for binary evidence selection in multi-document, multi-hop question answering. The goal: Given a question and a sentence from the context, predict whether this sentence is relevant evidence ("Yes") or irrelevant ("No"). This dataset is suitable for training: Sentence-level RAG rerankers Binary relevance classifiers Optimization-based truth discovery systems Multi-hop QA evidence… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/sentence-relevance-extractor.tabular1M<n<10M0 likes12 downloads10mo agoHugging Face27titou4ng /smolified-ocr-data-extractor-urssaf 🤏 smolified-ocr-data-extractor-urssaf Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf. 📦 Asset Details Origin: Smolify Foundry (Job ID: 6baf72cd) Records: 1288 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes12 downloads7mo agoHugging Face28shuaih777 /music-crs-state-extractor-datatext100K<n<1M0 likes12 downloads4mo agoHugging Face29Luimas /claim-extractor-detective-data0 likes12 downloads4mo agoHugging Face30OrganizedProgrammers /3gpp-innovation-extractor-dstabularn<1K0 likes10 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.