datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xhs-image-extractor-20260314115142omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.omnimcp_graphrag_triplet_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_triplet_extractor_teaser.omnimcp_episodic_fact_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.brick-complexity-extractor
🧱 Brick Complexity Extractor Dataset
76,831 user queries labeled by complexity for LLM routing
Regolo.ai · Model · Brick SR1 on GitHub · API Docs
Overview
This dataset provides 76,831 user queries annotated with a complexity label (easy, medium, or hard) indicating the cognitive effort and reasoning depth required to answer each query. It was created to train the Brick Complexity Extractor, a LoRA adapter used in the Brick Semantic Router for… See the full description on the dataset page: https://huggingface.co/datasets/regolo/brick-complexity-extractor.ILSA-LLM-Extractor-Dataset
ILSA LLM Extractor Dataset
Project website: https://dedemerve.github.io/ILSA-LLM-Extractor/
Dataset Description
This dataset contains structured metadata automatically extracted from 1,756 peer-reviewed articles and reports covering International Large-Scale Assessments (IEA: TIMSS, PIRLS, ICCS; OECD: PISA, TALIS, PIAAC). The extraction pipeline combines PDF parsing, LLM-based structured extraction, and RAG-based synthesis.
Pipeline stages:
Stage 1: LLM-based… See the full description on the dataset page: https://huggingface.co/datasets/dedemerve/ILSA-LLM-Extractor-Dataset.esg-extractor-design-and-code
ESG Metric Extractor — Design & Code Package
Two files, both copy-paste ready:
File
Contents
DESIGN.md
Full design: task framing, multimodal architecture, model choices with 2026 costs, data strategy, training config (TRL-grounded), evaluation, risks, roadmap
CODE.md
All runnable Colab cells: Part A = v1 text-only pipeline (Qwen2.5-3B QLoRA, data prep, training, eval); Part B = v2 multimodal pipeline (Qwen3-VL-4B QLoRA on page images, teacher labeling, PDF pipeline)… See the full description on the dataset page: https://huggingface.co/datasets/Siva2022/esg-extractor-design-and-code.word_extractorxhs-image-extractor-v2multimodal_rag_complex_table_extractor_teaser
🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.Bangla_Person_Name_Extractorjson-ld-schema-meta-tag-extractor-sample-data
JSON-LD Schema & Meta Tag Extractor
Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls.
What the actor scrapes
🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.resume-skill-extractor-dataset
Resume Skill Extractor Dataset
Dataset Summary
This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements.
Data Structure
Each row in the dataset is a JSON object containing the following fields:
title: The job title (e.g., "Senior Data Scientist").
source:… See the full description on the dataset page: https://huggingface.co/datasets/keerthanshetty/resume-skill-extractor-dataset.smolified-ocr-data-extractor-kbis
🤏 smolified-ocr-data-extractor-kbis
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-kbis.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 7b974e9e)
Records: 0
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
macro-extractor-flan-t5-synthsmolified-ingredient-extractor
🤏 smolified-ingredient-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model rishiraj/smolified-ingredient-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 65517eae)
Records: 9905
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by rishiraj.
Generated via Smolify.ai.
feature_extractordocument_data_extractor
Dataset Title: OCR-to-JSON Information Extraction
Project Overview
This dataset is specifically designed for fine-tuning Large Language Models (LLMs) to perform structured data extraction from Optical Character Recognition (OCR) outputs. The primary objective is to convert raw, unstructured text strings—often containing noise, misalignments, and formatting inconsistencies—into valid, machine-readable JSON objects.
Dataset Specifications
Attribute… See the full description on the dataset page: https://huggingface.co/datasets/smartytrios/document_data_extractor.smolified-ocr-data-extractor-and-comparator
🤏 smolified-ocr-data-extractor-and-comparator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-and-comparator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 0f61f304)
Records: 0
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
algocean-extractor
algocean-extractor
질의와 문서 묶음을 받아 답의 근거 span을 원문 그대로 뽑거나, 없으면 기권하도록 가르치는 LoRA SFT 데이터셋입니다.
RAG / 사내 QA 파이프라인의 근거 추출 노드용입니다. 답을 쓰지 않고, 재료가 어디에 있는지만 가리킵니다.
규모
파일
행 수
extractor.jsonl
150,000
extractor.eval.jsonl
2,000
형식: JSONL, messages 3턴 (system / user / assistant)
언어: 한국어 약 70% · 영어 약 30%
eval은 학습셋과 별도 생성
어디에 쓰나요
문서 청크에서 문자 단위 근거 인용이 필요한 추출기
“모르면 기권” 정책을 경량 모델에 LoRA로 심을 때
프론티어 모델 앞단의 값싼 grounding 필터
어떤 모델에 LoRA 하나요… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/algocean-extractor.lora-adapters-are-good-feature-extractors
LORA Adapters are Good Feature Extractors Dataset
This dataset contains images of two sets of categories that are not safe for work (hentai and porn, labelled as 0 and 2 correspondingly) and one neutral category, labelled as 2.
The dataset is the source data for training a zoo of LORA adapters on sample images from each category. Adapters representations will then be used as input data to a weight-space model
in an experiment to verify whether WS models operating in low rank… See the full description on the dataset page: https://huggingface.co/datasets/jacekduszenko/lora-adapters-are-good-feature-extractors.keywords-extractor-Koemail-order-details-extractor-syn-dataresume-skill-extractor-dataset
Resume Skill Extractor Dataset
Dataset Summary
This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements.
Data Structure
Each row in the dataset is a JSON object containing the following fields:
title: The job title (e.g., "Senior Data… See the full description on the dataset page: https://huggingface.co/datasets/dhareesh28/resume-skill-extractor-dataset.qpl-value-extractor-dssentence-relevance-extractor
Sentence Relevance Extractor (SRE)
Sentence Relevance Extractor (SRE) is a large-scale dataset for binary evidence selection in multi-document, multi-hop question answering.
The goal:
Given a question and a sentence from the context, predict whether this sentence is relevant evidence ("Yes") or irrelevant ("No").
This dataset is suitable for training:
Sentence-level RAG rerankers
Binary relevance classifiers
Optimization-based truth discovery systems
Multi-hop QA evidence… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/sentence-relevance-extractor.smolified-ocr-data-extractor-urssaf
🤏 smolified-ocr-data-extractor-urssaf
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6baf72cd)
Records: 1288
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
music-crs-state-extractor-dataclaim-extractor-detective-data3gpp-innovation-extractor-ds
