datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.omnimcp_graphrag_triplet_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_triplet_extractor_teaser.omnimcp_episodic_fact_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.ILSA-LLM-Extractor-Dataset
ILSA LLM Extractor Dataset
Project website: https://dedemerve.github.io/ILSA-LLM-Extractor/
Dataset Description
This dataset contains structured metadata automatically extracted from 1,756 peer-reviewed articles and reports covering International Large-Scale Assessments (IEA: TIMSS, PIRLS, ICCS; OECD: PISA, TALIS, PIAAC). The extraction pipeline combines PDF parsing, LLM-based structured extraction, and RAG-based synthesis.
Pipeline stages:
Stage 1: LLM-based… See the full description on the dataset page: https://huggingface.co/datasets/dedemerve/ILSA-LLM-Extractor-Dataset.multimodal_rag_complex_table_extractor_teaser
🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.resume-skill-extractor-dataset
Resume Skill Extractor Dataset
Dataset Summary
This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements.
Data Structure
Each row in the dataset is a JSON object containing the following fields:
title: The job title (e.g., "Senior Data Scientist").
source:… See the full description on the dataset page: https://huggingface.co/datasets/keerthanshetty/resume-skill-extractor-dataset.json-ld-schema-meta-tag-extractor-sample-data
JSON-LD Schema & Meta Tag Extractor
Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls.
What the actor scrapes
🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.smolified-ocr-data-extractor-kbis
🤏 smolified-ocr-data-extractor-kbis
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-kbis.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 7b974e9e)
Records: 0
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
macro-extractor-flan-t5-synthsmolified-ocr-data-extractor-and-comparator
🤏 smolified-ocr-data-extractor-and-comparator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-and-comparator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 0f61f304)
Records: 0
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
smolified-ingredient-extractor
🤏 smolified-ingredient-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model rishiraj/smolified-ingredient-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 65517eae)
Records: 9905
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by rishiraj.
Generated via Smolify.ai.
lora-adapters-are-good-feature-extractors
LORA Adapters are Good Feature Extractors Dataset
This dataset contains images of two sets of categories that are not safe for work (hentai and porn, labelled as 0 and 2 correspondingly) and one neutral category, labelled as 2.
The dataset is the source data for training a zoo of LORA adapters on sample images from each category. Adapters representations will then be used as input data to a weight-space model
in an experiment to verify whether WS models operating in low rank… See the full description on the dataset page: https://huggingface.co/datasets/jacekduszenko/lora-adapters-are-good-feature-extractors.keywords-extractor-Kosmolified-ocr-data-extractor-urssaf
🤏 smolified-ocr-data-extractor-urssaf
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6baf72cd)
Records: 1288
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
resume-skill-extractor-dataset
Resume Skill Extractor Dataset
Dataset Summary
This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements.
Data Structure
Each row in the dataset is a JSON object containing the following fields:
title: The job title (e.g., "Senior Data… See the full description on the dataset page: https://huggingface.co/datasets/dhareesh28/resume-skill-extractor-dataset.qpl-value-extractor-ds3gpp-innovation-extractor-dstest-updated-extractor-v2
test-updated-extractor-v2
LLM-based math span extraction with canonicalization
Dataset Info
Rows: 1
Columns: 26
Columns
Column
Type
Description
question
Value('string')
No description provided
metadata
Value('string')
No description provided
task_source
Value('string')
No description provided
formatted_prompt
List({'content': Value('string'), 'role': Value('string')})
No description provided
responses_by_sample
List(List(Value('string')))… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/test-updated-extractor-v2.test-updated-extractor
test-updated-extractor
LLM-based math span extraction with canonicalization
Dataset Info
Rows: 1
Columns: 26
Columns
Column
Type
Description
question
Value('string')
No description provided
metadata
Value('string')
No description provided
task_source
Value('string')
No description provided
formatted_prompt
List({'content': Value('string'), 'role': Value('string')})
No description provided
responses_by_sample
List(List(Value('string')))
No… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/test-updated-extractor.test-updated-extractor-v3
test-updated-extractor-v3
LLM-based math span extraction with canonicalization
Dataset Info
Rows: 1
Columns: 26
Columns
Column
Type
Description
question
Value('string')
No description provided
metadata
Value('string')
No description provided
task_source
Value('string')
No description provided
formatted_prompt
List({'content': Value('string'), 'role': Value('string')})
No description provided
responses_by_sample
List(List(Value('string')))… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/test-updated-extractor-v3.smolified-extractor
🤏 smolified-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 13d088d3)
Records: 15669
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
smolified-ocr-data-extractor-urssaf-2
🤏 smolified-ocr-data-extractor-urssaf-2
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf-2.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 1f7ab49c)
Records: 550
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
buzz_sources_100_extractor-00000-of-00001smolified-ingredient-extractor
🤏 smolified-ingredient-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Shubhankar444/smolified-ingredient-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 22fdd899)
Records: 250
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Shubhankar444.
Generated via Smolify.ai.
model_card_extractor_qwq
Dataset card for model_card_extractor_qwq
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"modelId": "digiplay/XtReMixAnimeMaster_v1",
"author": "digiplay",
"last_modified": "2024-03-16 00:22:41+00:00",
"downloads": 232,
"likes": 2,
"library_name": "diffusers",
"tags": [
"diffusers",
"safetensors",
"stable-diffusion",
"stable-diffusion-diffusers",
"text-to-image"… See the full description on the dataset page: https://huggingface.co/datasets/aravind-selvam/model_card_extractor_qwq.smolified-ingredient-extractor
🤏 smolified-ingredient-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-ingredient-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 1f92fa68)
Records: 600
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
table_extractor_vietnamesesmolified-calorie-nutrition-extractor
🤏 smolified-calorie-nutrition-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Shubhankar444/smolified-calorie-nutrition-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 54c6c8fc)
Records: 1080
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Shubhankar444.
Generated via Smolify.ai.
