datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.wooden_window_factory_01_enriched_v2
Real industrial data, AI-ready for Physical AI
ORION WWF1 – Certified Sample Pack v2.0 (Enriched)
Version
Status
Sector
Pipeline
v2.0-Enriched
🟢 Level 3 Certified
Industrial-Manufacturing
Orion Unified V5.2
🌟 The Evolution: Beyond Anonymization
The ORION WWF1 v2.0 Enriched pack represents the professional evolution of our baseline industrial dataset. While previous versions focused on privacy-first anonymization, v2.0 transforms raw video… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v2.wooden_window_factory_01_enriched_v1.5
Real industrial data, AI-ready for Physical AI
ORION WWF1 – Enriched Sample Pack v1.5 (Physical AI Edition)
🌟 The Evolution: Beyond Anonymization
The ORION WWF1 v1.5 Enriched pack is the professional evolution of our baseline industrial dataset. While version 1.0 focused on privacy-first anonymization, v1.5 transforms raw video into actionable intelligence.
This pack includes 10 representative clips from a high-intensity wood-processing facility, now… See the full description on the dataset page: https://huggingface.co/datasets/Orion-The-Lab/wooden_window_factory_01_enriched_v1.5.bigbenchhard-enrichedMAGE-enriched
Info
This is a version of the MAGE benchmark enriched with linguistic features.
The linguistic features were extracted with elfen.
Citation
If you use this enriched version of MAGE, please cite
@inproceedings{doenmez-maurer-2025-ai,
title = "AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts",
author = "Dönmez, Esra and
Maurer, Maximilian and
Lapesa, Gabriella and
Falenska, Agnieszka",
year = {2025},
booktitle = "To… See the full description on the dataset page: https://huggingface.co/datasets/mmmaurer/MAGE-enriched.gelbooru-characters-enriched
Gelbooru Characters Enriched
This dataset is an enriched, fully-mapped version of Gelbooru character tags. It contains resolved franchise (copyright) associations and core appearance features (core tags) for 263,441 unique characters.
Dataset Details
The dataset maps the original character list to their corresponding copyrights (franchises) and general core attributes. It was constructed using a multi-stage hybrid extraction pipeline:
Regex Extraction: Extracting… See the full description on the dataset page: https://huggingface.co/datasets/cloud19/gelbooru-characters-enriched.EnrichedMeaningDataset
Dataset Card for Chinese Degree Expressions for Pragmatic Reasoning (CDE-Prag), an ongoing project about Enriched Meaning.
Dataset Summary
CDE-Prag is a theory-driven evaluation dataset designed to probe the pragmatic competence of Large Language Models (LLMs) and Vision-Language Models (VLMs). It focuses specifically on manner implicatures and ambiguity detection through the lens of Chinese degree expressions (e.g., Kai gao, which is ambiguous between "Kai is tall" and… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/EnrichedMeaningDataset.mmlu-pro-enrichedreview-aspects-enrichedSynthetic ahh dataset
For each aspect:
0 = negative,
1 = not mentioned,
2 = positive.
For overall sentiment:
0 = negative,
2 = positive
pois-enriched-usa
Enriched Dataset with our stats
enriched-generated-arguments
Info
This is a version of a generated arguments corpus enriched with linguistic features and argument quality dimensions.
The linguistic features were extracted with elfen.
The argument quality dimensions were extracte with these adapters.
Citation
If you use this enriched version of the generated arguments corpus, please cite
@inproceedings{doenmez-maurer-2025-ai,
title = "AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive… See the full description on the dataset page: https://huggingface.co/datasets/mmmaurer/enriched-generated-arguments.restaurant_enrichedBooks_enriched
