datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
veracier-industries
EDiTh — Enterprise Digital Twin Benchmark
What is this dataset?
EDiTh (Enterprise Digital Twin) is an open benchmark for evaluating
enterprise search and RAG systems on documents that actually look like
the ones you deal with every day: multilingual, scanned, cross-referenced,
and full of the edge cases that break demos.
At its core is Véracier Industries S.A., a fictional but rigorously
grounded €1.8 B French industrial group: 7 subsidiaries across 5 countries
(France… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/veracier-industries.VeRA
VeRA: Reasoning Benchmarks as Executable Specifications
NeurIPS 2026 - Evaluations & Datasets Track (Anonymous Submission)
What is VeRA?
Most reasoning benchmarks are static: the same problems are reused indefinitely, enabling memorization, format exploitation, and saturation. VeRA redefines a benchmark as an executable specification - a triple (template, generator, verifier) that can be sampled post-training to produce unlimited fresh instances with programmatically… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-NeurIPS26-VeRA/VeRA.toxic_uncensored_LGBTQ_csvtafsir-dataset
📘 Quran Tafsir Dataset
🧾 Description
This dataset contains Persian text derived from Quran tafsir (interpretation) lectures and explanations. The data is structured for training language models in instruction-following and religious text understanding tasks.
The dataset includes detailed explanations of Quranic verses, focusing on conceptual understanding, theological insights, and contextual analysis.
🎯 Use Cases
This dataset can be used for:
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/vera110/tafsir-dataset.
