datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VeraCruz_PT-BR
Dataset Summary
The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are:
Portugal (PT): Samples with content URLs indicating a clear Portuguese origin.
Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.VeRA
VeRA: Reasoning Benchmarks as Executable Specifications
NeurIPS 2026 - Evaluations & Datasets Track (Anonymous Submission)
What is VeRA?
Most reasoning benchmarks are static: the same problems are reused indefinitely, enabling memorization, format exploitation, and saturation. VeRA redefines a benchmark as an executable specification - a triple (template, generator, verifier) that can be sampled post-training to produce unlimited fresh instances with programmatically… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-NeurIPS26-VeRA/VeRA.Vera-Agentic-Coder
Vera Agentic Coder - GLM-5.2 REAP-50 Calibration Corpus
Public-safe calibration corpus for a steep 50% REAP expert-pruning pass on GLM-5.2. The mix is weighted to preserve agent harness/tool behavior and coding capability, while retaining baseline general knowledge, basic math/reasoning, infra, security, shell, and home-lab competence.
This is intended for router/expert activation profiling only. It is not a supervised training set or benchmark.
Stats… See the full description on the dataset page: https://huggingface.co/datasets/hornsan1/Vera-Agentic-Coder.tafsir-dataset
📘 Quran Tafsir Dataset
🧾 Description
This dataset contains Persian text derived from Quran tafsir (interpretation) lectures and explanations. The data is structured for training language models in instruction-following and religious text understanding tasks.
The dataset includes detailed explanations of Quranic verses, focusing on conceptual understanding, theological insights, and contextual analysis.
🎯 Use Cases
This dataset can be used for:
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/vera110/tafsir-dataset.VeraDATAlrggeneral LLM dataset for a multitude of tasks including reasoning, general purpoe and agents.
