datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VeraDataMLRCVera-Agentic-Coder
Vera Agentic Coder - GLM-5.2 REAP-50 Calibration Corpus
Public-safe calibration corpus for a steep 50% REAP expert-pruning pass on GLM-5.2. The mix is weighted to preserve agent harness/tool behavior and coding capability, while retaining baseline general knowledge, basic math/reasoning, infra, security, shell, and home-lab competence.
This is intended for router/expert activation profiling only. It is not a supervised training set or benchmark.
Stats… See the full description on the dataset page: https://huggingface.co/datasets/hornsan1/Vera-Agentic-Coder.tafsir-dataset
📘 Quran Tafsir Dataset
🧾 Description
This dataset contains Persian text derived from Quran tafsir (interpretation) lectures and explanations. The data is structured for training language models in instruction-following and religious text understanding tasks.
The dataset includes detailed explanations of Quranic verses, focusing on conceptual understanding, theological insights, and contextual analysis.
🎯 Use Cases
This dataset can be used for:
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/vera110/tafsir-dataset.VeraDATAlrggeneral LLM dataset for a multitude of tasks including reasoning, general purpoe and agents.
VeraDataMLRGVera Data MLRG (Mega Large)
Under enclicas Open-model system which is either MIT, GPL, or under the enclica Open-model licence.
Used for text generation is SLMs or LLMs
used in enclica VERA or VERAAILabs aka EnclicaAILabs
sn120-screen-h3v8-vs-vera6Conscience_computationelmoonstone-datasets
