datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Void-Witch-Astra-Vanta
Void Witch Astra Vanta
Source-derived release with authored context (schema 4)
448 rows: 93 unchanged conversation exchanges and 355 document chunks.
All 1,623 nonblank authored source lines appear exactly once as body text.
No passages are omitted. The row count changed from 788 because passages,
headings and lists are now grouped by their source relationships.
The seven original .txt files are archived byte-for-byte in sources/ under
their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.IIRCturkish-llm-dataset
Turkish Pretraining Corpus
Dataset Description
This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.
This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-llm-dataset.cybersec-master-dataset
Cybersecurity Master Instruction Dataset
Overview
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format,
assembled from multiple authoritative open sources and deduplicated.
At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity
LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger
broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.coding-master-dataset
Coding Master Dataset
Overview
A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated.
Records: 766,987
Format: JSONL / ShareGPT
License: Apache 2.0
Sources
CodeX-2M-Thinking (430,542 records)
python-code-dataset-500k (559,515 records)
StackPulse high-quality subset (20,205 records)
CodeFeedback-Filtered-Instruction (156,525 records)
secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.bhagavad-gita-verses-sanskrit-translations
Bhagavad Gita – Sanskrit, Transliteration & Multi-Commentary Dataset
A complete dataset of all 700 verses of the Bhagavad Gita, sourced directly from the open-source VedicScriptures API (MIT-licensed).This dataset includes:
📜 Original Sanskrit slokas
🔡 IAST transliteration
🌐 Multiple English & Hindi translations
🧠 Traditional commentaries from many teachers
🔢 Structured metadata (chapter, verse, IDs, authors)
This dataset is ideal for NLP, LLM fine-tuning, translation… See the full description on the dataset page: https://huggingface.co/datasets/Voider22/bhagavad-gita-verses-sanskrit-translations.gemini-3.1-opus-4.6-reasoning-merged Merged from
https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning
https://huggingface.co/datasets/reedmayhew/gemini-3.1-pro-2048-reasoning-1100x
https://huggingface.co/datasets/crownelius/Opus-4.6-Reasoning-3300x
GSM-ICshivaay-identity-datasetqrecctulu-3-sft-mixture-50kEduQGgemini-3.1-opus-4.6-reasoning-merged_v2Voidfeet
Voidfeet
Synthetically generated insanity.
voidful__smol-360m-ft-details
Dataset Card for Evaluation run of voidful/smol-360m-ft
Dataset automatically created during the evaluation run of model voidful/smol-360m-ft
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/voidful__smol-360m-ft-details.tulu-3-sft-forkdistil-ent4saturn-void-corpus
