datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-training-dataset
LLM Fine-Tuning Dataset - 4,000,000+ logs, 32 languages
The dataset contains over 4 million+ logs written in 32 languages and is tailored for LLM training. It includes log and response pairs from 3 models, and is designed for language models and instruction fine-tuning to achieve improved performance in various NLP tasks - Get the data
Models used for text generation:
GPT-3.5
GPT-4
Uncensored GPT Version (is not included inthe sample)
Languages in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/llm-training-dataset.Gita-Train
Bhagavad-Gita-QA-Multilingual
Dataset Summary
Bhagavad-Gita-QA, is a carefully structured verse-aligned dataset that brings the timeless wisdom of the Bhagavad Gita into a modern question–answer framework.
This is the first open dataset that provides verse-level Q&A for the Gita with questions in Hindi and Gujarati along with English. This is not just a technical resource but also a cultural bridge, enabling new ways of studying, teaching, and exploring the Gita… See the full description on the dataset page: https://huggingface.co/datasets/rahul7star/Gita-Train.chess_puzzle_training_datasets_lt-2400
Chess puzzle training datasets: rating below 2400
This is a filtered derivative of
pavelslab-nyu/chess_puzzle_training_datasets.
Every retained row satisfies the exact condition:
Rating < 2400
Rating is the Lichess puzzle rating, not the Elo of either player in the
source game. The original column names, column order, directory layout, and CSV
schemas are preserved. As in the upstream repository, Hugging Face discovers
all three CSVs as one default configuration with one train… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/chess_puzzle_training_datasets_lt-2400.minicrit-training-12k
📝 Read the full blog post: MiniCrit: Adversarial AI Validation for Financial Decision-Making
MiniCrit Training Dataset: 12,132 Trading Rationale-Critique Pairs
A comprehensive dataset of realistic institutional trading rationales paired with adversarial critiques, designed to train AI models that can validate trading signals and reduce false positives.
Dataset Summary
Size: 12,132 unique rationale-critique pairsFormat: CSV (2 columns: rationale, rebuttal)License:… See the full description on the dataset page: https://huggingface.co/datasets/wmaousley/minicrit-training-12k.compliance_train
Compliance Database - Merged Profiles & Masters 📋
Dataset Description
This dataset is a comprehensive, synthetic collection of 200 merged records, combining detailed Company Profiles with their corresponding Compliance Master entries. Each row represents a specific compliance requirement for a particular company profile, meticulously linking all relevant attributes from both sources.
It serves as a foundational dataset for understanding the interplay between business… See the full description on the dataset page: https://huggingface.co/datasets/praveena0506/compliance_train.clinical-quad-protocol-deviation-cluster-staffing-load-training-gap-governance-pressure-v0.1Clarus Clinical Quad Coupling Protocol Deviation Cluster Staffing Load Training Gap Governance Pressure v0.1
What this dataset isThis dataset tests whether a model can detect clustered protocol deviations caused by four interacting nodes.
Quad coupling nodes
Deviation rate or severity cluster
Staffing or workload pressure
Training gap or outdated materials
Governance or compliance review pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-protocol-deviation-cluster-staffing-load-training-gap-governance-pressure-v0.1.BSG_CyLlama-computational-training
BSG CyLlama V8 - Computational Training Data
20,000 training samples for the computational theme, generated by DeepSeek.
Each sample contains source abstracts from a scientific cluster and target outputs (abstract summary, short summary, title, overview) for training the BSG CyLlama cluster description generator.
Format
TSV with columns: cluster_id, theme, source_abstracts, abstract_summary, short_summary, title, theme_token, short_title, overview
Related… See the full description on the dataset page: https://huggingface.co/datasets/jimnoneill/BSG_CyLlama-computational-training.leo-training-dataBSG_CyLlama-experimental-training
BSG CyLlama V8 - Experimental Training Data
20,000 training samples for the experimental theme, generated by DeepSeek.
Each sample contains source abstracts from a scientific cluster and target outputs (abstract summary, short summary, title, overview) for training the BSG CyLlama cluster description generator.
Format
TSV with columns: cluster_id, theme, source_abstracts, abstract_summary, short_summary, title, theme_token, short_title, overview
Related
Model:… See the full description on the dataset page: https://huggingface.co/datasets/jimnoneill/BSG_CyLlama-experimental-training.pre-train-testkurdish-train-smalltrain{
"use_cases": [
{
"use_case": "General Reporting",
"details": "N/A",
"process_flow": "N/A"
},
{
"use_case": "Voice CDR",
"details": "System shall handle errors and store CDRs in the database.",
"process_flow": [
"User uploads voice CDR files to the system.",
"System processes the files, validates, and ingests data into the database.",
"System logs any errors encountered during processing."
]
},
{… See the full description on the dataset page: https://huggingface.co/datasets/Aakanksha26/train.Naruto_training
