datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.tr-full-dataset
TR-Full_dataset
This is a merged speech dataset containing 41427 audio segments from 88 source datasets.
Dataset Information
Total Segments: 41427
Speakers: 222
Languages: tr
Emotions: neutral, angry, sad, happy
Original Datasets: 88
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.prompt_consistency_training_full_data
🚀 Load Dataset
from datasets import load_dataset
dataset = load_dataset("shuyuej/prompt_consistency_training_full_data")
dataset = dataset["train"]
print(dataset)
KUJIRA_DATASETS_FULL
Overview
This dataset is an English-language dataset created specifically for Reasoning models within the KUJIRA_v2 series.
The primary objective of the dataset is to enhance model performance while ensuring safety and removing censorship typical of Chinese-origin models.
In creating this dataset, references were made to the dataset recipes of r1-1776 and MAI-DS-R1.
Additionally, this dataset leverages existing Q&A data originally designed for intuitive models by employing the… See the full description on the dataset page: https://huggingface.co/datasets/Team-Kitsune/KUJIRA_DATASETS_FULL.Dataset_Robot_IA_full_v8alpaca-cleaned-serbian-full
Serbian Alpaca Cleaned Dataset
Original Repository: https://github.com/gururise/AlpacaDataCleaned
Original HF Repository: https://huggingface.co/datasets/yahma/alpaca-cleaned
Dataset Description
This is a serbian cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the… See the full description on the dataset page: https://huggingface.co/datasets/datatab/alpaca-cleaned-serbian-full.Mdist_Chem_LST_150k_ft_data_from_full_360kfull_data_15_negMdist_Chem_LST_Qwen3_14B_full_ft_datavulnerability_training_data_fullMdist_Chem_LST_250k_ft_data_from_full_360kMdist_Chem_FFT_Qwen3_14B_full_ft_dataReasoner-dataset-FULL-rolesMdist_Chem_LST_100k_ft_data_from_full_360kMdist_Chem_LST_300k_ft_data_from_full_360kArabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({
train: Dataset({
features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'],
num_rows: 12166802
})
})
ams_data_full_2000-2020Aerospace Mechanism Symposia PDF documents parsed by page. All symposia documents from the year 2000-2022 are included. No splitting was used.
Original documents here: https://github.com/dan-s-mueller/aerospace_chatbot/tree/main/data/AMS
sanskrit_full_datasetkyrgyz-asr-full-data
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
godlikehhd__alpaca_data_full_2-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_full_2
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_full_2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_full_2-details.godlikehhd__alpaca_data_full_3B-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_full_3B
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_full_3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_full_3B-details.distilled_dataset_full_cot_v1uzbek-asr-full-data
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
Francisco-Angulo-de-Lafuente-Full-Dataset
Francisco Angulo de Lafuente Full Training Dataset
Este dataset contiene la recopilación completa de la obra, investigación, código y biografía de Francisco Angulo de Lafuente.
Contenido
Biografía Detallada: Información sobre su trayectoria en biotecnología, ingeniería informática y literatura.
Proyectos Principales: Documentación técnica de P2PCLAW, EUHNN, CAJAL y otros.
Papers Científicos: Resúmenes y textos completos de sus investigaciones en computación neuromórfica… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/Francisco-Angulo-de-Lafuente-Full-Dataset.godlikehhd__qwen_full_data_alpaca-details
Dataset Card for Evaluation run of godlikehhd/qwen_full_data_alpaca
Dataset automatically created during the evaluation run of model godlikehhd/qwen_full_data_alpaca
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__qwen_full_data_alpaca-details.scta_full_dataset.jsonltajik-asr-full-data
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
Dataset_Robot_IA_full_v7kazakh-asr-full-data
Saidzoda Lab — Gated Research Dataset
Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz).
Dataset contents, provenance, and statistics are not publicly disclosed.
Access is granted manually on request.
