datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.qwen35-4b-drpo-vs0f49th-trainer-logprobs
Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th
This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587).
Contents
Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th
Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/
Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl
Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.fluency-trainertaskweft-fbd-trainer-train
taskweft-fbd-trainer-train
Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an
EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the
reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the
compiler refuses), and one score row per candidate from the trainer config: the calls were applied to the mjlab task config and the term table read back. Every row is
constructed from a template and a seed… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-trainer-train.fedgraph_citeseer_2trainer_1hop_iid_beta_1_trainer_id_0lulynx-trainer-offline-runtime-packetThis repository stores the pre-packaged offline runtime environment and pip dependency wheels for the lulynx trainer plugin ecosystem. It is designed to host environment freezing assets to ensure reproducible execution.
fedgraph_ogbn-products_10trainer_1hop_iid_beta_10000.0_trainer_id_2AI-Trainer-Studio
🇬🇧 English
|
🇹🇷 Türkçe
Code & Programming Q&A — SFT Dataset
A curated instruction-tuning dataset of 47,190 high-quality programming question-answer pairs, collected from StackOverflow and GitHub, cleaned through a multi-stage quality pipeline, and formatted in Alpaca style for supervised fine-tuning (SFT) of large language models.
Dataset Summary
Property
Value
Records
47,190
Format
Alpaca (instruction / output / system)
Total tokens
~23.0… See the full description on the dataset page: https://huggingface.co/datasets/hadilenya/AI-Trainer-Studio.fedgraph_ogbn-products_10trainer_1hop_iid_beta_10000.0_trainer_id_1minimax_music3_qlora_trainerfedgraph_ogbn-products_5trainer_1hop_iid_beta_1.0_trainer_id_1Transcription-Cleanup-Trainer
Text Cleanup Fine-Tuning Dataset
A curated dataset for training speech-to-text cleanup models to achieve optimal transcript refinement.
Dataset Description
This dataset contains paired examples of raw speech-to-text transcriptions and manually-cleaned versions, designed for fine-tuning models to clean up transcripts to a specific quality level ("Goldilocks" cleanup - not too much, not too little).
Dataset Structure
dataset/
├── data/
│ ├── audio/… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Transcription-Cleanup-Trainer.fedgraph_ogbn-arxiv_10trainer_1hop_iid_beta_100.0_trainer_id_6code-trainer-v9-mixed
code-trainer-v9-mixed
40,401-row mixed training dataset for supervised fine-tuning (SFT) in the
Code-Trainer / RTPI pipeline.
Used for both Qwen 14B (V9 SFT) and Gemma 26B (aggressive-full1 SFT) training.
Composition
Slice
Source
Rows (train)
Purpose
A -- Code generation
cmndcntrlcyber/code-trainer-offsec-dataset (8K subsample)
7,074
Preserve code-gen quality
B -- Tool calling
glaiveai/glaive-function-calling-v2 (19K cap)
~15,125
High-density tool… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v9-mixed.X-trainerx-trainer源文件
fedgraph_cora_1trainer_1hop_iid_beta_1.0_trainer_id_0code-trainer-v10-dpo-pairs
code-trainer-v10-dpo-pairs
Preference pair dataset for Direct Preference Optimization (DPO) training,
built from real offensive security agent sessions and synthetic degradations.
Used by both the Qwen and Gemma Code-Trainer pipelines for the DPO RL stage.
Part of the Code-Trainer / RTPI pipeline
(GitHub).
Dataset summary
Split
Pairs
Train
783
Validation
87
Total
870
Format
Each row is a preference triple:
{
"prompt": "..."… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v10-dpo-pairs.fedgraph_ogbn-papers100M_195trainer_0hop_iid_beta_100.0_trainer_id_112fedgraph_citeseer_2trainer_1hop_iid_beta_1_trainer_id_1fedgraph_ogbn-papers100M_195trainer_0hop_iid_beta_10000.0_trainer_id_61fedgraph_ogbn-papers100M_195trainer_0hop_iid_beta_100.0_trainer_id_127Saamayik
Sanskrit-English Parallel Translation Dataset (Saamayik)
Summary
The Saamayik Sanskrit-English Parallel Corpus is a contemporary prose-focused translation dataset containing around 53,000 parallel sentences in Sanskrit and English (48,326 in this main dataset, with an additional 4,047 Mann Ki Baat sentences). Sāmayik, meaning "sayings of the contemporary world" in Sanskrit, specifically addresses the gap in existing Sanskrit corpora which predominantly feature classical… See the full description on the dataset page: https://huggingface.co/datasets/Trainer2026/Saamayik.trainer_idolmastercinderellagirls
Dataset of trainer (THE iDOLM@STER: Cinderella Girls)
This is the dataset of trainer (THE iDOLM@STER: Cinderella Girls), containing 87 images and their tags.
The core tags of this character are black_hair, long_hair, hair_ornament, hairclip, breasts, black_eyes, brown_eyes, ponytail, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/trainer_idolmastercinderellagirls.Talker-T2AV-Data_trainer
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/prakhar-adaf/Talker-T2AV-Data_trainer.code-trainer-v7-mixedfedgraph_ogbn-papers100M_195trainer_0hop_iid_beta_10000.0_trainer_id_133fedgraph_ogbn-papers100M_195trainer_0hop_iid_beta_100.0_trainer_id_116code-trainer-offsec-datasetcode-trainer-v10-grpo-prompts
code-trainer-v10-grpo-prompts
500 curated prompts for GRPO (Group Relative Policy Optimization) training in the
Code-Trainer / RTPI pipeline.
Used by both Qwen and Gemma RL stages (Phase 4c) to train tool-call formatting
via a rule-based reward function.
Schema
Column
Type
Description
prompt
string
The user instruction/question
source
string
Origin: v10_eval, v9_training, or synthetic
expected_tool
string
Primary tool the prompt should invoke… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v10-grpo-prompts.fedgraph_ogbn-arxiv_10trainer_1hop_iid_beta_1.0_trainer_id_4
