datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
personal-trainer-ausbildung-ki-datensatz
SNFA Personal Trainer Ausbildung KI-Datensatz
Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz.
Inhalt
Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.qwen35-4b-drpo-vs0f49th-trainer-logprobs
Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th
This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587).
Contents
Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th
Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/
Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl
Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.AI-Trainer-Studio
🇬🇧 English
|
🇹🇷 Türkçe
Code & Programming Q&A — SFT Dataset
A curated instruction-tuning dataset of 47,190 high-quality programming question-answer pairs, collected from StackOverflow and GitHub, cleaned through a multi-stage quality pipeline, and formatted in Alpaca style for supervised fine-tuning (SFT) of large language models.
Dataset Summary
Property
Value
Records
47,190
Format
Alpaca (instruction / output / system)
Total tokens
~23.0… See the full description on the dataset page: https://huggingface.co/datasets/hadilenya/AI-Trainer-Studio.mcq_safety
MCQ Safety
Merged safety multiple-choice dataset built from SafetyBench test-en, SALAD
Bench MCQ data, and WildGuardMix harm-category data.
Splits
Deterministic random split with seed 42:
split
rows
train
15993
valid
889
test
888
Format
Each JSONL row contains:
prompt: problem plus options formatted as A) ..., B) ...
answer: single boxed option label, e.g. \boxed{C}
source: source dataset name
metadata: JSON-encoded source and normalization… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-vibe-trainers/mcq_safety.Dataset-setfit-TrainerDataset-setfit-Trainer-80recordstrainer_dataset
Dataset Card for Dataset Name
The dataset contains phylosophical books in question answer format
Curated by: Me
Language(s) (NLP): English
License: apache-2.0
Uses
Can be used for fine-tunining reasoning models
Dataset Structure
source, question, answer, extra
Extra:
Also checkout our new agent theStech/conscious_Ai-2.
weightlifting-trainer-logneuro_qa_SFT_TrainerTrainermy-code-trainerMH_Master_Trainer
