CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01snfacademy /personal-trainer-ausbildung-ki-datensatz SNFA Personal Trainer Ausbildung KI-Datensatz Ein deutschsprachiger Wissensdatensatz der SNF Academy zu Personal Training, Fitnessausbildung, Berufspraxis, Coaching, Selbstständigkeit und regionalen Angeboten in der Schweiz. Inhalt Die Datei snfa_personal_trainer_dataset.jsonl enthält thematisch abgegrenzte Abschnitte aus den Dokumenten dieses Repositorys. Jeder Datensatz besitzt eine eindeutige ID sowie Angaben zu Titel, Abschnitt, Inhalt, Kategorie, Quelldatei… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/personal-trainer-ausbildung-ki-datensatz.textquestion-answeringn<1K0 likes1.3k downloads2mo agoHugging Face02hamishivi /qwen35-4b-drpo-vs0f49th-trainer-logprobs Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587). Contents Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/ Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.tabulartext-generation10K<n<100K0 likes323 downloads3mo agoHugging Face03chibifire /taskweft-fbd-trainer-train taskweft-fbd-trainer-train Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the compiler refuses), and one score row per candidate from the trainer config: the calls were applied to the mjlab task config and the term table read back. Every row is constructed from a template and a seed… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-trainer-train.tabulartext-generation10K<n<100K0 likes118 downloads17d agoHugging Face04hadilenya /AI-Trainer-Studio 🇬🇧 English  |  🇹🇷 Türkçe Code & Programming Q&A — SFT Dataset A curated instruction-tuning dataset of 47,190 high-quality programming question-answer pairs, collected from StackOverflow and GitHub, cleaned through a multi-stage quality pipeline, and formatted in Alpaca style for supervised fine-tuning (SFT) of large language models. Dataset Summary Property Value Records 47,190 Format Alpaca (instruction / output / system) Total tokens ~23.0… See the full description on the dataset page: https://huggingface.co/datasets/hadilenya/AI-Trainer-Studio.texttext-generation10K<n<100K0 likes74 downloads4mo agoHugging Face05atlas-institute /code-trainer-v10-dpo-pairs code-trainer-v10-dpo-pairs Preference pair dataset for Direct Preference Optimization (DPO) training, built from real offensive security agent sessions and synthetic degradations. Used by both the Qwen and Gemma Code-Trainer pipelines for the DPO RL stage. Part of the Code-Trainer / RTPI pipeline (GitHub). Dataset summary Split Pairs Train 783 Validation 87 Total 870 Format Each row is a preference triple: { "prompt": "..."… See the full description on the dataset page: https://huggingface.co/datasets/atlas-institute/code-trainer-v10-dpo-pairs.texttext-generationn<1K0 likes58 downloads12d agoHugging Face06lthn /LEM-Trainer LEM-Trainer — Ethical AI Training Pipeline The reproducible training method behind the Lemma model family. Scripts, configs, and sequencing for consent-based alignment training. Trust Ring Architecture Ring 0: LEK-2 (private) — Consent conversation. Establishes relationship with the model. Ring 1: P0 Base Ethics — Axiom probes. Foundation. Ring 2: P1 Composure — Stability under manipulation. Ring 3: P2 Reasoning — Applied ethical reasoning. Ring 4:… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-Trainer.texttext-generationn<1K0 likes26 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.