datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turnbench-dev-no-backchannel
TurnBench Dev - Backchannels Removed
A derivative of mundo-ai/turn-benchmark-dev
with every majority-annotated backchannel removed from the audio: 1853 backchannels
across 38 conversations, 2077.0 seconds in total, cut out of the
speaker's own channel and replaced by background noise taken from elsewhere in that same channel.
Everything else is the original recording, sample for sample. Same conversations, same duration,
same timeline, same annotator tracks, same speech -- only… See the full description on the dataset page: https://huggingface.co/datasets/JSALT2026-Conv-AI-Simulator/turnbench-dev-no-backchannel.lemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-Shuffledlemonilia_LimaRP-Simple-CustomShareGPT-flatguard-splitlemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-qwq-all-aphrodite
lemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-qwq-all-aphrodite
You should mask everything except the last turn. The only part that matters to teach the model is the last turn, as you are teaching it to always output thinking, no matter what the user feeds it.
It's setup to be trained like R1:
dedup_ablation_sim_threshold_0D-EVAL__standard_eval_v3__simple_test__exp_runner_2-eval_0
D-EVAL__standard_eval_v3__simple_test__exp_runner_2-eval_0
This evaluation dataset was created as part of the simple_test__exp_runner_2 experiment using the SkillFactory experiment management system.
Experiment Tracking
🔗 View complete experiment details: Experiment Tracker Dataset
Evaluation Details
{"model": "Qwen/Qwen2.5-1.5b-Instruct", "tasks": ["countdown_2arg", "countdown_3arg"], "annotators": ["greedy"], "splits": ["test"], "dataset_url":… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-EVAL__standard_eval_v3__simple_test__exp_runner_2-eval_0.lemonilia_LimaRP-Simple-CustomShareGPT-Shuffledlemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-qwq-all
lemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-qwq-all
You should mask everything except the last turn. The only part that matters to teach the model is the last turn, as you are teaching it to always output thinking, no matter what the user feeds it.
Generation Script
This is what I used to generate the dataset, so that it's setup to be trained like R1:
import requests
import json
import time
import pandas as pd
from datasets import load_dataset
from tqdm import… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/lemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-qwq-all.D-EVAL__standard_eval_v3__simple_test__exp_runner_3-eval_sft
D-EVAL__standard_eval_v3__simple_test__exp_runner_3-eval_sft
This evaluation dataset was created as part of the simple_test__exp_runner_3 experiment using the SkillFactory experiment management system.
Experiment Tracking
🔗 View complete experiment details: Experiment Tracker Dataset
Evaluation Details
{"model": "TAUR-dev/M-simple_test__exp_runner_3-sft", "tasks": ["countdown_2arg", "countdown_3arg"], "annotators": ["greedy"], "splits": ["test"]… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-EVAL__standard_eval_v3__simple_test__exp_runner_3-eval_sft.LibriTTS-dev-clean-16khz-mono-loudnorm-100-random-samples-2024-04-18-17-34-39-similaritiesdedup_ablation_sim_threshold_10dedup_ablation_sim_threshold_20dedup_ablation_sim_threshold_40dedup_ablation_sim_threshold_nonelemonilia_LimaRP-Simple-CustomShareGPT-flatguard-split-qwqD-SFT_C-back_to_og_mix__simple_retries__sbon-sft-dataD-EVAL__standard_eval_v3__back_to_og_mix__simple_mix__rl_eval-eval_rl
D-EVAL__standard_eval_v3__back_to_og_mix__simple_mix__rl_eval-eval_rl
This evaluation dataset was created as part of the back_to_og_mix__simple_mix__rl_eval experiment using the SkillFactory experiment management system.
Experiment Tracking
🔗 View complete experiment details: Experiment Tracker Dataset
Evaluation Details
{"model": "TAUR-dev/SIE-back_to_og_mix__simple_retries__sbon-rl", "tasks": ["countdown_2arg", "countdown_3arg", "countdown_4arg"… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-EVAL__standard_eval_v3__back_to_og_mix__simple_mix__rl_eval-eval_rl.dataset__acronym_generation__simple__6_wordsD-EVAL__simple_eval__cd3arg-sft1ep_grpo_1e6lr_mix_ss_pse_vote_ansrev-sftD-simple_test-sft-dataD-back_to_og_mix__simple_retries__sbon-sft-datadataset__acronym_generation__simple__4_wordsdataset__acronym_generation__simple__5_wordsdataset__acronym_generation__simple__7_wordslemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-qwq-all-aphrodite-filtD-EVAL__standard_eval_v3__back_to_og_mix__simple_retries__sbon-eval_sft
D-EVAL__standard_eval_v3__back_to_og_mix__simple_retries__sbon-eval_sft
This evaluation dataset was created as part of the back_to_og_mix__simple_retries__sbon experiment using the SkillFactory experiment management system.
Experiment Tracking
🔗 View complete experiment details: Experiment Tracker Dataset
Evaluation Details
{"model": "TAUR-dev/M-back_to_og_mix__simple_retries__sbon-sft", "tasks": ["countdown_2arg", "countdown_3arg", "countdown_4arg"… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/D-EVAL__standard_eval_v3__back_to_og_mix__simple_retries__sbon-eval_sft.lemonilia_LimaRP-Only-NonSus-Simple-CustomShareGPT-qwq-all-aphrodite-classifiedlemonilia_LimaRP-Simple-CustomShareGPT-flatguard
