datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.fundusnap-fundustalk-v1-chatsft-11k
📢 Domain & Email Migration Notice
From May 30th, 2026, Fundusnap will transition to new domains as fundusnap.com will not be renewed:
🌐 Website: fundusnap.faizath.com (formerly fundusnap.com)
⚙️ API: fundusnap-api.faizath.com (formerly api.fundusnap.com)
📧 Email: contact@fundusnap.faizath.com (formerly contact@fundusnap.com)
🛰️ CDN: fundusnap-cdn.faizath.com (formerly cdn.fundusnap.com)
📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/fundusnap/fundusnap-fundustalk-v1-chatsft-11k.nanochat-d24-sft-chat-eval-v1
nanochat d24 SFT Chat Eval Capture
Per-example outputs for the nanochat chat eval tasks across dense, nested, and MatFormer SFT models.
Each dataset config corresponds to one model. Each split corresponds to one eval task.
Important columns include input_prompt, rendered_prompt, model_response, correct,
task_logical_index, example_key, and forward-inference FLOP estimates split into
flops_prefill, flops_decode, and flops_total.
Dataset repo: atrost/nanochat-d24-sft-chat-eval-v1… See the full description on the dataset page: https://huggingface.co/datasets/atrost/nanochat-d24-sft-chat-eval-v1.LogicMind-Chat-Reasoning-SFT-300K
Nemotron-Post-Training-Dataset-v2-chat Dataset Card
Overview 📌
This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line).
Highlights
Scale: 296,168 samples
Category: chat (100%)
Generator: qwen-3-32b (100%)
Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.llama3_additional_rr40k_non_delete_sft_chat_formatNemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.llama3_regular_balanced_sft_chat_formatcrisp-sft-chatllama31_sft_non_delete_300k_chat_formatskymizer__Llama2-7b-sft-chat-custom-template-dpo-details
Dataset Card for Evaluation run of skymizer/Llama2-7b-sft-chat-custom-template-dpo
Dataset automatically created during the evaluation run of model skymizer/Llama2-7b-sft-chat-custom-template-dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/skymizer__Llama2-7b-sft-chat-custom-template-dpo-details.math_70b_correct_llama3_8b_filtered_sft_chatazaria-mitchell_Yi-6B-Chat_sft_to_lieRM-Bench-chat-mistral-7b-sft-beta-comprehensivebigvul-distilled-sft-chatRM-Bench-chat-mistral-7b-sft-beta-entropyRM-Bench-chat-mistral-7b-sft-beta-yesnoRM-Bench-chat-Llama-3.1-Tulu-3-8B-SFT-yesnoshaer-eval-instruction-yehia-base-sft-chat-template-with-hitsshaer-eval-instruction-yehia-base-sft-chat-template
shaer-eval-instruction-yehia-base-sft-chat-template
Clean paper-facing evaluation dataset for the Shaer benchmark.
Rows: 3481
Split: test
Schema
id
base_meter
form
requested_bayts
requested_num_lines
description
enhanced_description
reference_completion
generated_text
meter
count_adherence
description_adherence
meaning
fluency
coherence
poeticness
Notes
meter is the row-level metrical conformity score used in the paper.
This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI/shaer-eval-instruction-yehia-base-sft-chat-template.
