datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.italic-extkd-pool
italic-extkd-pool
The stage-3 training data of idealab-cs2/zagreus-0.4B-italic-extkd: 57,563 Italian multiple-choice questions from public, in-distribution datasets with teacher soft labels. One soft-KD stage from the stage-2 checkpoint on the agreement-filtered subset (28,561 items where the teacher agrees with the gold answer) reaches 0.4921 / 0.4929 / 0.4932 on the full ITALIC 10K (official harness, 5-shot fast, temperature 0, three independent runs). Full lineage: 0.2802… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/italic-extkd-pool.italic-softkd-pool
italic-softkd-pool
The exact training data of idealab-cs2/zagreus-0.4B-italic-softkd: 21,606 Italian multiple-choice questions with committee soft labels. One soft-KD training run from mii-llm/zagreus-0.4B-ita on the train split reaches 0.4787 on the full ITALIC 10K (official harness, 5-shot fast, temperature 0), from a 0.2802 base.
train is the full pool; the other three splits partition it by provenance:
split
rows
contents
train
21,606
the full training file (union… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/italic-softkd-pool.deliberative-quality-72b
Reddit Deliberative Quality Labels (Qwen2.5-72B-Instruct)
This dataset has not been validated against human annotations. It is provided for illustrative and educational purposes only — specifically, as a training resource for building classifiers that rate the deliberative quality of Reddit comments. Scores should not be treated as ground-truth measures of discourse quality.
205,851 Reddit comment-parent pairs from r/worldnews and r/geopolitics (January 1 – February 3, 2026), each… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/deliberative-quality-72b.h-novel-alpacaConvert data to alpaca format based on https://huggingface.co/datasets/ystemsrx/Erotic_Literature_Collection and porn novel from other source.
Max length of the content is 7000
cs224w_final_proj_fxmodified_erotic_literature_collection-alpacaData converted to JSON from https://huggingface.co/datasets/li-long/modified_erotic_literature_collection
cs224n-dpo-3cs_2024_08_24cs224n-dpo-2cs2-coach-sftcs224n-dpo-1native_script_codemixedfull_native_scriptromanized_casualIndicCMix_WAarxiv-cs2021-corpusarxiv-cs2021-embeddings-bge-m3
