datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
instructtts-three-model-gemini-zh
InstructTTSEval 三模型 Gemini 评测数据
本目录整理了 InstructTTSEval 中文集上三个 TTS 模型的生成音频和 Gemini 一致性评测结果:Qwen3-TTS-12Hz-1.7B-VoiceDesign、Seed-Audio-1.0、VoxCPM2。
字段
records.jsonl 每行对应一个模型和一种控制格式(APS、DSD 或 RP):
id:InstructTTSEval 样本 ID
mode:控制格式
model、model_name:模型标识
text:合成文本
instruction:历史评测记录中的输入控制指令,按本行 APS/DSD/RP 格式保留;不等同于各模型 API 的完整请求封装
generated_audio:该模型生成音频的相对路径
reference_audio:原始参考音频的相对路径
gemini_consistent:Gemini judge 的一致性判断
inconsistency_reason:判断为不一致时的原因… See the full description on the dataset page: https://huggingface.co/datasets/zsy814/instructtts-three-model-gemini-zh.threejs-gamecode-instruct-v3-ultra
Three.js GameCode Instruct v3 Ultra
This is a large synthetic/original instruction dataset for training or testing LLM behavior around Three.js, browser game development, gameplay programming, debugging, optimization, architecture, and general coding.
Important note
This dataset is synthetic and programmatically generated from original templates. It is designed as a useful starting point for experiments, not as a fully hand-curated gold-standard benchmark.
No… See the full description on the dataset page: https://huggingface.co/datasets/agagasf123123/threejs-gamecode-instruct-v3-ultra.rr_three_tasks_v1
rr_three_tasks_v1
Three tasks on a Trossen AI solo arm, merged into one LeRobot v2.1 dataset.
task
episodes
frames
pick_specific_item_from_clutter
243
59088
pick_two_in_order
99
40478
open_pot_and_place
100
47288
meta/sources.jsonl maps every episode to its source dataset, episode and revision, with the
staging record (open_pot_and_place variant, pick_two second object, sheet row).
Held-out evaluation episodes
meta/eval_episodes_v1.json: 44… See the full description on the dataset page: https://huggingface.co/datasets/k1seul/rr_three_tasks_v1.threew
threew
threew adalah dataset pasangan instruksi–jawaban berbahasa Indonesia untuk eksperimen text generation dan instruction tuning. Dataset ini berisi contoh sintetis yang dikurasi secara programatis dan diarahkan agar jawaban bersifat jelas, aman, jujur tentang ketidakpastian, serta berguna untuk pembelajaran umum.
Struktur
Setiap baris JSONL memiliki kolom berikut:
Kolom
Tipe
Keterangan
id
string
Identitas unik contoh
instruction
string
Permintaan… See the full description on the dataset page: https://huggingface.co/datasets/ojiwzrd/threew.infrared_benchmarkfib
Dataset Card for FIB
Dataset Summary
The FIB benchmark consists of 3579 examples for evaluating the factual inconsistency of large language models. Each example consists of a document and a pair of summaries: a factually consistent one and a factually inconsistent one. It is based on documents and summaries from XSum and CNN/DM.
Since this dataset is intended to evaluate the factual inconsistency of large language models, there is only a test split.
Accuracies should be… See the full description on the dataset page: https://huggingface.co/datasets/r-three/fib.codellama-threejsinfrared-pretrain-500kpixeldit-xl-three-1p6m-checkpoints-20260922
Three full PixelDiT-XL checkpoints at optimizer step 1,600,000
These are distinct experimental branches, not three copies of one model.
Each file is a full training checkpoint (online weights, EMA, optimizer and
resumption state), not an inference-only weights export. For evaluation, use
the EMA weights and keep each branch label attached to its results.
Directory
Branch
h200_lr2e5_strict_rng_from1m
H200 strict-RNG continuation from the 1M LR2e-5 checkpoint… See the full description on the dataset page: https://huggingface.co/datasets/ziqiaow/pixeldit-xl-three-1p6m-checkpoints-20260922.threejs_30three_ch_bhagwatGeetaToolACE-masking-threeSFT_LLAMA_362_threeturninfrared-instruct-12ktrain_test_threeSFT_GPT_362_threeturnThreeBody-zhthree_keywordchina_three_kingdomsdata5sn96_2three-dbdata6korea_three_kingdomsthreesn96_3data3data8three_keyword_240923keyword_three_0926
