datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlam_hermes_validatedsynthvision-validated-qwen-by-kimi
synthvision-validated-qwen-by-kimi
Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate)
Records: 55,359
About
Cross-validated subset from the SynthVision pipeline. Kimi K2.5 reviewed all 59,476 Qwen 3.5 annotations and confirmed 55,359 as consistent with the source images (93.1% pass rate).
Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-validated-qwen-by-kimi.synthvision-validated-kimi-by-qwen
synthvision-validated-kimi-by-qwen
Kimi K2.5 annotations validated by Qwen 3.5 (93.0% pass rate)
Records: 55,382
About
Cross-validated subset from the SynthVision pipeline. Qwen 3.5 reviewed all 59,539 Kimi K2.5 annotations and confirmed 55,382 as consistent with the source images (93.0% pass rate).
Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-validated-kimi-by-qwen.CommonVoice17_validated_faSWE-bench_validated_12_18_style-3__fs-oracleSWE-bench_validated_12_18_bug_report_bug_report__fs-oracleSWE-bench_validated_12_18_unsolved_style-3__fs-oracleSWE-bench_validated_12_22_style-3__fs-oracleML-1M-Syntax-Validated-Python-Code
ML-1M Syntax-Validated Python Code
Dataset Summary
ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code.
The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.synthvision-validated-qwen-by-kimi
synthvision-validated-qwen-by-kimi
Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate)
Records: 55,359
About
Cross-validated subset from the SynthVision pipeline. Kimi K2.5 reviewed all 59,476 Qwen 3.5 annotations and confirmed 55,359 as consistent with the source images (93.1% pass rate).
Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/synthvision-validated-qwen-by-kimi.auto_validatedrunyankole-crowd-validated-paths
Dataset Card for "runyankole-crowd-validated-paths"
More Information needed
syn_validatedauto_validated2CV25.0-pt-validated-16kCV25.0-pt-validated-16k-filteredvalidated_12_18std_lang_wiki_qa_v4_validated_w_reasoning_effort_llmjp4
Standard Japanese Wiki QA v4 with Reasoning Effort — LLM-jp 4
概要
Imabari Wiki QA v4 with Reasoning Effort — LLM-jp 4 の thinking(思考文)と answer(最終回答)の言葉遣いを標準語(です・ます調)へ変換した派生データセットです。日本語QAの教師ありファインチューニング(SFT)や、reasoning effort に応じた説明文・回答の生成に利用できます。
質問に新たに回答したり、推論内容を追加したりすることは目的としていません。元の意味・結論・根拠を保持し、今治弁などの方言表現を標準語へ置き換えるよう指示しています。質問は変換対象に含めていません。
LLM-jp 4 のチャットテンプレート用に、assistant の content に最終回答、thinking に思考文を格納しています。「LLM-jp 4」は対象のデータ形式を示します。元の思考文生成モデルと標準語化の設定モデルは… See the full description on the dataset page: https://huggingface.co/datasets/ikedachin/std_lang_wiki_qa_v4_validated_w_reasoning_effort_llmjp4.bitflags_validatedtokio_validatedvalidated_12_18_unsolvedateso-crowd-validated-paths
Dataset Card for "ateso-crowd-validated-paths"
More Information needed
asterinas_validatedregex_validatedarrow_validatedhyper_validatedrand_validatedcrossbeam_validatedproc-macro2_validatedserde_validated
