datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
circleci-test-resultstest-results
SZL test-results — anatomy alive-harness sink
Public, DSSE-signed results sink for the SZL anatomy alive-harness. This
dataset was retired earlier in 2026 and stood back up on 2026-07-21 as the
harness's fail-closed publishing target — restoring the public proof loop
behind every "harness verified" claim in the estate.
What a run is
anatomy_alive_v6.py (in this repo) drives live assertions across the whole
substrate — organ liveness, live formula-gate executions… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/test-results.test-resultscausalgraph-results-testWSV_AVIS_test_results_v1.7benchio-results-expert-test
benchio
WSV_AVIS_test_results_v1.6lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test-private
Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test
Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test-private.Garments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed item… See the full description on the dataset page: https://huggingface.co/datasets/ArtmeScienceLab/Garments2Look-Test-Set-Results.medical-test-resultsexp10_promptC_results_testmoderation-test-resultsqwen35-arabic-test-resultsunderwater_test_enhance_results
Underwater Test Enhancement Results
This private dataset repository stores the packaged outputs of the underwater image enhancement reproduction run dated 2026-07-09.
Archive
File: results/underwater_test_enhance_20260709/underwater_test_enhance_20260709_results.zip
Size: 2,196,977,738 bytes
SHA256: 2a68c68a75aa853d8eb7008db3f9ccce647fd349639a9fa4f65d4b812f03041a
Source dataset: BitStrawber/underwater_test
The archive contains outputs/, manifests/… See the full description on the dataset page: https://huggingface.co/datasets/Mortallll/underwater_test_enhance_results.test-bench-blind-resultsRIAID_test_10k_results_2024WSV_AVIS_test_results_v1.3stage2_csqa_eval_test_resultsWSV_AVIS_test_results_v1.5sdft-tat-dqa-test-resultsexp10_promptB_results_testtest_resultstag-validation-results-test
タグ検証結果
このデータセットは説明生成におけるタグ検証の結果を含んでいます。
概要
処理したサンプル数: 100
説明終了タグ成功率: 92.00%
検証対象タグ: <|end_of_explanation|>
トークン数統計
最小トークン数: 226
最大トークン数: 30070
平均トークン数: 1773.4
トークン数分布
0-100トークン: 0件 (0.0%)
101-500トークン: 44件 (44.0%)
501-1000トークン: 37件 (37.0%)
1001-2000トークン: 13件 (13.0%)
2001-5000トークン: 1件 (1.0%)
5001+トークン: 5件 (5.0%)
データセット構造
id: 各サンプルの一意識別子
user_content: モデルにユーザーメッセージとして送信した内容(質問のみ)
assistant_content: モデルにアシスタントメッセージとして送信した内容(タグ付き解答 +… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/tag-validation-results-test.model2-test-resultsblack_box_test_resultsblack_box_test_updated_results_t0_token_calculatedleaderboard-test-resultsRIAID_balanced_test_results2_with_baselines_qwentemplate-test-results-exploration-round2-v2template-test-results-first-person-v2
