datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
musicAnime-2026anonymous-phystc-2026
PhysTC-v1 Dataset
PhysTC-v1 is a physics-enhanced tropical cyclone forecasting dataset for track and intensity prediction. It provides storm-centered slow-dynamics features, 3D synoptic fields, environmental features, and 1D best-track records for evaluating machine learning models under cross-scale atmospheric dynamics.
This repository is anonymized for double-blind review. Author and institution information will be added in the final public release if the paper is accepted.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-phystc-2026/anonymous-phystc-2026.twoframe-eval-artifacts-20260505
TwoFrame Eval Artifacts 2026-05-05
Generated image artifacts for TwoFrame image-editing evaluation. The large image payload is stored as tar archives under archives/ to avoid uploading tens of thousands of loose PNG files.
Layout
archives/single_ref.tar: all complete single-reference outputs.
archives/multiref_part*.tar: complete K=2/K=3 multi-reference runs, sharded by run name.
archives/metadata.tar: README, index, manifests, and metrics as an archive.… See the full description on the dataset page: https://huggingface.co/datasets/wyhhey/twoframe-eval-artifacts-20260505.multilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.YouTubeVideoMetadata-2026-04-1Metadata from the Scrape Exchange for:
40k YouTube channels, including counts for subscribers, views and videos and merch, courses, posts, playlists. JSONSchema for the YouTube channels is on Github
9m YouTube videos, including views, likes, formats, thumbnail URLs, captions. JSONSchema for these YouTube videos is on Github
This is just the metadata; no actual videos, images, etc. are included. This dataset consists of a single tar file that contains Brotli-compressed JSON files. The… See the full description on the dataset page: https://huggingface.co/datasets/Boinko/YouTubeVideoMetadata-2026-04-1.IALP-2026-data
IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets
Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC
Adaptation." This repository holds the fixed target-query, validation, and
evaluation sets used across all experiments. Each part is a self-contained
.tar.gz.
All audio is 16 kHz mono. Each split ships with:
audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16)
wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.UAVM_2026_testvalkeir-public-derived-mimic3-matched-2p7tb-pass-2026-04-08-shard-de2ebe49
mimic3-matched-2p7tb-pass-2026-04-08 shard 09 of 03
This dataset is published from Project ValkEir's governed internal ML platform.
Publication ID: 8aaf9743b5064d5e
Role: source_dataset
Category: public-derived
License ID: odbl-1.0
Sanitization state: approved_public_release
Source classification: public_safe
See manifest/publication_manifest.json for the machine-readable package manifest.
robotwin_stack_blocks_two_clean50_2step_20260425
RobotWin Stack Blocks Two Clean50 2-Step Cache
This dataset repo contains a compressed latent-cache artifact for Video2VLA:
Archive: robotwin_stack_blocks_two_clean50_2step_20260425.tar.gz
Extracted directory: data/cache/robotwin_stack_blocks_two_clean50_2step_20260425
File count: 3757 .pt files
Sample latent shape: (1, 32, 4, 15, 20)
Dtype: torch.bfloat16
Representation: latent
This cache corresponds to the RobotWin stack_blocks_two clean50 setup using the 2-step denoising… See the full description on the dataset page: https://huggingface.co/datasets/RoMALab/robotwin_stack_blocks_two_clean50_2step_20260425.valkeir-public-derived-mimic3-matched-2p7tb-pass-2026-04-08-shard-a7157b9b
mimic3-matched-2p7tb-pass-2026-04-08 shard 02 of 03
This dataset is published from Project ValkEir's governed internal ML platform.
Publication ID: 28717c28a776d01e
Role: source_dataset
Category: public-derived
License ID: odbl-1.0
Sanitization state: approved_public_release
Source classification: public_safe
See manifest/publication_manifest.json for the machine-readable package manifest.
2026_2_5_smalldatavalkeir-public-derived-mimic3-matched-2p7tb-pass-2026-04-08-shard-1802c50f
mimic3-matched-2p7tb-pass-2026-04-08 shard 11 of 03
This dataset is published from Project ValkEir's governed internal ML platform.
Publication ID: bfb751396c75db05
Role: source_dataset
Category: public-derived
License ID: odbl-1.0
Sanitization state: approved_public_release
Source classification: public_safe
See manifest/publication_manifest.json for the machine-readable package manifest.
media-abstract-2026-03-0320260125score-imagesmedia-backup-2026-03-032026.04.082026.04.152026.04.202026.06.09reg2026-metric-a-data-feats2026.06.102026.06.122026_02_08_bigdata
PreciseCam Dataset - 2026_02_08_bigdata
数据集概述
本数据集包含全景图像的处理结果,合并了两个独立处理的数据集。
统计信息
总条目数: 5,212 个数据对
不同的组: 1,535 个
Group ID 范围: 1 ~ 1,535
图片总数: 约 114,664 张(去重后)
数据来源
数据集 1: ori_output_full (918 个组,3,134 条数据)
数据集 2: ori_output_full_gpu1 (617 个组,2,078 条数据)
文件说明
images_simplified.tar.gz (14GB)
压缩包包含:
images/ - 所有图片文件
每张图片以 group_xxx_ 作为前缀
例如:group_0001_083_yaw02_pair_left.jpg
AB_format_simplified.json - AB 格式数据
每个条目包含 4… See the full description on the dataset page: https://huggingface.co/datasets/Samp1eTree/2026_02_08_bigdata.2026020920260210202602122026021320260214
