datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HR-VILAGE-3K3M
HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression
This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.HRVideoBench
HRVideoBench
This repo contains the test data for HRVideoBench, which is released under the paper "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation". VISTA is a video spatiotemporal augmentation method that generates long-duration and high-resolution video instruction-following data to enhance the video understanding capabilities of video LMMs.
🌐 Homepage | 📖 arXiv | 💻 GitHub | 🤗 VISTA-400K | 🤗 Models | 🤗 HRVideoBench… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/HRVideoBench.HRVQAHR-VILAGE-3K3M
HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression
This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a benchmark… See the full description on the dataset page: https://huggingface.co/datasets/maywovel/HR-VILAGE-3K3M.gsr-hrv-stress-detection-dataset
GSR + HRV Stress Detection — Cleaned Volunteer Dataset
Cleaned, anonymized physiological recordings from 124 volunteers used to train the GSR+HRV stress detection model. Collected as part of an applied research project at Almutlaq United Company.
Full data-cleaning pipeline and training code: github.com/Alansi775/GSR-COSSINUS_VOLUNTEERS_DATA_FOR_ML
Protocol
Each volunteer completed an identical 4-phase session, recorded at 1 Hz:
Stage
Duration
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/malansi/gsr-hrv-stress-detection-dataset.genesis-hr-bench-rebuttal-full19-hrv2-eval-20260812
Genesis HR Bench — preliminary HR v2 train19 rebuttal evaluation
This is an immutable, sanitized snapshot of the currently available artifacts from:
hf19_eval5_act080000_full95_fast_20260812T0236Z
It contains per-episode videos, state traces, result JSON, validation records, task/model manifests, policy-server logs, and Slurm logs for five policies across the 19 HR v2 tasks used for this rebuttal evaluation. It also includes both focused seed-7 skillet diagnostic pilots.… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/genesis-hr-bench-rebuttal-full19-hrv2-eval-20260812.HRVLAHRVQA-2kA 2k subset of the validation split of the HRVQA dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation.
For more information and attribution please refer to the original dataset: https://hrvqa.nl/
HR-VITON
Dataset Card for "HR-VITON"
More Information needed
hrvatski-dataset
Hrvatski dataset
Širok hrvatski korpus za jezičnu prilagodbu i fino ugađanje malih jezičnih
modela, osobito Gemma 3 1B, Gemma 3 4B te kompatibilnih Gemma 4 modela.
Skup nije samo zbirka kratkih uputa. Sastoji se od dva komplementarna dijela:
cpt: 22,5 milijuna riječi književnog, enciklopedijskog i autentičnog
govornog hrvatskog za continued pretraining
sft: 47.588 razgovora za praćenje uputa, prirodne odgovore, dulji tekst,
književni nastavak i razgovorne replike
Za najbolji… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset.hrvatski-dataset-v2
Hrvatski Dataset v2
Comprehensive Croatian (hrvatski) language dataset with 455 examples, optimized for
chatbot training and article generation. Every example is a question/answer (or
instruction/output) pair written in natural Croatian.
This repo is a single-source-of-truth catalog: all format files are generated from one
canonical file (source/hrvatski_dataset_v2.jsonl), so every format is guaranteed to
contain the exact same 455 examples. If a format is ever out of sync, that… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset-v2.HRVQAHR-VSGeo
📝 This dataset is associated with our TGRS submission: "HR-VSGeo: HR-VSGeo: A High-Resolution Visible-SAR Geo-Referenced Dataset
and Evaluation Benchmark for Complex Geological Scenes"
HR-VSGeo Series Dataset
Dataset Summary
The HR-VSGeo Series dataset contains high-resolution Synthetic Aperture Radar (SAR) imagery, suitable for tasks such as object detection, semantic segmentation, and image classification. The data is collected from Google Earth and the Gaofen-3… See the full description on the dataset page: https://huggingface.co/datasets/RoyiLau/HR-VSGeo.HR-VVT
3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
Min Wei, Chaohui Yu, Jingkai Zhou, and Fan Wang. 2025.
3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models.
In Proceedings of the 33rd ACM International Conference on Multimedia (MM ’25),
October 27–31, 2025, Dublin, Ireland. ACM, New York, NY, USA, 10 pages.
https://doi.org/10.1145/3746027.3754754
Installation
git clone https://github.com/2y7c3/3DV-TON.git
cd 3DV-TON
pip… See the full description on the dataset page: https://huggingface.co/datasets/2y7c3/HR-VVT.HRVubuntu-vmHRVVShrv-qa-datasethrv-qna-dataset-v3hrv2hrv-qna-pipelinehr-viton-finetunedbabylm-hrv
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: hrv
Script: Latn
Tier: 1M
Byte Premium Factor: 0.989673
Size (MB): 5.39
Expected Size (MB): 5.37
Number of Documents: 466
Total Tokens: 915,054
Tokenizer: separate by whitespace
Tokens Per Category
child-directed-speech: 469,078 tokens
padding-opensubtitles: 445,976 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-hrv.hr_v2HrvluNQO
