datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenVid-60k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 60k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
from datasets import load_dataset, disable_caching, DownloadMode
from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-60k-split.OpenVid-10k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 10k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
from datasets import load_dataset, disable_caching, DownloadMode
from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-10k-split.azerbaijan-court-data
Azerbaijan Court System Dataset
The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations.
Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale.
Quick Start
Load with Hugging Face datasets
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.OpenVid-1k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 1k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
Loading the data:
from datasets import load_dataset, disable_caching, DownloadMode
from… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-1k-split.clustered-reference-voices
Clustered Reference Voices (EMOLIA 3K)
3,000 enhanced reference voice MP3s — one high-quality representative sample per speaker cluster, selected and scored by a multi-expert neural quality model.
Overview
Property
Value
Total clips
3,000
Total duration
11.3 hours
Mean duration
13.5 s (range: 3.5 – 29.9 s)
Format
MP3, 192 kbps, 48 kHz
Language
English
Naming
{cluster_id}.mp3 (0 – 2999)
Source
The source data is… See the full description on the dataset page: https://huggingface.co/datasets/laion/clustered-reference-voices.
