datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.OpenVid-10k-split
Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M.
This is a 10k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered.
from datasets import load_dataset, disable_caching, DownloadMode
from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-10k-split.emilia_clean_10k
EMILIA Clean 10k
A filtered subset of the amphion/Emilia-Dataset (English split), designed for single-speaker TTS training.
Dataset Statistics
Total clips: 10,000
Speakers: 200 (single-speaker English)
Train / Val split: 8,000 / 2,000
Duration per clip: 3–10 seconds
Sample rate: 24 kHz (mono)
Language: English (EN)
Filtering Pipeline
Candidate selection — Filtered EMILIA EN clips for duration (3–10s) and DNSMOS quality (≥3.2). Selected top 400 speakers with… See the full description on the dataset page: https://huggingface.co/datasets/lonesamurai/emilia_clean_10k.obelisc_embedded_10k_v0real-estate-10k-rawsokoban-10k-vjepa2-tokenized-shardsgenimage-midjourney-10kTAPVid360-10kTSPO_10ksynthetic-derm-10kvlm3r_sample_10kopenwebtext-10kConcept-10k-imgsTAP360-10k-zippedSinGAN-Seg-10k-polypsConcept-10k-imgs
