CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes5.8k downloads7mo agoHugging Face02artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes530 downloads7mo agoHugging Face03finetrainers /OpenVid-10k-split Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M. This is a 10k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered. from datasets import load_dataset, disable_caching, DownloadMode from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-10k-split.tabulartext-to-video1K<n<10K2 likes370 downloads1y agoHugging Face04lonesamurai /emilia_clean_10k EMILIA Clean 10k A filtered subset of the amphion/Emilia-Dataset (English split), designed for single-speaker TTS training. Dataset Statistics Total clips: 10,000 Speakers: 200 (single-speaker English) Train / Val split: 8,000 / 2,000 Duration per clip: 3–10 seconds Sample rate: 24 kHz (mono) Language: English (EN) Filtering Pipeline Candidate selection — Filtered EMILIA EN clips for duration (3–10s) and DNSMOS quality (≥3.2). Selected top 400 speakers with… See the full description on the dataset page: https://huggingface.co/datasets/lonesamurai/emilia_clean_10k.audio10K<n<100K1 likes42 downloads5mo agoHugging Face05VictorSanh /obelisc_embedded_10k_v0text10K<n<100K0 likes39 downloads3y agoHugging Face06henryyzhaoo /real-estate-10k-rawtext10K<n<100K0 likes37 downloads6mo agoHugging Face07qinglinhou /sokoban-10k-vjepa2-tokenized-shardstext10K<n<100K0 likes34 downloads5mo agoHugging Face08FerasMad /genimage-midjourney-10kimage10K<n<100K0 likes30 downloads4mo agoHugging Face09fhudson96 /TAPVid360-10kimage100K<n<1M2 likes24 downloads11mo agoHugging Face10yilingwang /TSPO_10ktext1K<n<10K0 likes18 downloads10mo agoHugging Face11tbuckley /synthetic-derm-10kimage10K<n<100K0 likes16 downloads2y agoHugging Face12Embodied1 /vlm3r_sample_10ktext1K<n<10K0 likes16 downloads8mo agoHugging Face13wangzn2001 /openwebtext-10ktextn<1K0 likes14 downloads3y agoHugging Face14sapienzanlp /Concept-10k-imgsimage10K<n<100K2 likes7 downloads11mo agoHugging Face15fhudson96 /TAP360-10k-zippedimage100K<n<1M0 likes6 downloads1y agoHugging Face16deepsynthbody /SinGAN-Seg-10k-polypsimage10K<n<100K0 likes3 downloads2y agoHugging Face17KarimGhon /Concept-10k-imgsimage10K<n<100K0 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.