CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ksolovev /FineNews30 likes551k downloads6mo agoHugging Face02ksolovev /FineNewsTestSampletext10M<n<100M0 likes14k downloads7mo agoHugging Face03KShivendu /dbpedia-entities-openai-1M1M OpenAI Embeddings -- 1536 dimensions Created: June 2023. Text used for Embedding: title (string) + text (string) Embedding Model: text-embedding-ada-002 First used for the pgvector vs VectorDB (Qdrant) benchmark: https://nirantk.com/writing/pgvector-vs-qdrant/ Citation @dataset{dbpedia-entities-openai-1M, doi = {10.57967/hf/6768}, url = {https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M}, author = {{Kumar Shivendu} and {Nirant Kasliwal}}, title =… See the full description on the dataset page: https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M.textfeature-extraction1M<n<10M26 likes3.8k downloads11mo agoHugging Face04ks46 /urls URLs 74,918,894,107 deduplicated, validated URLs, sorted by SURT key and split into 2,334 range shards. As plain text the URLs are 5.8 TiB, averaging 84 characters each. Sorted by SURT key and delta-encoded they fit in 665.6 GiB — 9.54 bytes per URL, a 8.85× reduction. That is the whole point of the ordering: SURT puts URLs from the same site next to each other, DELTA_LENGTH_BYTE_ARRAY then stores only where each row differs from the one above it, and zstd compresses what is… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls.texttext-generation10B<n<100B1 likes3k downloads1mo agoHugging Face05ks48 /url-atlas URL Atlas 257,548,097,528 URLs from 105 web corpora, each kept as its own separately-loadable config, plus the raw source dumps two of them were extracted from. 4.35 TiB across 52,244 files. This is the input side of a URL-compression corpus: every source reduced to its URL column and nothing else. It is deliberately not deduplicated or merged — sources are kept intact and overlapping so you can measure what each one contributes, pick the subset you want, and dedup on your own… See the full description on the dataset page: https://huggingface.co/datasets/ks48/url-atlas.texttext-retrieval100B<n<1T0 likes2.9k downloads1mo agoHugging Face06ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes2.6k downloads15d agoHugging Face07ksshumab /hf_ckpt0 likes2.3k downloads1y agoHugging Face08kshitijd /platonic-embeddingstimeseries1M<n<10M0 likes2.2k downloads5mo agoHugging Face09Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2k downloads4y agoHugging Face10ksokolovic /dec-pdfsdocumentn<1K0 likes1.9k downloads1y agoHugging Face11ksbai123 /Chime4text100K<n<1M1 likes1.8k downloads3y agoHugging Face12KSE-RESEARCH-Group /USL-Suspilne USL-Suspilne A small-scale Ukrainian Sign Language dataset for text-to-pose research, built from publicly broadcast news clips on Suspilne Mовлення (Ukrainian public broadcaster). Each clip pairs a Ukrainian sentence with the corresponding interpreter's signing, provided as a video clip and as MediaPipe pose sequences. Layout usl-suspilne/ ├── README.md ├── train.csv # 80% — model training ├── dev.csv # 10% — validation /… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/USL-Suspilne.0 likes1.7k downloads4mo agoHugging Face13ksaml /Stanford_dogs Context The Stanford Dogs dataset contains images of 120 breeds of dogs from around the world. This dataset has been built using images and annotation from ImageNet for the task of fine-grained image categorization. It was originally collected for fine-grain image categorization, a challenging problem as certain dog breeds have near identical features or differ in colour and age. I have used only images, so this does not contain any labels . Content Number of images:… See the full description on the dataset page: https://huggingface.co/datasets/ksaml/Stanford_dogs.image0 likes1.5k downloads4y agoHugging Face14JetBrains /KStack Dataset Summary KStack is the largest collection of permissively licensed Kotlin code. Comparison with The Stack v2 In the table below one can find the comparsion between the Kotlin part of The Stack v2 and KStack: Files Repositories Lines Tokens Kotlin in The Stack v2 2M 109,457 162M 1.7B Kstack 4M 168,902 292M 3.1B Dataset Creation Collection procedure We collected repositories from GitHub with the main language being… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/KStack.tabulartext-generation1M<n<10M14 likes1.3k downloads1y agoHugging Face15ksterx /hle-no-img-prompt-completion-formatimage1K<n<10K0 likes1.2k downloads1y agoHugging Face16kshitijd /pu-regress-results1M<n<10M0 likes1k downloads5mo agoHugging Face17DragonLine /ksponspeechaudio100K<n<1M2 likes940 downloads3y agoHugging Face18kshitijd /platonic-all-experimentstabularn<1K0 likes915 downloads1mo agoHugging Face19Bingsu /KSS_Dataset Description of the original author KSS Dataset: Korean Single speaker Speech Dataset KSS Dataset is designed for the Korean text-to-speech task. It consists of audio files recorded by a professional female voice actoress and their aligned text extracted from my books. As a copyright holder, by courtesy of the publishers, I release this dataset to the public. To my best knowledge, this is the first publicly available speech dataset for Korean. File Format Each… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/KSS_Dataset.audiotext-to-speech10K<n<100K20 likes869 downloads4y agoHugging Face20kshitijrajsharma /xview2-xbd xView2 / xBD (mirror) Mirror of the xView2 / xBD building damage assessment dataset (parquet), used to train kshitijrajsharma/dinov3-damage-assessment. Source Mirrored from EVER-Z/torchange_xView2. Attribution xBD / xView2 dataset from the xView2 Building Damage Assessment Challenge (Gupta et al., 2019). Imagery from the Maxar Open Data Program. See https://xview2.org/. License CC BY-NC-SA 4.0 (non-commercial, share-alike), following… See the full description on the dataset page: https://huggingface.co/datasets/kshitijrajsharma/xview2-xbd.geospatialimage-segmentation0 likes817 downloads3mo agoHugging Face21ksuinbragimova /AOPS_Full_Verified_sfrtext100K<n<1M0 likes810 downloads1y agoHugging Face22kszucs /opendaltext100M<n<1B0 likes802 downloads16d agoHugging Face23SebastianNi /KS-dataset-processed0 likes786 downloads6mo agoHugging Face24ksikka /test Fly Anipose — Lightning Pose Multiview Dataset 6-camera pose estimation dataset for Drosophila leg keypoints, packaged for use with Lightning Pose. Dataset Description Head-fixed flies run on a spherical treadmill while 6 synchronized cameras capture locomotion at 300 Hz. Each frame is labeled with 30 keypoints — 5 joint segments (A–E) on each of 6 legs (left legs L1–L3, right legs R1–R3). Labels are filtered Anipose predictions, not hand-labeled frames. They were… See the full description on the dataset page: https://huggingface.co/datasets/ksikka/test.imagekeypoint-detection1K<n<10K0 likes778 downloads6mo agoHugging Face25kshitijd /mmu-norm-legacy-north Legacy Survey DR9 North image cutouts — L1 (release v1) This L1 repository contains 2,191,927 objects matched across the release, in 726 shards (about 700 GB). Each object has 152×152-pixel cutouts at 0.262″ per pixel in three bands. Band names: DR9 North imaging uses BASS g/r and MzLS z. This repository uses bass-g, bass-r, and mzls-z; the original des-g/des-r/des-z tokens are kept in native_band_tokens. Schema image struct: band (3), flux (3×152×152… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-legacy-north.tabular1M<n<10M0 likes746 downloads1mo agoHugging Face26kshitijd /mmu-norm-jwst-wht JWST DAWN weight-map cutouts — L1 supplement (release v1) This L1 supplement provides per-pixel full weight maps (wht_full) for objects matched across the release in the six MMU JWST fields: CEERS, GDN, GDS, NGDEEP, PRIMER-COSMOS, and PRIMER-UDS. The maps come from DAWN JWST Archive v7 Grizli mosaics and are distributed in 93 chunks across eight mosaics. They add pixel-level uncertainty information to the science cutouts. Schema (one row per object-filter)… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-jwst-wht.text1M<n<10M0 likes723 downloads1mo agoHugging Face27jp1924 /KsponSpeechgatedaudioautomatic-speech-recognition100K<n<1M12 likes722 downloads9mo agoHugging Face28DongyunZou /minimax-h3-qkv-ksnr-202609170 likes718 downloads6d agoHugging Face29ksasse /government-ai-detection Government AI Text Detection — Results Completed AI-detection output from the government-ai pipeline, which measures the prevalence of AI-generated/edited text across four kinds of US government media, 2000–2026: source what bills Congressional bill text (as-introduced versions), from govinfo speeches Floor speeches + Extensions of Remarks from the Congressional Record comments Public comments on regulations.gov (via the Mirrulations mirror) documents The… See the full description on the dataset page: https://huggingface.co/datasets/ksasse/government-ai-detection.tabular1M<n<10M0 likes639 downloads1mo agoHugging Face30kshift /ahr999-dataset AHR999 BTC Hoarding Index Dataset Open, daily-updated AHR999 BTC hoarding index dataset, self-computed from Binance BTCUSDT daily closes and published as CSV and JSON. This Hugging Face repository is a mirror. The canonical dataset endpoints are: Dashboard: https://ahr999.aix4u.com/ GitHub: https://github.com/RuochenLyu/ahr999-dataset CSV endpoint: https://ahr999.aix4u.com/datasets/ahr999.csv JSON endpoint: https://ahr999.aix4u.com/datasets/ahr999.json Kaggle discovery mirror:… See the full description on the dataset page: https://huggingface.co/datasets/kshift/ahr999-dataset.tabular1K<n<10K1 likes638 downloads2h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.