CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LokaalHub /sv-SE-asr-cv Swedish ASR (Common Voice 22, filtered + rebalanced) Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 22166 26.1 dev 694 0.8 test 1602 2.0 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/sv-SE-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes101 downloads4mo agoHugging Face02SVSG /BigEarthNet.txt BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation BigEarthNet.txt is a large-scale multi-sensor image–text dataset for Earth observation, designed to advance vision–language learning on remote sensing data. It comprises 464,044 co-registered Sentinel-1 (SAR) and… See the full description on the dataset page: https://huggingface.co/datasets/SVSG/BigEarthNet.txt.tabularimage-text-to-text1M<n<10M0 likes62 downloads17d agoHugging Face03RLVR-SvS /Variational-DAPO Dataset Card for SvS/Variational-DAPO [🌐 Website] • [🤗 Dataset] • [📜 Paper] • [🐱 GitHub] • [🐦 Twitter] • [📕 Rednote] This dataset consists of 314k variational problems synthesized by the Qwen2.5-32B-Instruct policy during RLVR training on DAPO-17k using the SvS strategy for 600-step training, each accompanied by reference answers.The variational problems undergo a min_hash deduplication with a threshold of 0.85. Data Loading from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/RLVR-SvS/Variational-DAPO.textquestion-answering100K<n<1M4 likes51 downloads1y agoHugging Face04nathbns /SVS-TCGA-2048imagen<1K0 likes40 downloads5mo agoHugging Face05svshrithik12 /brainrottext1K<n<10K3 likes19 downloads1y agoHugging Face06discrete-speech /interspeech2024_discrete_speech_svs_resultstextn<1K0 likes16 downloads2y agoHugging Face07Refrainkana33 /sln-breast-tcia-svstextn<1K0 likes12 downloads3mo agoHugging Face08nathbns /SVS-TCGA-BRimage10K<n<100K0 likes9 downloads5mo agoHugging Face09SVSA /data_jobs 🧠 data_jobs Dataset A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse. Background I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources. You can find the full dataset at my app datanerd.tech. Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/SVSA/data_jobs.tabular100K<n<1M0 likes4 downloads5mo agoHugging Face10Refrainkana33 /sln-breast-tcia-svs-lfstextn<1K0 likes4 downloads3mo agoHugging Face11folkSci /my_svstabular100K<n<1M0 likes2 downloads5mo agoHugging Face12HLovisiEnnes /SVsDatasetgatedtext10K<n<100K1 likes1 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.