datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sv-SE-asr-cv
Swedish ASR (Common Voice 22, filtered + rebalanced)
Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via
the open fsicoli/common_voice_22_0 mirror. Built to fine-tune
streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
train
22166
26.1
dev
694
0.8
test
1602
2.0
train = the official validated train split + the filtered other bucket + the excess
dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/sv-SE-asr-cv.BigEarthNet.txt
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
BigEarthNet.txt is a large-scale multi-sensor image–text dataset for Earth observation, designed to advance vision–language learning on remote sensing data. It comprises 464,044 co-registered Sentinel-1 (SAR) and… See the full description on the dataset page: https://huggingface.co/datasets/SVSG/BigEarthNet.txt.Variational-DAPO
Dataset Card for SvS/Variational-DAPO
[🌐 Website] •
[🤗 Dataset] •
[📜 Paper] •
[🐱 GitHub] •
[🐦 Twitter] •
[📕 Rednote]
This dataset consists of 314k variational problems synthesized by the Qwen2.5-32B-Instruct policy during RLVR training on DAPO-17k using the SvS strategy for 600-step training, each accompanied by reference answers.The variational problems undergo a min_hash deduplication with a threshold of 0.85.
Data Loading
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/RLVR-SvS/Variational-DAPO.SVS-TCGA-2048brainrotinterspeech2024_discrete_speech_svs_resultssln-breast-tcia-svsSVS-TCGA-BRdata_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/SVSA/data_jobs.sln-breast-tcia-svs-lfsmy_svsSVsDataset
