datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NS-SVS
Short Description
Based on the incompressible Navier-Stokes equations, this dataset contains trajectories starting from sinusoidal vortex sheet initial conditions, see https://arxiv.org/abs/2405.19101.
It also contains a passive tracer that is carried by the flow.
Dimensions
The assembled NetCDF file has a single variable called velocity with dimensionality
20000 (number of trajectories)
21 (time steps)
3 (horizontal velocity, vertical velocity, passive tracer)
128… See the full description on the dataset page: https://huggingface.co/datasets/camlab-ethz/NS-SVS.svs-lame-compression-jpeg-vs-neural
Comparaison visuelle : compression neuronale vs JPEG sur lames histopathologiques
Ce dataset permet à un anatomopathologiste de juger à l'œil nu si une image
de lame numérique compressée par un réseau de neurones est visuellement
équivalente à la même lame compressée en JPEG (qui est le standard)
En une phrase
On a pris 5 lames histopathologiques au format SVS, on les a compressées avec
4 modèles neuronaux et avec JPEG Q75, à
deux niveaux d'agressivité (q5 ≈… See the full description on the dataset page: https://huggingface.co/datasets/nathbns/svs-lame-compression-jpeg-vs-neural.sv-SE-asr-cv
Swedish ASR (Common Voice 22, filtered + rebalanced)
Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via
the open fsicoli/common_voice_22_0 mirror. Built to fine-tune
streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
train
22166
26.1
dev
694
0.8
test
1602
2.0
train = the official validated train split + the filtered other bucket + the excess
dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/sv-SE-asr-cv.BigEarthNet.txt
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
BigEarthNet.txt is a large-scale multi-sensor image–text dataset for Earth observation, designed to advance vision–language learning on remote sensing data. It comprises 464,044 co-registered Sentinel-1 (SAR) and… See the full description on the dataset page: https://huggingface.co/datasets/SVSG/BigEarthNet.txt.midi-svs
[WIP] MIDI SVS
A richly annotated English Suno vocal+MIDI dataset featuring 3k+ curated tracks with stems, transcriptions, lyrics, and structural metadata for SVS music AI and MIR purposes
Attribution
Suno Various 94k
Facebook Demucs
ASLP-lab SongFormer
LinTO AI Whisper Timestamped
ROSVOT
Project Los Angeles
Tegridy Code 2026
Variational-DAPO
Dataset Card for SvS/Variational-DAPO
[🌐 Website] •
[🤗 Dataset] •
[📜 Paper] •
[🐱 GitHub] •
[🐦 Twitter] •
[📕 Rednote]
This dataset consists of 314k variational problems synthesized by the Qwen2.5-32B-Instruct policy during RLVR training on DAPO-17k using the SvS strategy for 600-step training, each accompanied by reference answers.The variational problems undergo a min_hash deduplication with a threshold of 0.85.
Data Loading
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/RLVR-SvS/Variational-DAPO.svs-subspace-validity-suite
Subspace Validity Suite (SVS)
Diagnostic toolkit for validating "visual directions" in Vision-Language Models.
Paper: "What PCA-Based Visual Directions in VLMs Actually Capture" (WACV 2027)
Installation
git clone https://huggingface.co/datasets/Anonymousblind/svs-subspace-validity-suite
cd svs-subspace-validity-suite
pip install .
Quick Start
from svs import SubspaceValiditySuite
svs = SubspaceValiditySuite()
report = svs.full_report(… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousblind/svs-subspace-validity-suite.SVS-TCGA-2048brainrotinterspeech2024_discrete_speech_svs_resultsPathGene-CSU_svsWithout the consent of Xiangya Hospital of Central South University or its Pathology Department, the use of the raw data is prohibited. Should any data be used for commercial purposes, we will hold the user legally accountable.
We support the following 21 pre-trained foundation models to extract the feature representation of WSI. Please contact us by email before using. (Strongly recommended!!)
Patch Encoder
Embedding Dim
Args
Link
UNI
1024
--patch_encoder uni_v1 --patch_size 256… See the full description on the dataset page: https://huggingface.co/datasets/LiangruiPan/PathGene-CSU_svs.sln-breast-tcia-svsSVS-TCGA-BRReal-Time-SVSDF-PlannerSVS_huststeepHãy git clone hoặc tải từ huggingface chứ đừng tải tay :skull:
hffsssvs-qdora-datasln-breast-tcia-svs-lfssv-stat-tablesgewd791data_jobs
🧠 data_jobs Dataset
A dataset of real-world data analytics job postings from 2023, collected and processed by Luke Barousse.
Background
I've been collecting data on data job postings since 2022. I've been using a bot to scrape the data from Google, which come from a variety of sources.
You can find the full dataset at my app datanerd.tech.
Serpapi has kindly supported my work by providing me access to their API. Tell them I sent you and get 20% off paid plans.… See the full description on the dataset page: https://huggingface.co/datasets/SVSA/data_jobs.text-to-svspritesgggdsvdssdvhhrthrtrhrtdrrrhdhrdhrrdhrthhthdrdrrhhhdhbxxxbbhhhdhdrhdrdrhddh
