datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaiaThis catalog is developed for use with the Siril 1.4 series as a public reference database. Hugging Face is one of several mirrors used to distribute the data. This database is provided for both offline download and also for online access. This dataset is provided for scientific and reproducibility purposes.
This is an extract of the Gaia DR3 catalog optimized for spectrophotometric color calibration. The catalog is indexed at HEALpix level 8 and selects up to the 127 brightest sources in each… See the full description on the dataset page: https://huggingface.co/datasets/siril-spcc/gaia.SPC
Dataset Card for "SPC-v2"
More Information needed
SP_CountingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 37501,
"total_tasks": 5,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Counting.spc_r
Dataset Card: Swiss Parliaments Corpus — SPC_R v1.0
Background
The aim is to build a large, high-quality dataset. To get there, we correct pseudo-labeled transcriptions of parliamentary debates with an LLM. The model receives semantically relevant chunks from a manually prepared session protocol as context and then produces the corrected transcription.
We also show that Whisper’s average log probability can be used to predict BLEU. This lets us estimate transcription… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/spc_r.SP_Counting_200This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Counting_200.MF_RPN_convection_super_param_CAM5_SPCAM5
Probabilistic Multi-fidelity climate model parameterization for better generalization and extrapolation
Code and data accompanying the manuscript titled "Multi-fidelity climate model parameterization for better generalization and extrapolation", authored by Mohamed Aziz Bhouri, Liran Peng, Michael S Pritchard and Pierre Gentine.
Abstract
Machine-learning-based parameterizations (i.e. representation of sub-grid processes) of global climate models or turbulent simulations… See the full description on the dataset page: https://huggingface.co/datasets/MohamedAzizBhouri/MF_RPN_convection_super_param_CAM5_SPCAM5.arxiv-for-fanns-large
arXiv Dataset to Evaluate (Filtered) Approximate Nearest Neighbor Search
This dataset is intended to benchmark Approximate Nearest Neighbor Search (ANNS) and Filtered Approximate Nearest Neighbor Search (FANNS) algorithms. It is based on the arXiv Dataset whose paper abstracts are embedded using the stella_en_400M_v5 embedding model. 10,000 unique arXiv search terms were generated by GPT-4 and embedded using the same Stella model to obtain the query vectors. Query attributes for… See the full description on the dataset page: https://huggingface.co/datasets/SPCL/arxiv-for-fanns-large.arxiv-for-fanns-small
arXiv Dataset to Evaluate (Filtered) Approximate Nearest Neighbor Search
This dataset is intended to benchmark Approximate Nearest Neighbor Search (ANNS) and Filtered Approximate Nearest Neighbor Search (FANNS) algorithms. It is based on the arXiv Dataset whose paper abstracts are embedded using the stella_en_400M_v5 embedding model. 10,000 unique arXiv search terms were generated by GPT-4 and embedded using the same Stella model to obtain the query vectors. Query attributes for… See the full description on the dataset page: https://huggingface.co/datasets/SPCL/arxiv-for-fanns-small.arxiv-for-fanns-medium
arXiv Dataset to Evaluate (Filtered) Approximate Nearest Neighbor Search
This dataset is intended to benchmark Approximate Nearest Neighbor Search (ANNS) and Filtered Approximate Nearest Neighbor Search (FANNS) algorithms. It is based on the arXiv Dataset whose paper abstracts are embedded using the stella_en_400M_v5 embedding model. 10,000 unique arXiv search terms were generated by GPT-4 and embedded using the same Stella model to obtain the query vectors. Query attributes for… See the full description on the dataset page: https://huggingface.co/datasets/SPCL/arxiv-for-fanns-medium.spcThis is a collection of parallel corpora collected by Hercules Dalianis and his research group for bilingual dictionary construction.
More information in: Hercules Dalianis, Hao-chun Xing, Xin Zhang: Creating a Reusable English-Chinese Parallel Corpus for Bilingual Dictionary Construction, In Proceedings of LREC2010 (source: http://people.dsv.su.se/~hercules/SEC/) and Konstantinos Charitakis (2007): Using Parallel Corpora to Create a Greek-English Dictionary with UPLUG, In Proceedings of NODALIDA 2007. Afrikaans-English: Aldin Draghoender and Mattias Kanhov: Creating a reusable English – Afrikaans parallel corpora for bilingual dictionary construction
4 languages, 3 bitexts
total number of files: 6
total number of tokens: 1.32M
total number of sentence fragments: 0.15Mxiang2021-spcas9
🧬 Xiang 2021 SpCas9 On-Target Efficiency
Merged dataset from Luo 2020 and Kim 2019 for SpCas9 on-target efficiency prediction. Published in Nature Communications. Contains 30mer sequences with measured CRISPR-Cas9 cutting efficiency.
Dataset Details
Task: On Target Efficiency
Nuclease: SpCas9 UniProt CasPEDIA
Organism: Human
Cell Line: HEK293T
Paper: https://doi.org/10.1038/s41467-021-23576-0
Usage
from datasets import load_dataset
# Standard… See the full description on the dataset page: https://huggingface.co/datasets/saiden89/xiang2021-spcas9.spc-factor-results
Results for Keypoint-based Stereophotoclinometry for Characterizing and Navigating Small Bodies: A Factor Graph Approach presented at the 2024 AIAA SciTech Forum
See example.ipynb for instructions on loading and manipulating the reconstructions.
If you utilize our reconstructions or data, please our paper:
@inproceedings{driver2023spc,
title={Keypoint-based Stereophotoclinometry for Characterizing and Navigating Small Bodies: A Factor Graph Approach},
author={Driver, Travis and… See the full description on the dataset page: https://huggingface.co/datasets/travisdriver/spc-factor-results.SPC
Dataset Card: Swiss Parliaments Corpus — Train v0.9
Summary
The SPC Train v0.9 release pairs Swiss German speech with Standard German transcriptions, providing a high‑quality resource for training and evaluating automatic speech‑recognition (ASR) or speech‑translation systems.
If you intend to fine‑tune Whisper, we recommend the companion project i4Ds/whisper‑finetune, which is fully compatible with the data structure produced here.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/SPC.rsna2025-aneurysm-26class-seg
RSNA 2025 Aneurysm — 26-class Vessel-Anatomy and Aneurysm Segmentation Labels
26-class vessel-anatomy and aneurysm segmentation labels (13 vessel anatomy
classes + 13 aneurysm location classes, values 0–26, see labels.json) for
the RSNA 2025 Intracranial Aneurysm Detection challenge, placed back into the
original image space (pure voxel placement, no resampling).
Paper: arXiv:2606.26706
Contents
Folder
Count
Aligned to
labelsTr_26classes_in_orig_space/… See the full description on the dataset page: https://huggingface.co/datasets/spc819/rsna2025-aneurysm-26class-seg.spc-tornado-history
NOAA / SPC Tornado History (1950–2024)
Bundled for Aerostratospheric / UOGW tornado-trend research. Source of record: NOAA Storm Prediction Center Severe Weather Database.
Contents
1950-2024_all_tornadoes.csv / 1950-2024_actual_tornadoes.csv — SPC actual tornado segment CSV
schema.json — provenance metadata
Bundled UTC: 2026-09-18T15:47:02Z
License / attribution
US Government work redistributed by SPC for public use. You must attribute NOAA/SPC… See the full description on the dataset page: https://huggingface.co/datasets/aerostratospheric/spc-tornado-history.eurosat_spcspc_r_segmented
i4ds/spc_r_segmented
Diarized and segmented speech dataset derived from i4ds/spc_r.
Description
Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline:
Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap.
Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap between… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/spc_r_segmented.spcas9Fifa_Playersspc_r_whisperspc_r_segmented
eko57/spc_r_segmented
Diarized and segmented speech dataset derived from i4ds/spc_r.
Description
Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline:
Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap.
Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap… See the full description on the dataset page: https://huggingface.co/datasets/eko57/spc_r_segmented.afternight-spcc-responses
AfterNight SPCC Response Repositories
This dataset contains reviewed AfterNight Spectrophotometric Color Calibration response repositories staged for direct
HTTPS download.
Publication channel: preview
Direct download base: https://huggingface.co/datasets/emruiz81/afternight-spcc-responses/resolve/main
Response Repositories
Repository ID
Version
Status
Profiles
Size
Rights
Path
afternight-spcc-response-baseline
2026a
current
2,342
26.3 MiB… See the full description on the dataset page: https://huggingface.co/datasets/emruiz81/afternight-spcc-responses.spc-pick-screwdriverThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_slider_follower",
"total_episodes": 51,
"total_frames": 12233,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/spc-pick-screwdriver.SPC_test
Dataset Card: Swiss Parliaments Corpus — Test
Summary
The SPC Train v0.9 release pairs Swiss German speech with Standard German transcriptions, providing a high‑quality resource for training and evaluating automatic speech‑recognition (ASR) or speech‑translation systems.
Dataset Details
Maintainer
Curated by: Vincenzo Timmel (@vincenzo.timmel)
Intended Use & Scope
Primary use‑case: Evaluation for Swiss-German STT Systems.… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/SPC_test.sp-chinese-translated-annotationsspc-pick-stuffThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_slider_follower",
"total_episodes": 154,
"total_frames": 39338,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:154"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/spc-pick-stuff.SP_Counting_200
SP_Counting_200 (TsFile)
Apache TsFile version of justintiensmith/SP_Counting_200.
Overview
A LeRobot robot-manipulation dataset. The source card is auto-generated ("This dataset was created using LeRobot") and does not add a free-text description; the facts below are taken from the source meta/info.json.
Robot: so_follower
Episodes: 200
Frames: 96721
Sampling rate: 30 fps
Splits: a single train split
Schema (TsFile structure)
Time (INT64… See the full description on the dataset page: https://huggingface.co/datasets/THULab/SP_Counting_200.CALE-SPCD
SPCD (Semcor Pairs for Concept Differentiation)
africa-strategic-partnership-cooperation-framework-spcf
Strategic Partnership Cooperation Framework SPCF | Africa (original)
Size category: n<1K - Formats: parquet - Sector: humanitarian_development - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-strategic-partnership-cooperation-framework-spcf.SP_Counting_200_20260724_122814This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Counting_200_20260724_122814.
