datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mroot-q17-scoreband-runtime-v1mroot-q17-scoreband-runtime-bridge-v1astroPT_euclid_Q1_desi_dr1_dataset
DESI DR1 × Euclid Q1 Dataset
Dataset Description
This dataset contains the cross-matched catalog between DESI Data Release 1 (DR1) and Quick Data Release 1 (Q1) data.
Dataset Summary
Total Sources: 39,850 galaxies
Cross-match Radius: 0.5 arcsec
Data Products:
Euclid VIS imaging
Euclid NISP imaging (Y, J, H bands)
DESI spectra
Source Data
DESI Sample Selection: Applied quality cuts to DESI DR1:
ZCAT_PRIMARY == True
ZWARN == 0 or ZWARN == 4… See the full description on the dataset page: https://huggingface.co/datasets/msiudek/astroPT_euclid_Q1_desi_dr1_dataset.common_voice_26_0_de
Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset)
Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence.
📊 Dataset parameters
Origin: Mozilla Common Voice 26.0 (version 18/06/2026).
Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.mmu_euclid_q1german-golden-audio_speech-IPA
🌟 German Golden Speech & IPA Corpus (FLEURS + Multilingual TEDx)
An ultra-clean, high-standard curated German speech dataset combining Google FLEURS (de_de) and Multilingual TEDx German (mTEDx), fully embedded with 16kHz WAV audio bytes, normalized orthographic text, and pre-computed International Phonetic Alphabet (IPA) transcriptions.
📊 Dataset Summary
Total Samples: 1,354 high-quality audio recordings.
Total Size: ~419 MB (Compressed Parquet format).
Audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-golden-audio_speech-IPA.euclid-Q1-VF
Euclid Q1: Generative Neural Network Training Database for Shear Estimation
Dataset Summary
This dataset contains a collection of image stamps extracted from the Euclid (Q1)
space mission data. It was specifically designed for training generative neural
networks applied to cosmic shear estimation within the context of weak lensing
studies.
The images are centered on pre-filtered galaxies (spatially isolated, without major
pixel defects) and provide all the… See the full description on the dataset page: https://huggingface.co/datasets/VincentB03/euclid-Q1-VF.OSHA_Citations_Q1_2026OSHA Citations — Q1 2026 (Citation-Level)
This is a mirror. Canonical home: https://www.fastdol.com/datasets/osha-citations-2026-q1
License: CC BY 4.0
Version DOI: 10.5281/zenodo.20138442
Concept DOI: 10.5281/zenodo.20138441
Download CSV: https://www.fastdol.com/datasets/osha-citations-2026-q1/data.csv
Visit the canonical page for the full schema, methodology, BibTeX citation, and most recent version.
OSHA Citations — Q1 2026 (Citation-Level)
Every OSHA citation issued in the first… See the full description on the dataset page: https://huggingface.co/datasets/FastDOLz/OSHA_Citations_Q1_2026.common_voice_spontaneous_speech_4_0_de
Mozilla Common Voice Spontaneous Speech 4.0 - German (Parquet & IPA)
Clean Parquet format of Mozilla Spontaneous Speech 4.0 German (264 samples).
euclid_q1euclid-Q1-1obs-RepoTest
Build info
Automatically appended by src/main.py at push time (2026-09-04 16:04 UTC). Everything above this line is preserved as-is; only this block is regenerated on each push.
Rows in this dataset: 8223
Observation IDs (3): 2696, 2697, 2698
Command used to build this dataset
python src/main.py --obs-ids 2696 2697 2698 --drop-duplicates --zero-flagged-pixels --push --public --repo-id VincentB03/euclid-Q1-1obs-RepoTest
Columns
column… See the full description on the dataset page: https://huggingface.co/datasets/VincentB03/euclid-Q1-1obs-RepoTest.euclid-Q1-3obs-DropDupDropEmpty
Build info
Automatically appended by src/main.py at push time (2026-09-05 12:12 UTC). Everything above this line is preserved as-is; only this block is regenerated on each push.
Rows in this dataset: 7876
Observation IDs (3): 2696, 2697, 2698
Command used to build this dataset
python src/main.py --obs-ids 2696 2697 2698 --drop-duplicates --zero-flagged-pixels --drop-empty-stamps --push --public --repo-id VincentB03/euclid-Q1-3obs-DropDupDropEmpty… See the full description on the dataset page: https://huggingface.co/datasets/VincentB03/euclid-Q1-3obs-DropDupDropEmpty.Q_10_datasetq1-bin-prediction
Q1 — Out-of-View Object Direction Prediction (Bin Classification)
Anonymized release for double-blind NeurIPS 2026 Evaluations &
Datasets review. Author identity will be revealed at acceptance time;
see the paper for full citation.
Companion eval-code repo
Reference implementations of the loaders, sighted-prompt eval, and
metric pipelines (calibrated NLL family, group-KL/JSD vs GT pairwise
distribution) live in a separate anonymous repo:
→… See the full description on the dataset page: https://huggingface.co/datasets/nipsedtrack2026/q1-bin-prediction.Euclid-Q1-postage-stamps
Build info
Automatically appended by src/main.py at push time (2026-09-22 18:33 UTC). Everything above this line is preserved as-is; only this block is regenerated on each push.
Rows in this dataset: 265196
Observation IDs (76): 2681, 2682, 2683, 2684, 2685, 2686, 2687, 2688, 2689, 2690, 2691, 2692, 2693, 2694, 2695, 2696, 2697, 2698, 2699, 2700, 2701, 2702, 2703, 2704, 2705, 2706, 2707, 2708, 2709, 2710, 2711, 2713, 2714, 2715, 2716, 2717, 2718, 2719, 2720, 2721, 2722, 2723… See the full description on the dataset page: https://huggingface.co/datasets/VincentB03/Euclid-Q1-postage-stamps.synpre_extract_q10_a5_1M_a_first
Dataset Card for "synpre_extract_q10_a5_1M_a_first"
More Information needed
german-pronuncheck-mega-dataset
German PronunCheck Mega Dataset 🇩🇪
Dataset Summary
This is a highly curated, 123GB+ mega-dataset designed specifically for training and fine-tuning German Automatic Speech Recognition (ASR) and Computer-Assisted Pronunciation Training (CAPT) models, such as HuBERT and Wav2Vec2.
Composition
This dataset is a clean concatenation of three distinct open-source datasets:
Mozilla Common Voice 26.0 (German): Standard scripted crowdsourced speech.… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-pronuncheck-mega-dataset.oalz-1788-q1-ner-annotations-merged-union-dataset
OALZ/1788/Q1/NER
Postprocessing
Training
Published datasets (union, merged union) and models (EVENT, LOC, MISC, ORG, PER, TIME)
A named entity recognition system (NER) was trained on text extracted from Oberdeutsche Allgemeine Litteraturueitung (OALZ) of the first quarter (January, Febuary, March) of 1788. The scans from which text was extracted can be found at Bayerische Staatsbibliothek using the extraction strategy of the KEDiff project, which can be found at cborgelt/KEDiff.… See the full description on the dataset page: https://huggingface.co/datasets/LelViLamp/oalz-1788-q1-ner-annotations-merged-union-dataset.helpsteer2_tail_modeldep_q10synpre_extract_q10_a5_1M
Dataset Card for "synpre_extract_q10_a5_1M"
More Information needed
oalz-1788-q1-ner-annotations-union-dataset
OALZ/1788/Q1/NER
Postprocessing
Training
Published datasets (union, merged union) and models (EVENT, LOC, MISC, ORG, PER, TIME)
A named entity recognition system (NER) was trained on text extracted from Oberdeutsche Allgemeine Litteraturueitung (OALZ) of the first quarter (January, Febuary, March) of 1788. The scans from which text was extracted can be found at Bayerische Staatsbibliothek using the extraction strategy of the KEDiff project, which can be found at cborgelt/KEDiff.… See the full description on the dataset page: https://huggingface.co/datasets/LelViLamp/oalz-1788-q1-ner-annotations-union-dataset.FinQA_Q1africa-somalia-3w-q1-2015
Somalia - Who is doing what and where (3W) - 2015 | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: humanitarian_development - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-3w-q1-2015.institutional-core-2024-q1
Institutional Core 2024 Q1 — Market Data Product
Production-grade, backtest-ready market data for 12 core asset proxies spanning
FX, Equities, Crypto, and Commodities. Built on the Universal Market Data Platform
with full lineage tracing, quality certification, and reproducibility receipts.
Dataset Details
Property
Value
Period
2024-01-01 → 2024-01-14
Base timeframe
m1 (1-minute candles)
Total rows
238,260
Certified
✅ is_certified=True
Completeness
1.0… See the full description on the dataset page: https://huggingface.co/datasets/TommyKwok/institutional-core-2024-q1.Natural-Farming-Real-QandA-Conversations-Q1-2024-Updatevietnam_traffic_signq1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/angjminer/q1.sp500_dataset_annotated_2025_earnings_q1euclid_q1_embeddingshelpsteer2_tail_intrinsic_q10
