datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ResearchData_P1mmBERT-pretrain-p1-fineweb2-langs
mmBERT Pre-training Data P1
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P1 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p1-fineweb2-langs.p1-segments
DR P1 speech segments
Dataset
Danish speech clips from DR P1, in mono 16 kHz OGG/Opus, with verbatim text, timing, and speaker metadata. Transcript text and speaker attribution may contain automated errors.
Source
The recordings cover roughly 2006–2022 and come from DR P1 recordings in kb.dk’s DR archive. Audio is sourced through the pinned syvai/p1 revision 449b9c2294026df6d0d37538f279fdec03f565ff. Transcripts were generated with ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1-segments.HBGsn38-r11-p1SDG-30K
SDG-30K — Structured Defect Grounding Dataset
A 30,000-image dataset for structured defect grounding in text-to-image
generations. Each image is annotated with bounding-box-level defects, where
each defect carries:
a category (artifact for visual flaws / misalignment for caption-image
mismatches),
a natural-language description, and
a chain-of-thought reasoning trace.
This is the public release accompanying the NeurIPS 2026 anonymous submission
"SDG: Structured Defect… See the full description on the dataset page: https://huggingface.co/datasets/P1n3/SDG-30K.p1
DR P1 Audio Archive
Danish public radio (DR) P1 audio recordings sourced from the kb.dk DR-arkivet (Royal Danish Library DR archive), covering roughly 2006–2022.
Format
Audio: Opus, 24 kbps, mono, in OGG container (transcoded from DR's mp3 archive)
Parquet shards (~500 items each), small row groups for streaming compatibility
Sortable by year / month / start_time
Schema
Each row is one broadcast item with the full audio bytes inline plus rich… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1.egopi_latal_gr1_p1stack-exchange-preferences-code-v2
Dataset Card for "stack-exchange-preferences-code-v2"
More Information needed
stack-exchange-preferences-code
Dataset Card for "stack-exchange-preferences-code"
More Information needed
P1
all_in_one.zarr.zip
<xarray.Dataset> Size: 25GB
Dimensions: (Timestamp: 245376, station: 537)
Coordinates:
Timestamp (Timestamp) datetime64[ns] 2MB 2017-01-01 ... 2023-12-31T23:...
address (station) <U187 402kB ...
city (station) <U18 39kB ...
latitude (station) float64 4kB ...
longitude (station) float64 4kB ...
state (station) <U17 37kB 'Chhattisgarh' ... 'West Bengal'
station (station) <U64 137kB '32Bungalows, Bhilai - CECB'… See the full description on the dataset page: https://huggingface.co/datasets/Zeel/P1.viVoice-v1-p1tatoeba-nusax-mt-p1viVoice-v1-p1danbooru-2024chess_datasets
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.Insectscomputational_limits_of_implicit_deductive_reasoningbbh-train-p1.0-bm25afd_mix_p100_verified
Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning
Overview
We identify the Flow Moment, a new confident reasoning format in LLM reasoning, and we explore the influence of this format on On-Policy Self-Distillation (OPSD). We propose Aha-Flow Distillation, a dual-mode extension of OPSD that combines Aha branch and Flow branch to distill the student model. This training on Qwen3-8B model only takes ~30 minutes on 8×H20 and peaks within 100… See the full description on the dataset page: https://huggingface.co/datasets/Xiaodong/afd_mix_p100_verified.MIT10M-refine目前有测试集英文的标注集合test.json,图片在data文件夹下,分small, base, large尺寸
visually-dependent-ambiguity
VIDA: Visually-Dependent Ambiguity for Multimodal MT
VIDA is an English-Chinese multimodal machine translation dataset for visual ambiguity resolution.Each instance contains an English source sentence, its paired image, and Chinese references that resolve annotated ambiguity spans using visual evidence.
Paper: A Multimodal Dataset for Visually Grounded Ambiguity in Machine Translation
Dataset composition
This release contains four splits:
Split
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/p1k0/visually-dependent-ambiguity.stackexchangesannas-archive-index
Dataset Card for "annas-archive-index"
More Information needed
b7x92-kf841-p109zpick_place_avoid_and_not_avoid_calculator_p1_molmobot
pick_place_avoid_and_not_avoid_calculator_p1_molmobot
MolmoBot-format dataset for Synthmanip/MolmoBot training.
Layout
dataset_manifest.json
train/valid_trajectory_index.json and val/valid_trajectory_index.json
train/house_*/*.h5 and val/house_*/*.h5
HDF5 video sidecars under the same split/house directories
normalization stats: pick_place_avoid_and_not_avoid_calculator_molmobot_norm_stats.yaml
Split Summary
split
houses
h5 files… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/pick_place_avoid_and_not_avoid_calculator_p1_molmobot.LMA_INDIVIDUAL_PROJECT_P1books-3-textbooks
Dataset Card for "books-3-textbooks"
More Information needed
isbndb-full-database
Dataset Card for "isbndb-annas"
More Information needed
College-Texts-pt1
