p1
Models
All models matching “p1”Datasets
All datasets matching “p1”ResearchData_P1mmBERT-pretrain-p1-fineweb2-langs
mmBERT Pre-training Data P1
Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite.
NOTE: this is only P1 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p1-fineweb2-langs.p1-segments
DR P1 speech segments
Dataset
Danish speech clips from DR P1, in mono 16 kHz OGG/Opus, with verbatim text, timing, and speaker metadata. Transcript text and speaker attribution may contain automated errors.
Source
The recordings cover roughly 2006–2022 and come from DR P1 recordings in kb.dk’s DR archive. Audio is sourced through the pinned syvai/p1 revision 449b9c2294026df6d0d37538f279fdec03f565ff. Transcripts were generated with ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1-segments.HBGsn38-r11-p1SDG-30K
SDG-30K — Structured Defect Grounding Dataset
A 30,000-image dataset for structured defect grounding in text-to-image
generations. Each image is annotated with bounding-box-level defects, where
each defect carries:
a category (artifact for visual flaws / misalignment for caption-image
mismatches),
a natural-language description, and
a chain-of-thought reasoning trace.
This is the public release accompanying the NeurIPS 2026 anonymous submission
"SDG: Structured Defect… See the full description on the dataset page: https://huggingface.co/datasets/P1n3/SDG-30K.
