CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VideoUFO /ResearchData_P10 likes19k downloads24d agoHugging Face02orionweller /mmBERT-pretrain-p1-fineweb2-langs mmBERT Pre-training Data P1 Phase 1 of 3: Diverse multilingual pre-training data mixture (trained for 2.3T tokens) used to train the mmBERT model suite. NOTE: this is only P1 of the pre-training data due to HF limits, you need to download and combine all three into one folderThis dataset contains the pre-training phase data used to train all mmBERT encoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository.… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-pretrain-p1-fineweb2-langs.fill-mask7 likes5.5k downloads11mo agoHugging Face03syvai /p1-segmentsgated DR P1 speech segments Dataset Danish speech clips from DR P1, in mono 16 kHz OGG/Opus, with verbatim text, timing, and speaker metadata. Transcript text and speaker attribution may contain automated errors. Source The recordings cover roughly 2006–2022 and come from DR P1 recordings in kb.dk’s DR archive. Audio is sourced through the pinned syvai/p1 revision 449b9c2294026df6d0d37538f279fdec03f565ff. Transcripts were generated with ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1-segments.audioautomatic-speech-recognition1M<n<10M4 likes2.2k downloads4d agoHugging Face04p1k0 /HBGimagen<1K0 likes2k downloads11d agoHugging Face05TaoMagnet /sn38-r11-p1textn<1K0 likes1.2k downloads13d agoHugging Face06P1n3 /SDG-30K SDG-30K — Structured Defect Grounding Dataset A 30,000-image dataset for structured defect grounding in text-to-image generations. Each image is annotated with bounding-box-level defects, where each defect carries: a category (artifact for visual flaws / misalignment for caption-image mismatches), a natural-language description, and a chain-of-thought reasoning trace. This is the public release accompanying the NeurIPS 2026 anonymous submission "SDG: Structured Defect… See the full description on the dataset page: https://huggingface.co/datasets/P1n3/SDG-30K.imageimage-text-to-text10K<n<100K1 likes1.1k downloads3mo agoHugging Face07syvai /p1 DR P1 Audio Archive Danish public radio (DR) P1 audio recordings sourced from the kb.dk DR-arkivet (Royal Danish Library DR archive), covering roughly 2006–2022. Format Audio: Opus, 24 kbps, mono, in OGG container (transcoded from DR's mp3 archive) Parquet shards (~500 items each), small row groups for streaming compatibility Sortable by year / month / start_time Schema Each row is one broadcast item with the full audio bytes inline plus rich… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1.audioautomatic-speech-recognition100K<n<1M2 likes889 downloads4mo agoHugging Face08junbrro /egopi_latal_gr1_p10 likes834 downloads2mo agoHugging Face09P1ayer-1 /stack-exchange-preferences-code-v2 Dataset Card for "stack-exchange-preferences-code-v2" More Information needed text1M<n<10M2 likes750 downloads3y agoHugging Face10P1ayer-1 /stack-exchange-preferences-code Dataset Card for "stack-exchange-preferences-code" More Information needed text1M<n<10M3 likes691 downloads3y agoHugging Face11Zeel /P1 all_in_one.zarr.zip <xarray.Dataset> Size: 25GB Dimensions: (Timestamp: 245376, station: 537) Coordinates: Timestamp (Timestamp) datetime64[ns] 2MB 2017-01-01 ... 2023-12-31T23:... address (station) <U187 402kB ... city (station) <U18 39kB ... latitude (station) float64 4kB ... longitude (station) float64 4kB ... state (station) <U17 37kB 'Chhattisgarh' ... 'West Bengal' station (station) <U64 137kB '32Bungalows, Bhilai - CECB'… See the full description on the dataset page: https://huggingface.co/datasets/Zeel/P1.0 likes482 downloads2y agoHugging Face12vuhoanhuy /viVoice-v1-p1audio100K<n<1M0 likes474 downloads9mo agoHugging Face13thonyyy /tatoeba-nusax-mt-p1text100M<n<1B0 likes402 downloads2y agoHugging Face14vlinhd11 /viVoice-v1-p1audio100K<n<1M0 likes392 downloads2y agoHugging Face15p1atdev /danbooru-2024image1M<n<10M14 likes382 downloads2y agoHugging Face16p11-p11 /chess_datasets Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.texttext-generation1M<n<10M0 likes329 downloads2y agoHugging Face17P147852369 /Insectsimage10K<n<100K1 likes327 downloads2y agoHugging Face18p12315132 /computational_limits_of_implicit_deductive_reasoningtabular1M<n<10M0 likes313 downloads5mo agoHugging Face19CCDS-Scaling /bbh-train-p1.0-bm25text100K<n<1M0 likes274 downloads2y agoHugging Face20Xiaodong /afd_mix_p100_verified Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning Overview We identify the Flow Moment, a new confident reasoning format in LLM reasoning, and we explore the influence of this format on On-Policy Self-Distillation (OPSD). We propose Aha-Flow Distillation, a dual-mode extension of OPSD that combines Aha branch and Flow branch to distill the student model. This training on Qwen3-8B model only takes ~30 minutes on 8×H20 and peaks within 100… See the full description on the dataset page: https://huggingface.co/datasets/Xiaodong/afd_mix_p100_verified.texttext-generation10K<n<100K0 likes230 downloads13d agoHugging Face21p1k0 /MIT10M-refine目前有测试集英文的标注集合test.json,图片在data文件夹下,分small, base, large尺寸 0 likes224 downloads1y agoHugging Face22p1k0 /visually-dependent-ambiguity VIDA: Visually-Dependent Ambiguity for Multimodal MT VIDA is an English-Chinese multimodal machine translation dataset for visual ambiguity resolution.Each instance contains an English source sentence, its paired image, and Chinese references that resolve annotated ambiguity spans using visual evidence. Paper: A Multimodal Dataset for Visually Grounded Ambiguity in Machine Translation Dataset composition This release contains four splits: Split Rows Description… See the full description on the dataset page: https://huggingface.co/datasets/p1k0/visually-dependent-ambiguity.imagetranslation1K<n<10K2 likes219 downloads5mo agoHugging Face23p1atdev /stackexchangestabular100K<n<1M1 likes217 downloads3y agoHugging Face24P1ayer-1 /annas-archive-index Dataset Card for "annas-archive-index" More Information needed tabular10M<n<100M1 likes214 downloads3y agoHugging Face25manhasambitionstrengthispoll /b7x92-kf841-p109ztext10K<n<100K0 likes202 downloads9d agoHugging Face26ccwatson /pick_place_avoid_and_not_avoid_calculator_p1_molmobot pick_place_avoid_and_not_avoid_calculator_p1_molmobot MolmoBot-format dataset for Synthmanip/MolmoBot training. Layout dataset_manifest.json train/valid_trajectory_index.json and val/valid_trajectory_index.json train/house_*/*.h5 and val/house_*/*.h5 HDF5 video sidecars under the same split/house directories normalization stats: pick_place_avoid_and_not_avoid_calculator_molmobot_norm_stats.yaml Split Summary split houses h5 files… See the full description on the dataset page: https://huggingface.co/datasets/ccwatson/pick_place_avoid_and_not_avoid_calculator_p1_molmobot.videon<1K0 likes198 downloads3mo agoHugging Face27priyanshu389 /LMA_INDIVIDUAL_PROJECT_P1text1M<n<10M0 likes194 downloads1mo agoHugging Face28P1ayer-1 /books-3-textbooks Dataset Card for "books-3-textbooks" More Information needed text1K<n<10K13 likes181 downloads3y agoHugging Face29P1ayer-1 /isbndb-full-database Dataset Card for "isbndb-annas" More Information needed text10M<n<100M8 likes180 downloads3y agoHugging Face30P1ayer-1 /College-Texts-pt1text10K<n<100K1 likes168 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.