CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /Wikipedia-AbstractWikipedia Abstract Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards. A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.texttext-classification10M<n<100M9 likes4.2k downloads2y agoHugging Face02laion /biorXiv-pdf BiorXiv Pdf BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets. BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.documentfeature-extraction1K<n<10K4 likes2.1k downloads2y agoHugging Face03laion /voice-acting-cutscene-prompts Cut-Scene Voice-Acting Prompts Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting models. Each prompt describes a single speaker across two sharply contrasting emotional moments separated by a CUT TO: transition, in a voice-acting stage-direction format (spoken lines in "quotes", performance notes in (parentheses)). Total prompts: 4,057,000 Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.tabulartext-generation1M<n<10M2 likes993 downloads13d agoHugging Face04laion /llama-nemotron-science-reasoning-on-canonical-think-full Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter) The complete reasoning:on science split of nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical Delphi chat-template thinking format. 708,920 rows. Unlike the cold-start warmup slice open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k (and its -canonical-think variant), this build applies no length cap and no subsample — every long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.texttext-generation100K<n<1M0 likes379 downloads20d agoHugging Face05laion /terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827 TaskTrove DQ unitsyn-python training traces (step 20, 30B-A3B) Terminus-2 agent rollouts recorded while training laion/tasktrove-dq-unitsyn-python-step20-30b-a3b with SkyRL from Qwen/Qwen3-Coder-30B-A3B-Instruct. Each row is the last episode of one trial: the full agent transcript, the task instruction, the scalar reward, and the verifier's output. Source run: rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-163115-1ae770. Coverage This dataset is the complete… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827.texttext-generation1K<n<10K0 likes326 downloads2mo agoHugging Face06laion /Wikipedia-ExoticWikipedia Exotic is an extensive dataset designed to offer Wikipedia corpus in uncommon European and religious languages. Wikimedia and other reputable sources maintain existing corpuses, housing 300+ languages and their respective corpora. Additionally, datasets such as Wikipedia 40B from Google provide corpora in 40 widely spoken languages. LAION AI sought to make a distinct and long-lasting contribution to the field by filling a void that would enable researchers to work with exotic… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Exotic.summarization2 likes249 downloads2y agoHugging Face07laion /Project-Gutenberg Project Gutenberg Introducing Project Gutenberg, a dataset that provides access to all the books available in that project. In our dataset, we wanted to provide a bulk download option to have access to Gutenberg books in ten different languages such as English, German, French, Polish, Portuguese, Dutch, Spanish, Hebrew, Russian and Chinese. English has the largest collection of books, followed by German. We are releasing this dataset for researchers and engineers to integrate… See the full description on the dataset page: https://huggingface.co/datasets/laion/Project-Gutenberg.summarization8 likes245 downloads2y agoHugging Face08laion /COREX-18textCORE-18 Fulltext Introducing the CORE-18 Full Text dataset, among the first well-maintained public datasets of CORE. CORE offers one of the largest collections of research papers, including supplementary metadata, to support Artificial Intelligence, Machine Learning research, and engineering projects. This dataset has gained significant attention among major corporations and research laboratories for Natural Language Processing research. Recognizing the importance of accessibility… See the full description on the dataset page: https://huggingface.co/datasets/laion/COREX-18text.translation1 likes210 downloads2y agoHugging Face09laion /tasktrove-bugsinpy-v4-oracle17-apptainer-v1 TaskTrove BugsInPy v4: oracle-verified Apptainer subset v1 This derived Harbor release contains 17 of 479 upstream tasks. Every included reference solution is grounded in the official BugsInPy patch and selected by verifier execution. The remaining tasks are retained in the exclusion ledger; they are not silently discarded. The release targets offline, rootless Apptainer on aarch64. Validation evidence is stored under validation/. Do not describe the full 479-task source as… See the full description on the dataset page: https://huggingface.co/datasets/laion/tasktrove-bugsinpy-v4-oracle17-apptainer-v1.text-generation0 likes57 downloads2d agoHugging Face10laion /Sera-4.5A-Full-T1-v3 laion/Sera-4.5A-Full-T1-v3 Subset of allenai/Sera-4.5A-Full-T1. Size: 72,118 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3.texttext-generation10K<n<100K0 likes53 downloads5mo agoHugging Face11laion /sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13 Nemotron Terminal SFT reproduction evaluation artifacts This repository contains the complete Harbor artifact tree for the 300-trial OpenThoughts-TBLite evaluation of laion/sft-repro-thinking-step630-nemotron-terminal-step1888. The checkpoint was trained from the Grug stage-2 thinking checkpoint on the Nemotron Terminal corpus for 1,888 steps. Result Measure Value Attempted / completed 300 / 300 Verifier-scoreable 259 (86.33%) Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.texttext-generationn<1K0 likes51 downloads1mo agoHugging Face12laion /Sera-4.5A-Full-T1-v3-1000 laion/Sera-4.5A-Full-T1-v3-1000 Subset of allenai/Sera-4.5A-Full-T1. Size: 1,000 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-1000.texttext-generation1K<n<10K1 likes41 downloads5mo agoHugging Face13laion /Sera-4.5A-Full-T1-v3-3160 laion/Sera-4.5A-Full-T1-v3-3160 Subset of allenai/Sera-4.5A-Full-T1. Size: 3,160 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-3160.texttext-generation1K<n<10K0 likes32 downloads5mo agoHugging Face14laion /Sera-4.6-Lite-T2-v4-1000 laion/Sera-4.6-Lite-T2-v4-1000 Row-subset of allenai/Sera-4.6-Lite-T2 (the dataset upstream SERA-8B was trained on), with OpenAI tool_calls pre-rendered into the content string as Hermes/Qwen3-style <tool_call>...</tool_call> wire tokens and tool responses wrapped as <tool_response>...</tool_response>. This mirrors Ai2's sera/datagen/data/postprocess/utils.py::transform_traj_hermes (default tool_call_format: "hermes") which is the missing step between the public Sera-4.6-Lite-T2… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.6-Lite-T2-v4-1000.texttext-generation1K<n<10K0 likes29 downloads5mo agoHugging Face15laion /CoderForge-Preview-v6-1000 laion/CoderForge-Preview-v6-1000 Row-subset of togethercomputer/CoderForge-Preview (trajectories split, filtered_reward1), rendered into Qwen3-compatible think-first OpenHands-XML wire format. Why v6? v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced garbage at eval time (8888..., 0.0.0.0...) despite clean training losses. Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-1000.texttext-generation1K<n<10K0 likes28 downloads5mo agoHugging Face16laion /Sera-4.5A-Full-T1-v3-10000 laion/Sera-4.5A-Full-T1-v3-10000 Subset of allenai/Sera-4.5A-Full-T1. Size: 10,000 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-10000.texttext-generation10K<n<100K0 likes27 downloads5mo agoHugging Face17laion /Sera-4.5A-Full-T1-v3-316 laion/Sera-4.5A-Full-T1-v3-316 Subset of allenai/Sera-4.5A-Full-T1. Size: 316 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag for… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-316.texttext-generationn<1K0 likes26 downloads5mo agoHugging Face18laion /exp_rpt_manybugs-v2 exp_rpt_manybugs-v2 A fixed, non-gameable release of DCAgent/exp_rpt_manybugs: 164 C bug-repair tasks (ManyBugs) packaged as Harbor sandbox tasks. Tasks 164 Projects php (102), libtiff (24), python (15), wireshark (7), lighttpd (9), gzip (5), gmp (2) Unique environments 7 snapshots (ubuntu:22.04 base, one -dev set per project) Format tasks.parquet — columns path (task id), task_binary (gzipped task tarball) Each task contains a single buggy .c file from a… See the full description on the dataset page: https://huggingface.co/datasets/laion/exp_rpt_manybugs-v2.texttext-generationn<1K0 likes25 downloads2mo agoHugging Face19laion /CoderForge-Preview-v3-100000 laion/CoderForge-Preview-v3-100000 Row-subset of the pre-tokenized trajectories in togethercomputer/CoderForge-Preview (trajectories-tokenized_qwencoder subset). Size: 100,000 rows (source: 155,144 across 4 slugs). Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B). Per row columns: input_ids: list[int32] attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's auto-detection of pre-tokenized datasets triggers… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v3-100000.texttext-generation100K<n<1M0 likes23 downloads5mo agoHugging Face20laion /Sera-4.6-Lite-T2-v4-316 laion/Sera-4.6-Lite-T2-v4-316 Row-subset of allenai/Sera-4.6-Lite-T2 (the dataset upstream SERA-8B was trained on), with OpenAI tool_calls pre-rendered into the content string as Hermes/Qwen3-style <tool_call>...</tool_call> wire tokens and tool responses wrapped as <tool_response>...</tool_response>. This mirrors Ai2's sera/datagen/data/postprocess/utils.py::transform_traj_hermes (default tool_call_format: "hermes") which is the missing step between the public Sera-4.6-Lite-T2… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.6-Lite-T2-v4-316.texttext-generationn<1K0 likes23 downloads5mo agoHugging Face21laion /NeurIPSpdf-2024NeurIPS 2024 (paper's pdf) text-generation1 likes21 downloads2y agoHugging Face22laion /Sera-4.5A-Full-T1-v3-31600 laion/Sera-4.5A-Full-T1-v3-31600 Subset of allenai/Sera-4.5A-Full-T1. Size: 31,600 rows (full dataset: 72,118 rows). Format: Raw JSONL, OpenAI-native messages layout. Preserves the original messages field (as JSON string), instance_id, rollout_patch, func_name, func_path, problem_statement, target_patch, docker_image. Adds a source field pointing back to the parent dataset. Each assistant message carries a native tool_calls array (OpenAI tool-calling format) and a train: bool flag… See the full description on the dataset page: https://huggingface.co/datasets/laion/Sera-4.5A-Full-T1-v3-31600.text-generation0 likes21 downloads5mo agoHugging Face23laion /CoderForge-Preview-v3-10000 laion/CoderForge-Preview-v3-10000 Row-subset of the pre-tokenized trajectories in togethercomputer/CoderForge-Preview (trajectories-tokenized_qwencoder subset). Size: 10,000 rows (source: 155,144 across 4 slugs). Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B). Per row columns: input_ids: list[int32] attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's auto-detection of pre-tokenized datasets triggers —… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v3-10000.texttext-generation10K<n<100K0 likes21 downloads5mo agoHugging Face24laion /CoderForge-Preview-v3-31600 laion/CoderForge-Preview-v3-31600 Row-subset of the pre-tokenized trajectories in togethercomputer/CoderForge-Preview (trajectories-tokenized_qwencoder subset). Size: 31,600 rows (source: 155,144 across 4 slugs). Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B). Per row columns: input_ids: list[int32] attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's auto-detection of pre-tokenized datasets triggers —… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v3-31600.texttext-generation10K<n<100K0 likes19 downloads5mo agoHugging Face25laion /CoderForge-Preview-v6-316 laion/CoderForge-Preview-v6-316 Row-subset of togethercomputer/CoderForge-Preview (trajectories split, filtered_reward1), rendered into Qwen3-compatible think-first OpenHands-XML wire format. Why v6? v3 (pre-tokenized) and v5 (wrapper-stripped, no think-block) both produced garbage at eval time (8888..., 0.0.0.0...) despite clean training losses. Root cause: stock Qwen3-8B assigns ~100% prior to <think> as the first token after <|im_start|>assistant. CoderForge's… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v6-316.texttext-generationn<1K0 likes18 downloads5mo agoHugging Face26laion /CoderForge-Preview-v3 laion/CoderForge-Preview-v3 Row-subset of the pre-tokenized trajectories in togethercomputer/CoderForge-Preview (trajectories-tokenized_qwencoder subset). Size: 155,144 rows (source: 155,144 across 4 slugs). Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B). Per row columns: input_ids: list[int32] attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's auto-detection of pre-tokenized datasets triggers —… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v3.texttext-generation100K<n<1M0 likes17 downloads5mo agoHugging Face27laion /CoderForge-Preview-v3-316 laion/CoderForge-Preview-v3-316 Row-subset of the pre-tokenized trajectories in togethercomputer/CoderForge-Preview (trajectories-tokenized_qwencoder subset). Size: 316 rows (source: 155,144 across 4 slugs). Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B). Per row columns: input_ids: list[int32] attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's auto-detection of pre-tokenized datasets triggers —… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v3-316.texttext-generationn<1K0 likes15 downloads5mo agoHugging Face28laion /CoderForge-Preview-v3-3160 laion/CoderForge-Preview-v3-3160 Row-subset of the pre-tokenized trajectories in togethercomputer/CoderForge-Preview (trajectories-tokenized_qwencoder subset). Size: 3,160 rows (source: 155,144 across 4 slugs). Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B). Per row columns: input_ids: list[int32] attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's auto-detection of pre-tokenized datasets triggers —… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v3-3160.texttext-generation1K<n<10K0 likes15 downloads5mo agoHugging Face29laion /CoderForge-Preview-v3-1000 laion/CoderForge-Preview-v3-1000 Row-subset of the pre-tokenized trajectories in togethercomputer/CoderForge-Preview (trajectories-tokenized_qwencoder subset). Size: 1,000 rows (source: 155,144 across 4 slugs). Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B). Per row columns: input_ids: list[int32] attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's auto-detection of pre-tokenized datasets triggers —… See the full description on the dataset page: https://huggingface.co/datasets/laion/CoderForge-Preview-v3-1000.texttext-generation1K<n<10K0 likes14 downloads5mo agoHugging Face30laion /edrXiv-pdfEdArXiv Pdf is a renowned preprint server dedicated to publishing manuscripts in the Education domain. Managed by the Centre of Open Science and a team of well-qualified professionals from prestigious universities, it aims to encourage and promote high-quality research in Education. As part of our open science initiative, we aspired to provide training resources to ignite artificial intelligence research not only in traditional science but also across a broad range of scientific disciplines.… See the full description on the dataset page: https://huggingface.co/datasets/laion/edrXiv-pdf.documenttext-generationn<1K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.