CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexshpunt /explicit-edit-benchmark Explicit Edit Benchmark 226 deterministic exact-edit tasks, run by different agents, harnesses, models and configurations. Every observation records what the harness did and whether the resulting files matched byte for byte. Source code and benchmark runner: GitHub — Explicit Edit Benchmark Open the interactive Explorer to compare agents, harnesses, models, versions, reasoning modes, correctness, recovery, time, cost and tokens. Leaderboard by model route Score v2… See the full description on the dataset page: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark.tabulartext-generationn<1K2 likes9k downloads4d agoHugging Face02HiTZ /casimedicos-exp Antidote CasiMedicos Dataset - Possible Answers Explanations in Resident Medical Exams We present a new multilingual parallel medical dataset of commented medical exams which includes not only explanatory arguments for the correct answer but also arguments to explain why the remaining possible answers are incorrect. This dataset can be used for various NLP tasks including: Medical Question Answering, Explanatory Argument Extraction or Explanation Generation. The… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/casimedicos-exp.tabulartext-generation1K<n<10K4 likes2k downloads3y agoHugging Face03eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes911 downloads4mo agoHugging Face04placeholderlabs /exp-pool-repository-code-dolma2-tokenized Locus EXP Repository Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-repository-code-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes736 downloads1mo agoHugging Face05eczech /marinfold-exp11-protein-docs marinfold-exp11-pdocs Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from timodonnell/protein-docs, partitioned by the source round column: Config Source rounds Approx rows high round 0 ~1.68M medium round 1 ~1.42M low round 2–4 ~2.29M Train/val/test split assignment is inherited from the source dataset (leakage-resistant structural-cluster hashing). All columns from the source are preserved; rows are simply partitioned by round. See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.tabulartext-generation1M<n<10M0 likes576 downloads4mo agoHugging Face06tankalapavankalyan /exp01-eeg-to-text-sentences Exp01 — Sentence-level EEG-to-text training data (unified) This is a private working corpus for experiment 1 (fine-tuning EEG / time-series foundation models on EEG-to-English-text). It bundles several public EEG-while-reading datasets into a single, raw-lossless parquet schema where one row = one sentence read by one participant. ⚠️ License: Per-source licenses are preserved verbatim in each row's license column and source_url. Do not re-distribute publicly without re-checking the… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/exp01-eeg-to-text-sentences.tabulartext-generation10K<n<100K0 likes350 downloads5mo agoHugging Face07placeholderlabs /exp-pool-commit-code-dolma2-tokenized Locus EXP Commit Code - Dolma 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes305 downloads1mo agoHugging Face08JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes209 downloads9d agoHugging Face09placeholderlabs /exp-pool-academic-dolma2-tokenized Locus EXP Academic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-academic-dolma2-tokenized.tabulartext-generation100K<n<1M0 likes196 downloads1mo agoHugging Face10ExponentialScience /DLT-Tweets DLT-Tweets [Paper] • [Code] Dataset Description Dataset Summary DLT-Tweets is a large-scale corpus of social media posts related to Distributed Ledger Technology (DLT). This dataset is part of the larger DLT-Corpus collection, designed to support NLP research, social computing studies, and public discourse analysis in the DLT domain. It was introduced in the paper DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain.… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/DLT-Tweets.tabulartext-generation10M<n<100M0 likes195 downloads7mo agoHugging Face11b-mc2 /cli-commands-explained Overview This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.tabulartext-generation10K<n<100K5 likes151 downloads2y agoHugging Face12placeholderlabs /exp-pool-olmo-web-dolma2-tokenized Locus EXP OLMo Web - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-olmo-web-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes149 downloads1mo agoHugging Face13placeholderlabs /exp-pool-encyclopedic-dolma2-tokenized Locus EXP Encyclopedic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.tabulartext-generation1M<n<10M0 likes136 downloads1mo agoHugging Face14placeholderlabs /exp-pool-nemotron-math-dolma2-tokenized Locus EXP Nemotron Math - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes123 downloads1mo agoHugging Face15shunanhe /NPM-Artifact-Explanation-Benchmark NPM-Artifact-Explanation-Benchmark English NPM-Artifact-Explanation-Benchmark is a cross-category multimodal corpus and benchmark resource for Chinese cultural artifact understanding and explanation. This release contains 28,826 cleaned artifact records derived from National Palace Museum source records' opendata (https://digitalarchive.npm.gov.tw/opendata/). Each record includes structured artifact metadata, image URLs, source record URLs, and human-written… See the full description on the dataset page: https://huggingface.co/datasets/shunanhe/NPM-Artifact-Explanation-Benchmark.tabularimage-to-text10K<n<100K1 likes123 downloads9d agoHugging Face16YYYYYYibo /alfworld-expert-prefix-rollouts ALFWorld Expert-Prefix Rollout Landscape This dataset measures how a frozen language-model actor's probability of solving an ALFWorld task changes after replaying different-length prefixes of a successful expert trajectory. The collection contains all 3,553 ALFWorld training tasks from the Agent-G2 SFT data. Eight independent actor rollouts were sampled from the initial state for every task. For the 2,307 low-signal tasks with at most one root success, eight additional… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/alfworld-expert-prefix-rollouts.tabulartext-generation100K<n<1M0 likes89 downloads22d agoHugging Face17YYYYYYibo /alfworld-experimenter-gpt5mini-sft-1k ALFWorld Experimenter GPT-5 mini SFT 1K This dataset contains 1,000 blind GPT-5 mini reasoning demonstrations for an ALFWorld expert-prefix selection task. The intended use is to give a 7B experimenter model a structured reasoning warm start before reinforcement learning, not to treat GPT-5 mini's selected depths as ground-truth labels. Task For each ALFWorld task, the experimenter receives eight failed trajectories from a frozen Qwen2.5-7B-Instruct actor and one… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/alfworld-experimenter-gpt5mini-sft-1k.tabulartext-generation1K<n<10K0 likes85 downloads19d agoHugging Face18JackHsieh /spam.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained spam.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A length-matched null control for JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids, stored directly in the kv-tags-explained training format (there is no separate base repo). Every thought is this one sentence, repeated: We are thinking hard about what comes in the next 8 tokens by reasoning correctly and carefully about what comes before it in the document. It is fluent, on-topic and… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/spam.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes81 downloads24d agoHugging Face19jash404 /emergent-misalignment-experiment-1-data Emergent Misalignment Experiment 1 Data Artifacts Curated SFT data and diagnostics for an awareness-stratified code experiment on emergent misalignment. This artifact contains the exact trainable JSONL branches used for the reported n=1000 and n=3452 runs, plus the small manifests and balance summaries needed to audit the data mixture. The paired model adapters are available at jash404/emergent-misalignment-experiment-1-adapters. The source code and reports are in… See the full description on the dataset page: https://huggingface.co/datasets/jash404/emergent-misalignment-experiment-1-data.tabulartext-generationn<1K0 likes80 downloads4mo agoHugging Face20mencosk /gomodel-go-expert-v4 GoModel Go Expert v4 Dataset Description A high-quality dataset for fine-tuning Qwen2.5-Coder-7B to be an expert Go software engineer with tool-calling capabilities. This is version 4, substantially rebuilt from v3 with: Structured messages format (not pre-rendered ChatML text) Go AST-extracted code from real repositories using go/parser Go 1.26 feature coverage (February 2026 release) Senior/staff-level engineering content (architecture, distributed systems, API… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v4.tabulartext-generation10K<n<100K0 likes76 downloads2mo agoHugging Face21JackHsieh /4B-Instruct-distill-3rkrz2vo.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-distill-3rkrz2vo.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-Instruct-distill-3rkrz2vo.stride-1.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: {last 8 prefix tokens} VALUE: {thought_text} <|/note|> and stored both as text… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-distill-3rkrz2vo.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes69 downloads24d agoHugging Face22expertdata-factory /cybersecurity-reasoning-cot-v1 🛡️ Expert Cybersecurity Reasoning Dataset (CoT) This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic. 💎 Key Highlights Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets. Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.tabulartext-generationn<1K2 likes63 downloads7mo agoHugging Face23almador2002 /tripalchemy-experiences 🧪 TripAlchemy — Synthetic Travel Experiences 10,396 rich, vibe-scored travel experiences across 30 cities — generated by a pre-trained Hugging Face model and served through a live recommender app. 🚀 Live demo: huggingface.co/spaces/almador2002/tripalchemy ✨ What makes it special Every experience is scored 0–1 across all six categories at once — 🍽️ culinary, 🏛️ historical, 🛍️ shopping, 🌲 nature, 🌃 nightlife, 🎨 art & culture. That multi-label… See the full description on the dataset page: https://huggingface.co/datasets/almador2002/tripalchemy-experiences.tabulartext-generation10K<n<100K0 likes57 downloads2mo agoHugging Face24leonli66 /stage3-real-expansion-agent-teacher-separated-pilot Teacher-Separated Expansion Agent Pilot A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks. The teacher-only trajectory-generation system prompt is recorded in metadata/generation-manifest.json for auditability, but is absent from every saved training trajectory. Each final messages list begins with the real memory-wrapped task user message, followed by native assistant expand calls, exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.tabularquestion-answeringn<1K0 likes57 downloads23d agoHugging Face25FamilyLinks /prompts-export-dataset 🔥 Prometheus Prompts The Definitive Prompt Engineering Corpus v0.1 "Just as Prometheus stole fire from the gods to empower humanity, this corpus steals the spark of perfect prompting to ignite the next generation of AI." By FAMILY LINK 📊 Dataset Stats 1,347,933 Prompts • 54,743 Topics • 1.43GB • 1147 Char Avg 100% Human-Reviewed • v0.1 • Educational License 🎖️ Featured Sample Prompts 🏃‍♂️ Running Training & Genetics… See the full description on the dataset page: https://huggingface.co/datasets/FamilyLinks/prompts-export-dataset.tabulartext-generation1M<n<10M3 likes56 downloads10mo agoHugging Face26JackHsieh /luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: {last 8 prefix tokens} VALUE: {thought_text} <|/note|> and stored both as text… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation100K<n<1M0 likes56 downloads24d agoHugging Face27AmanPriyanshu /tool-reasoning-sft-RESEARCH-explorations Explorations Trajectories — Cleaned & Stripped 149,025 multi-turn code exploration agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Origin Derived from AmanPriyanshu/random-small-github-repositories and AmanPriyanshu/random-python-github-repositories. Each trajectory is a search session where an agent navigates a GitHub repository using terminal commands to locate a target file. The agent reasons about project… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-explorations.tabulartext-generation100K<n<1M1 likes55 downloads6mo agoHugging Face28leonli66 /stage3-real-expansion-agent Stage 3 Real-Source Expansion Agents — Pilot This inspection pilot converts pinned training examples from real legal, financial, biomedical, and grounded-QA corpora into native selective-expansion traces. It is not the final-scale mixture. Each row contains eight positional seg_i blocks. Every initial segment holds 512–896 words of real source material wrapped in <|memory_start|>...<|memory_end|>. Qwen3-235B-A22B-Instruct-2507 receives a native expand({"segment_id": "seg_i"})… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent.tabularquestion-answeringn<1K0 likes54 downloads29d agoHugging Face29JackHsieh /32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: {last 8 prefix tokens} VALUE: {thought_text} <|/note|> and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes51 downloads24d agoHugging Face30MichaelAnthony /lemonseed-codex-cogen-expansion lemonseed-codex-cogen-expansion LemonSeed — Codex-teacher expansion data, reviewed (v2). Contents intelligent_codex_expansion_reviewed_v2.jsonl (336 rows) Format JSON Lines (.jsonl), one example per line. Provenance LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning. tabulartext-generationn<1K0 likes51 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.