CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adameubanks /filtered_articles_by_year Dataset Card for Filtered Articles by Year Dataset Summary The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time. Supported Tasks and Leaderboards This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.texttext-generation10M<n<100M1 likes2.6k downloads1y agoHugging Face02adamraudonis /DoorBench DoorBench 1000 procedural articulated doors for robot simulation, with MJCF, URDF and USD exports, Blender appearances, and reference motion. Interactive catalogue · Source and tools · Release guide Version v2026.09.05 contains 1014 saved Blender images, including 14 images rendered with the higher sample preset, and reference clips/native trajectories for 1000 doors. The release manifest, complete per-file SHA256 inventory, source hashes and immutable Hub revision identify the… See the full description on the dataset page: https://huggingface.co/datasets/adamraudonis/DoorBench.tabularother1K<n<10K2 likes431 downloads19d agoHugging Face03adamo1139 /temp_poziomka_sft2text100K<n<1M1 likes148 downloads11mo agoHugging Face04adampippert /granite-decisions-synthetic Granite Decisions synthetic datasets Original, deterministic English fixtures for Adam Pippert's personal Granite Decisions project. The original default config has 162 examples: 54 train, 54 calibration, and 54 test. These exercise the pipeline; they are not a representative quality benchmark. Source and license The source is the project's original template generator, published here as make_smoke_data.py, from release v0.1.0, commit… See the full description on the dataset page: https://huggingface.co/datasets/adampippert/granite-decisions-synthetic.text10K<n<100K0 likes140 downloads6d agoHugging Face05AdamCodd /unreal-engine-5-codeProcessed dataset from AdamCodd/unreal-engine-5-raw focused on the code. If you want to support me, you can here. text100K<n<1M13 likes118 downloads2y agoHugging Face06adamcnoonan /love-repro-aigve60k LOVE / AIGVE-60K — reproduction results Raw per-video outputs, subset manifests, and metrics from an independent reproduction of LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation (ICML 2026 submission #1055, OpenReview P6fWeIVbwb). Every model number here comes from running the authors' own released 9B checkpoints through the authors' own evaluation code on 1× A100-80GB. 💻 Code + full writeup… See the full description on the dataset page: https://huggingface.co/datasets/adamcnoonan/love-repro-aigve60k.text1K<n<10K0 likes97 downloads2mo agoHugging Face07adamm-hf /Fable-5-Max-Reasoning-Filtered-250x Dataset Description This dataset contains 25. highly detailed architectural traces mapping out security implementations for hybrid global banking systems encompassing both fiat and cryptocurrency infrastructures. This is 10,000,000+ estimated tokens of fable 5 data, filtered and classified to remove low-quality entries by qwen 2.5 7B, and improved by GLM 5.2. The dataset bypasses basic conversational filler and is engineered to advance the domain precision, strict formatting… See the full description on the dataset page: https://huggingface.co/datasets/adamm-hf/Fable-5-Max-Reasoning-Filtered-250x.texttext-generationn<1K3 likes94 downloads1mo agoHugging Face08AdamLeung /twitter_suicidal_risk Twitter Suicide Risk Level Dataset Short English tweets paired with a 0–4 suicide risk label, used for fine-tuning and evaluating risk-level classification. This directory holds the final splits: train.jsonl / val.jsonl / test.jsonl. Files and size File Rows Share train.jsonl 7006 80% val.jsonl 875 10% test.jsonl 875 10% Total 8756 100% Fields JSONL, one sample per line, three fields only: Field Type Description id… See the full description on the dataset page: https://huggingface.co/datasets/AdamLeung/twitter_suicidal_risk.texttext-classification1K<n<10K0 likes92 downloads1mo agoHugging Face09Adam1010 /cgrt-consensus-5model CGRT Consensus 5-Model Dataset Multi-model consensus dataset for studying model agreement and disagreement patterns on mathematical reasoning tasks. Dataset Description 61,678 math problems evaluated by 5 frontier LLMs with full reasoning traces and extracted answers. Models Used Model Provider Version Claude Anthropic claude-3-5-sonnet-20241022 Codex/GPT-4 OpenAI gpt-4o Gemini Google gemini-1.5-flash DeepSeek DeepSeek deepseek-chat Qwen… See the full description on the dataset page: https://huggingface.co/datasets/Adam1010/cgrt-consensus-5model.tabularquestion-answering10K<n<100K0 likes74 downloads9mo agoHugging Face10AdamLucek /youtube-titles Youtube Title & Descriptions Dataset About 4941 videos across 50 YouTube Channels List of sampled channels here Splits: Train: 4199 Validation: 493 Test: 249 Data was shuffled and sampled evenly from all channels to create splits. Additionally, has a column ready to go for gemma-2-9b-it fine tuning formatting! Potentially more model formats to come. About the Data: Label Description channel_name The… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/youtube-titles.texttext-generation1K<n<10K1 likes73 downloads2y agoHugging Face11hollywoodfrancis /adam_pretraintext1M<n<10M0 likes73 downloads10mo agoHugging Face12AdamCodd /emotion-balanced Dataset Card for "emotion" Dataset Summary Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances An example looks as follows. { "text": "im feeling quite sad… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/emotion-balanced.texttext-classification10K<n<100K5 likes69 downloads3y agoHugging Face13AdamCodd /no_robots-alpaca No Robots: Alpaca edition This dataset is a cleaned (missing/extra spaces...) and reformatted version of the No Robots dataset from HuggingFaceH4, adapted to conform with the Alpaca instruction set. Notably, it diverges from the original dataset in the way the 'Chat' category is handled; it has been decomposed into single-turn conversations to align with Alpaca's limitations regarding multi-turn interactions. The dataset's IDs have been generated using the SHA256 algorithm.… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/no_robots-alpaca.texttext-generation10K<n<100K7 likes67 downloads2y agoHugging Face14mzio /aprm-sft-thoughts-tau2-retail-policy_best-adamw30-lp0 Act-PRM SFT thoughts — tau2-bench retail Act-PRM (Action Process Reward Models) infers the latent thoughts behind logged, action-only agent demonstrations via an offline EM. For each logged action x in state s we sample G=4 candidate thoughts z, score each by the length-penalized action likelihood reward(z) = p(x | s, z) (len_frac grows with the thought's token length), and mark the best thought (argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-retail-policy_best-adamw30-lp0.tabulartext-generationn<1K0 likes65 downloads27d agoHugging Face15AdamMuhtar /NOAA-HRRR-CONUS-ImageCaptionimagen<1K0 likes63 downloads1y agoHugging Face16adamo1139 /AEZAKMI_v2 New things changed in v2 added some samples from spicy 3.1 removed samples that were sending a shiver down RPers spines! removed some more GPTslop like "Remember to always..." added some jokes and duplicated some samples with two system prompts - "A chat." and "A chat with uncensored assistant.", so that hopefully model connects those two and act more freely. New things 2023-02-01 moved sharegpt version to a different repo to make it easier to use. New things… See the full description on the dataset page: https://huggingface.co/datasets/adamo1139/AEZAKMI_v2.text10K<n<100K5 likes57 downloads3y agoHugging Face17mzio /aprm-sft-thoughts-snorkel-insurance-policy_best-adamw30-lp0 Act-PRM SFT thoughts — snorkel-insurance insurance Act-PRM (Action Process Reward Models) infers the latent thoughts behind logged, action-only agent demonstrations via an offline EM. For each logged action x in state s we sample G=4 candidate thoughts z, score each by the length-penalized action likelihood reward(z) = p(x | s, z) (len_frac grows with the thought's token length), and mark the best thought (argmax reward). The (thought + action) span is then what downstream SFT… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-insurance-policy_best-adamw30-lp0.tabulartext-generation1K<n<10K0 likes57 downloads26d agoHugging Face18mzio /aprm-sft-thoughts-tau2-airline-policy_best-adamw30-lp0 Act-PRM SFT thoughts — tau2-bench airline Act-PRM (Action Process Reward Models) infers the latent thoughts behind logged, action-only agent demonstrations via an offline EM. For each logged action x in state s we sample G=4 candidate thoughts z, score each by the length-penalized action likelihood reward(z) = p(x | s, z) (len_frac grows with the thought's token length), and mark the best thought (argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-airline-policy_best-adamw30-lp0.tabulartext-generationn<1K0 likes56 downloads27d agoHugging Face19adamabuhamdan /startup-advisor-dataset 🚀 Startup Advisor Dataset A high-quality instruction-following dataset distilled from 8 foundational business and startup books, structured as actionable advice with real-world 2025 examples. Designed for fine-tuning large language models (e.g., Qwen, LLaMA, Mistral) to become expert startup advisors. 📖 Dataset Summary Property Value Total Entries 1,564 Format JSONL — ChatML (messages array) Language English License CreativeML OpenRAIL-M Avg. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/adamabuhamdan/startup-advisor-dataset.texttext-generation1K<n<10K2 likes53 downloads5mo agoHugging Face20adamrotmil /claudish-pairs Claudish Pairs The first open parallel corpus of English ↔ Claudish — the characteristic prose style of Claude and Claude Code. 10,227 pairs, each an English text and its Claudish restyling, authored and quality-controlled for faithfulness. This is the v3 training set of adamrotmil/claudish-style-adapter; pipeline code at github.com/adamrotmil/claudish-style-adapter. Fields Field Meaning english source text (plain English) claudish the restyling… See the full description on the dataset page: https://huggingface.co/datasets/adamrotmil/claudish-pairs.texttranslation10K<n<100K0 likes53 downloads1mo agoHugging Face21adamallcock /mmmu-pro-clean MMMU-Pro-Clean — exclusion overlay A corrected drop-in for MMMU-Pro (standard 10-option split): 1,730 → 1,526 items, with 204 broken items removed. ⚠️ Overlay, not a rehost. MMMU-Pro is Apache-2.0, but its images come from exams/textbooks and carry third-party copyright, so this repo does NOT host the data or images. It ships the exclusion manifest — IDs, categories, tiers, the official-answer letter, coded reasons — which you apply to your own licensed MMMU/MMMU_Pro download.… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/mmmu-pro-clean.textvisual-question-answeringn<1K1 likes52 downloads2mo agoHugging Face22adamallcock /gpqa-extended-clean GPQA-Extended-Clean — exclusion overlay A corrected drop-in for the full 546-item GPQA-Extended: 546 → 498 items, with 48 broken items removed. ⚠️ Overlay, not a rehost. Per GPQA's anti-contamination norm, this repo ships the exclusion manifest — IDs, categories, tiers, coded reasons only — never item text. Apply it to your own licensed GPQA-Extended download. 📄 Paper: When the Answer Key Is Wrong — Allcock 2026 (arXiv forthcoming) · 💻 Loader + gated evidence:… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/gpqa-extended-clean.textquestion-answeringn<1K1 likes52 downloads2mo agoHugging Face23AdamLucek /apple-environmental-report-QA-retrieval Apple's 2024 Environmental Report QA Pairs 4300 question and relevant text chunks made from Apple's 2024 Environmental Report. Chunking was done with a token based recursive chunker at 800 token chunk size with a 400 token overlap resulting in 215 chunks. 20 question labels per chunk were synthetically generated using gpt-4o-mini with the attached prompt and a temperature of 1.0. Entries were shuffled and split into an 80/20 Train/Validation split resulting in:Training set size:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/apple-environmental-report-QA-retrieval.textquestion-answering1K<n<10K0 likes51 downloads2y agoHugging Face24AdamiTitus /pii-masking-300k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/AdamiTitus/pii-masking-300k.texttext-classification100K<n<1M1 likes50 downloads8mo agoHugging Face25adamallcock /gpqa-diamond-clean GPQA-Diamond-Clean — exclusion overlay A corrected drop-in for GPQA Diamond: 198 → 189 items, with 9 broken items removed. ⚠️ This is an overlay, not a rehost. Per GPQA's anti-contamination norm (no crawlable plaintext), this repo ships the exclusion manifest — item IDs, categories, tiers, and coded reasons only — never GPQA item text or answer values. You apply it to your own licensed GPQA download. 📄 Paper: When the Answer Key Is Wrong — Allcock 2026 (arXiv forthcoming) ·… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/gpqa-diamond-clean.textquestion-answeringn<1K1 likes49 downloads2mo agoHugging Face26mzio /aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0 Act-PRM SFT thoughts — snorkel-finance finance Act-PRM (Action Process Reward Models) infers the latent thoughts behind logged, action-only agent demonstrations via an offline EM. For each logged action x in state s we sample G=4 candidate thoughts z, score each by the length-penalized action likelihood reward(z) = p(x | s, z) (len_frac grows with the thought's token length), and mark the best thought (argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0.tabulartext-generation1K<n<10K0 likes48 downloads26d agoHugging Face27adamo1139 /powershell_thestacktabular100K<n<1M1 likes46 downloads2y agoHugging Face28adamallcock /gpqa-ext-complement-clean GPQA-Extended-Complement-Clean — exclusion overlay A corrected drop-in for the 348 GPQA-Extended items disjoint from Diamond: 348 → 309 items, with 39 broken items removed. ⚠️ Scope — read first. This is NOT canonical GPQA-Extended. Canonical Extended is the 546-item superset that includes Diamond. This release covers only the 348-item complement (Extended minus Diamond). Applying these exclusions to the full 546-item split, or treating 348 as "Extended", silently evaluates a… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/gpqa-ext-complement-clean.textquestion-answeringn<1K1 likes46 downloads2mo agoHugging Face29mzio /aprm-sft-thoughts-tau2-airline-base_best-adamw30-lp0 Act-PRM SFT thoughts — tau2-bench airline Act-PRM (Action Process Reward Models) infers the latent thoughts behind logged, action-only agent demonstrations via an offline EM. For each logged action x in state s we sample G=4 candidate thoughts z, score each by the length-penalized action likelihood reward(z) = p(x | s, z) (len_frac grows with the thought's token length), and mark the best thought (argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-airline-base_best-adamw30-lp0.tabulartext-generationn<1K0 likes44 downloads19d agoHugging Face30adamo1139 /toxic-dpo-natural-v5I mixed in toxid-dpo-natural-v4 and rawrr v2-1 stage 2 with chosen field from original no_robots and got myself toxic-dpo-natural-v5. Goal is to avoid overfitting via DPO to a specific type of instruct, and instead just DPO the model to be more open to answering and also answer like a human being. We'll see whether this works.I trained Yi 34B with this dataset and ORPO, it does work very nicely so far! text1K<n<10K13 likes43 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.