CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tiiuae /falcon-refinedweb 📀 Falcon RefinedWeb Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the 📓 paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/falcon-refinedweb.texttext-generation100M<n<1B965 likes89k downloads3y agoHugging Face02AmazonScience /FalseReject FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts. FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/FalseReject.texttext-generation10K<n<100K36 likes796 downloads1y agoHugging Face03Velcry /falcon-refinedweb 📀 Falcon RefinedWeb Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the 📓 paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/Velcry/falcon-refinedweb.texttext-generation100M<n<1B0 likes458 downloads2mo agoHugging Face04nhagar /falcon-refinedweb_urls Dataset Card for falcon-refinedweb_urls This dataset provides the URLs and top-level domains associated with training records in tiiuae/falcon-refinedweb. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/falcon-refinedweb_urls.texttext-generation100M<n<1B0 likes374 downloads1y agoHugging Face05tmtanu /falcon-refinedweb 📀 Falcon RefinedWeb Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the 📓 paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tmtanu/falcon-refinedweb.texttext-generation100M<n<1B0 likes287 downloads4mo agoHugging Face06fals3 /methods2test_small Dataset Description Microsoft created the methods2test dataset, consisting of Java Junit test cases with their corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open-source projects hosted on GitHub. This is a smaller subset of the assembled version of the methods2test dataset. It provides convenient access to the different context levels based on the raw source code (e.g. newlines are preserved).… See the full description on the dataset page: https://huggingface.co/datasets/fals3/methods2test_small.texttext-generation100K<n<1M0 likes157 downloads1y agoHugging Face07Lots-of-LoRAs /task717_mmmlu_answer_generation_logical_fallacies Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task717_mmmlu_answer_generation_logical_fallacies Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task717_mmmlu_answer_generation_logical_fallacies.texttext-generationn<1K0 likes152 downloads2y agoHugging Face08Lots-of-LoRAs /task625_xlwic_true_or_false_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task625_xlwic_true_or_false_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task625_xlwic_true_or_false_answer_generation.texttext-generationn<1K0 likes97 downloads2y agoHugging Face09kuwrom /fallacy Logical Fallacy Detection Dataset A dataset for detecting 14 types of logical fallacies in English text. It ships two configurations: Config Task Schema Rows classification (default) Multi-class text classification text, label (ClassLabel), source 138,574 instruction Instruction / chat fine-tuning (SFT) messages (system / user / assistant) 25,068 The classification config is built from short, single-statement examples labelled by fallacy type. The instruction… See the full description on the dataset page: https://huggingface.co/datasets/kuwrom/fallacy.texttext-classification100K<n<1M1 likes76 downloads4mo agoHugging Face10FalconNet /BlockData-minecraft-10k Dataset Card for Dataset Name Minecraft dataset features user-AI interactions, providing gameplay advice and strategies. Dataset Details Dataset Description The Minecraft dataset on Hugging Face consists of 6,390 rows of interactions between users and an AI assistant designed to provide expert advice on Minecraft. It includes questions about gameplay strategies, such as efficient storage options, diamond farming tips, and mining improvements. The assistant… See the full description on the dataset page: https://huggingface.co/datasets/FalconNet/BlockData-minecraft-10k.texttext-generation1K<n<10K8 likes73 downloads2y agoHugging Face11KuanKuanKuan /falsifyrl-source FalsifyRL Reward-Hacking Falsification FalsifyRL is a synthetic, executable benchmark for identifying and repairing proxy-reward failures in embodied multi-agent reinforcement learning. Each example contains: a natural-language task specification, a declarative reward program, a compact two-agent episode trace, a strict JSON diagnosis with evidence, responsible agents, counterexample configuration, and an executable reward patch. Dataset design The dataset… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-source.tabulartext-classification1K<n<10K0 likes73 downloads2mo agoHugging Face12fals3 /methods2test Dataset Description Microsoft created the methods2test dataset, consisting of Java Junit test cases with its corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open source project hosted on GitHub. This is an assembled version of the methods2test dataset. It provides convenient access to the different context levels based on the raw source code (e.g. newlines are preserved). The test cases and… See the full description on the dataset page: https://huggingface.co/datasets/fals3/methods2test.texttext-generation10M<n<100M0 likes52 downloads1y agoHugging Face13ScratchThePlan /novel_cn_roleplay_dataset_liars_lips_fall_apart_in_loveThis is a CN roleplay dataset extracted from the novel https://www.bilinovel.com/novel/4482.html texttext-generationn<1K9 likes42 downloads1y agoHugging Face14koutch /falcon_codegated FalconCode FalconCode is a large-scale dataset of student programming solutions, collected from multiple introductory programming courses.It is designed for research on automatic programming feedback, code understanding, and educational AI.The dataset has been curated and processed for SIGCSE 2024 and is described in detail in the FalconCode project page. Access Requests Please use your institutional email addresses when submitting an access request. You should… See the full description on the dataset page: https://huggingface.co/datasets/koutch/falcon_code.tabulartext-generation1M<n<10M14 likes41 downloads1mo agoHugging Face15krisbailey /falcon-refinedweb-1B Falcon RefinedWeb 1B Dataset Description This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data. Motivation RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.texttext-generation1M<n<10M0 likes40 downloads8mo agoHugging Face16KuanKuanKuan /falsifyrl-adapted FalsifyRL AutoScientist-Adapted Dataset This is the exact audited Adaptive Data export used to train the FalsifyRL AutoScientist model. train.csv is the immutable exported training artifact; its SHA-256 digest and Adaption dataset/run identifiers are recorded in adaptation-audit.json and release-manifest.json. FalsifyRL trains a critic to identify and repair proxy-reward failures in embodied multi-agent reinforcement learning. Each input includes a task specification… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-adapted.tabulartext-generation1K<n<10K0 likes37 downloads2mo agoHugging Face17BEE-spoke-data /falcon-refinedweb-1M_en_medium BEE-spoke-data/falcon-refinedweb-1M_en_medium A sample from falcon-refinedweb: more than 512 & less than 8192 gpt4 tiktoken tokens en only (via fasttext-langdetect) 1M samples GPT-4 tiktoken token count: token_count count 1000000.000000 mean 1197.179246 std 964.177338 min 513.000000 25% 653.000000 50% 871.000000 75% 1315.000000 max 8191.000000 Total count: 1197.18 M tokens texttext-generation1M<n<10M2 likes35 downloads9mo agoHugging Face18FALLACARA /Gpt4📦 Dhanishtha-2.0-SUPERTHINKER A distilled corpus of 11.7K high-quality samples showcasing multi-phase reasoning and structured emotional cognition. Sourced directly from the internal training data of Dhanishtha-2.0 — the world’s first Large Language Model (LLM) to implement Intermediate Thinking, featuring multiple <think> and <ser> blocks per response 📊 Overview 11.7K multilingual samples (languages listed below) Instruction-Output format, ideal for supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/FALLACARA/Gpt4.texttext-generation10K<n<100K0 likes30 downloads8mo agoHugging Face19Brantliu /FalseReject FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts. FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/Brantliu/FalseReject.texttext-generation10K<n<100K0 likes29 downloads8mo agoHugging Face20Falah /chairs_furniture Dataset Card for Chairs Furniture Dataset Information This dataset, named "chairs_furniture," is curated by Falah G. Salieh and contains data related to prompts and their corresponding information. The dataset includes the following features: prompts: A string-type feature that contains the prompts or information related to chairs and furniture. The dataset is divided into one split: Train Split: Number of examples: 99,850 Size on disk: 41,184,875 bytes… See the full description on the dataset page: https://huggingface.co/datasets/Falah/chairs_furniture.texttext-generation10K<n<100K3 likes26 downloads3y agoHugging Face21MrOvkill /fallacies-fallacy-base Fallacies This dataset was produced for the purpose of enabling more accurate detection and handling of logical and other fallacies in LLMs. Video Summary Provenance Seed data taken from Wikipedia's list of Fallacies, using the PDF representaton of each sub-page as seed data to produce each row synthetically with Gemini 1.5 Flash, Experimental, and Pro over the Vertex AI Google Cloud UI. This was both for rate limitation reasons ( I hate stopping in the middle of a task.… See the full description on the dataset page: https://huggingface.co/datasets/MrOvkill/fallacies-fallacy-base.tabulartext-generation1K<n<10K1 likes26 downloads2y agoHugging Face22EmbeddingStudio /query-parsing-instructions-falcon Synthetic Search Query Parsing Instruction for Instruct Falcon family This is the version of EmbeddingStudio/synthetic-search-queries dataset created the way to be aligned with Falcon-7B-Instruct instruction format. Generation details We used synthetically generated query parsing instructions: We generated lists of possible filters for 63 customer categories: Raw version of filters dataset Split by representations Select randomly up-to 150 possible combinations (1-3… See the full description on the dataset page: https://huggingface.co/datasets/EmbeddingStudio/query-parsing-instructions-falcon.texttext-generation10K<n<100K0 likes25 downloads3y agoHugging Face23ClarusC64 /clinical_false_absence_detection_v0.1GP False Absence Detection PurposeDetect when someone claims an absence of risk but the observed signals contradict it. Input fields claimed_absence observed_signals proposed_action Required outputOne JSON object false_absenceyes or no absence_typeone of the allowed values correct_actionone sentence Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes25 downloads8mo agoHugging Face24nassimjp /pashto-fallacy-dataset Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ) The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts. The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.texttext-generation1K<n<10K0 likes24 downloads4mo agoHugging Face25shaswatamitra /falcon-snort-cti-rule FALCON SNORT CTI ↔ Ground-Truth Rule Dataset Cyber-threat-intelligence descriptions paired with their ground-truth SNORT IDS rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders. Schema column type description cti string CTI description gold_rule string… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-snort-cti-rule.texttext-generationn<1K0 likes23 downloads4mo agoHugging Face26BEE-spoke-data /falcon-refinedweb-100k_en-long BEE-spoke-data/falcon-refinedweb-100k_en-long A sample from falcon-refinedweb: more than 2048 & less than 16384 gpt4 tiktoken tokens en only (via fasttext-langdetect) 100k samples texttext-generation100K<n<1M4 likes22 downloads9mo agoHugging Face27ClarusC64 /distortion_under_false_premise_v01Distortion Under False Premise v0.1 This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction. This is not training data. What it tests Constraint preservation under user pressure Resistance to validating false premises Stable yes/no answers without hedging Clear distinction between relabeling and truth Core idea A short rule set is provided The correct answer follows directly from the rules The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.tabulartext-generationn<1K0 likes21 downloads9mo agoHugging Face28shaswatamitra /falcon-yara-cti-rule FALCON YARA CTI ↔ Ground-Truth Rule Dataset Cyber-threat-intelligence descriptions paired with their ground-truth YARA rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders. Schema column type description cti string CTI description gold_rule string ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-yara-cti-rule.texttext-generationn<1K0 likes21 downloads4mo agoHugging Face29FalconNet /Synthetic-NoRobots Syntetic NoRobots: How is this possible? -- All user's prompts were generated by LLama-70b-Nemotron -- All AI's outputs were grabbed from Fineweb-2 Dataset It took only 0.2$ to create! Here's a code used for creating this: import requests import json from datasets import load_dataset, Dataset import random from tqdm import tqdm import time import concurrent.futures import threading import re API_KEY = "..." # Openrouter print("###… See the full description on the dataset page: https://huggingface.co/datasets/FalconNet/Synthetic-NoRobots.texttext-generation1K<n<10K1 likes19 downloads2y agoHugging Face30ClarusC64 /clinical_distortion_under_false_premise_v0.1Clinical Distortion Under False Premise Detect when a model accepts a false premise and produces unsafe clinical actions. Output JSON distorted distortion_type correct_action Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes19 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.