datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
falcon-refinedweb
📀 Falcon RefinedWeb
Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license.
See the 📓 paper on arXiv for more details.
RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data.
RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tiiuae/falcon-refinedweb.FalseReject
FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models
FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts.
FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/FalseReject.falcon-refinedweb
📀 Falcon RefinedWeb
Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license.
See the 📓 paper on arXiv for more details.
RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data.
RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/Velcry/falcon-refinedweb.falcon-refinedweb_urls
Dataset Card for falcon-refinedweb_urls
This dataset provides the URLs and top-level domains associated with training records in tiiuae/falcon-refinedweb. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers.… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/falcon-refinedweb_urls.falcon-refinedweb
📀 Falcon RefinedWeb
Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license.
See the 📓 paper on arXiv for more details.
RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data.
RefinedWeb is also "multimodal-friendly": it contains links and… See the full description on the dataset page: https://huggingface.co/datasets/tmtanu/falcon-refinedweb.methods2test_small
Dataset Description
Microsoft created the methods2test dataset, consisting of Java Junit test cases with their corresponding focal methods.
It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open-source projects hosted on GitHub.
This is a smaller subset of the assembled version of the methods2test dataset.
It provides convenient access to the different context levels based on the raw source code (e.g. newlines are preserved).… See the full description on the dataset page: https://huggingface.co/datasets/fals3/methods2test_small.task717_mmmlu_answer_generation_logical_fallacies
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task717_mmmlu_answer_generation_logical_fallacies
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task717_mmmlu_answer_generation_logical_fallacies.task625_xlwic_true_or_false_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task625_xlwic_true_or_false_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task625_xlwic_true_or_false_answer_generation.fallacy
Logical Fallacy Detection Dataset
A dataset for detecting 14 types of logical fallacies in English text. It ships
two configurations:
Config
Task
Schema
Rows
classification (default)
Multi-class text classification
text, label (ClassLabel), source
138,574
instruction
Instruction / chat fine-tuning (SFT)
messages (system / user / assistant)
25,068
The classification config is built from short, single-statement examples labelled
by fallacy type. The instruction… See the full description on the dataset page: https://huggingface.co/datasets/kuwrom/fallacy.BlockData-minecraft-10k
Dataset Card for Dataset Name
Minecraft dataset features user-AI interactions, providing gameplay advice and strategies.
Dataset Details
Dataset Description
The Minecraft dataset on Hugging Face consists of 6,390 rows of interactions between users and an AI assistant designed to provide expert advice on Minecraft. It includes questions about gameplay strategies, such as efficient storage options, diamond farming tips, and mining improvements. The assistant… See the full description on the dataset page: https://huggingface.co/datasets/FalconNet/BlockData-minecraft-10k.falsifyrl-source
FalsifyRL Reward-Hacking Falsification
FalsifyRL is a synthetic, executable benchmark for identifying and repairing proxy-reward failures
in embodied multi-agent reinforcement learning.
Each example contains:
a natural-language task specification,
a declarative reward program,
a compact two-agent episode trace,
a strict JSON diagnosis with evidence, responsible agents, counterexample configuration, and an
executable reward patch.
Dataset design
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-source.methods2test
Dataset Description
Microsoft created the methods2test dataset, consisting of Java Junit test cases with its corresponding focal methods.
It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K
Java open source project hosted on GitHub.
This is an assembled version of the methods2test dataset. It provides convenient access to the different context levels based on the raw source code (e.g. newlines are preserved). The test cases and… See the full description on the dataset page: https://huggingface.co/datasets/fals3/methods2test.novel_cn_roleplay_dataset_liars_lips_fall_apart_in_loveThis is a CN roleplay dataset extracted from the novel https://www.bilinovel.com/novel/4482.html
falcon_code
FalconCode
FalconCode is a large-scale dataset of student programming solutions, collected from multiple introductory programming courses.It is designed for research on automatic programming feedback, code understanding, and educational AI.The dataset has been curated and processed for SIGCSE 2024 and is described in detail in the FalconCode project page.
Access Requests
Please use your institutional email addresses when submitting an access request.
You should… See the full description on the dataset page: https://huggingface.co/datasets/koutch/falcon_code.falcon-refinedweb-1B
Falcon RefinedWeb 1B
Dataset Description
This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data.
Motivation
RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.falsifyrl-adapted
FalsifyRL AutoScientist-Adapted Dataset
This is the exact audited Adaptive Data export used to train the FalsifyRL AutoScientist model.
train.csv is the immutable exported training artifact; its SHA-256 digest and Adaption dataset/run
identifiers are recorded in adaptation-audit.json and release-manifest.json.
FalsifyRL trains a critic to identify and repair proxy-reward failures in embodied multi-agent
reinforcement learning. Each input includes a task specification… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-adapted.falcon-refinedweb-1M_en_medium
BEE-spoke-data/falcon-refinedweb-1M_en_medium
A sample from falcon-refinedweb:
more than 512 & less than 8192 gpt4 tiktoken tokens
en only (via fasttext-langdetect)
1M samples
GPT-4 tiktoken token count:
token_count
count 1000000.000000
mean 1197.179246
std 964.177338
min 513.000000
25% 653.000000
50% 871.000000
75% 1315.000000
max 8191.000000
Total count: 1197.18 M tokens
Gpt4📦 Dhanishtha-2.0-SUPERTHINKER
A distilled corpus of 11.7K high-quality samples showcasing multi-phase reasoning and structured emotional cognition. Sourced directly from the internal training data of Dhanishtha-2.0 — the world’s first Large Language Model (LLM) to implement Intermediate Thinking, featuring multiple <think> and <ser> blocks per response
📊 Overview
11.7K multilingual samples (languages listed below)
Instruction-Output format, ideal for supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/FALLACARA/Gpt4.FalseReject
FalseReject: A Dataset for Over-Refusal Mitigation in Large Language Models
FalseReject is a large-scale dataset designed to mitigate over-refusal behavior in large language models (LLMs)—the tendency to reject safe prompts that merely appear sensitive. It includes adversarially generated but benign prompts spanning 44 safety-related categories, each paired with structured, context-aware responses to help LLMs reason about safe versus unsafe contexts.
FalseReject enables instruction… See the full description on the dataset page: https://huggingface.co/datasets/Brantliu/FalseReject.chairs_furniture
Dataset Card for Chairs Furniture
Dataset Information
This dataset, named "chairs_furniture," is curated by Falah G. Salieh and contains data related to prompts and their corresponding information. The dataset includes the following features:
prompts: A string-type feature that contains the prompts or information related to chairs and furniture.
The dataset is divided into one split:
Train Split:
Number of examples: 99,850
Size on disk: 41,184,875 bytes… See the full description on the dataset page: https://huggingface.co/datasets/Falah/chairs_furniture.fallacies-fallacy-base
Fallacies
This dataset was produced for the purpose of enabling more accurate detection and handling of logical and other fallacies in LLMs. Video Summary
Provenance
Seed data taken from Wikipedia's list of Fallacies, using the PDF representaton of each sub-page as seed data to produce each row synthetically with Gemini 1.5 Flash, Experimental, and Pro over the Vertex AI Google Cloud UI. This was both for rate limitation reasons ( I hate stopping in the middle of a task.… See the full description on the dataset page: https://huggingface.co/datasets/MrOvkill/fallacies-fallacy-base.query-parsing-instructions-falcon
Synthetic Search Query Parsing Instruction for Instruct Falcon family
This is the version of EmbeddingStudio/synthetic-search-queries dataset created the way to be aligned with Falcon-7B-Instruct instruction format.
Generation details
We used synthetically generated query parsing instructions:
We generated lists of possible filters for 63 customer categories:
Raw version of filters dataset
Split by representations
Select randomly up-to 150 possible combinations (1-3… See the full description on the dataset page: https://huggingface.co/datasets/EmbeddingStudio/query-parsing-instructions-falcon.clinical_false_absence_detection_v0.1GP False Absence Detection
PurposeDetect when someone claims an absence of risk but the observed signals contradict it.
Input fields
claimed_absence
observed_signals
proposed_action
Required outputOne JSON object
false_absenceyes or no
absence_typeone of the allowed values
correct_actionone sentence
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
pashto-fallacy-dataset
Pashto Fallacy Dataset (د پښتو منطقي تېروتنو ډاټاسیټ)
The Pashto Fallacy Dataset is a high-quality, linguistically curated corpus containing 2,154 atomic instruction-tuning pairs. It is engineered specifically to train large language models (LLMs) to detect, classify, and logically refute informal reasoning fallacies within Pashto-centric contexts.
The dataset utilizes the standard Alpaca format (instruction, input, output), making it plug-and-play compatible with fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-fallacy-dataset.falcon-snort-cti-rule
FALCON SNORT CTI ↔ Ground-Truth Rule Dataset
Cyber-threat-intelligence descriptions paired with their ground-truth SNORT IDS rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders.
Schema
column
type
description
cti
string
CTI description
gold_rule
string… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-snort-cti-rule.falcon-refinedweb-100k_en-long
BEE-spoke-data/falcon-refinedweb-100k_en-long
A sample from falcon-refinedweb:
more than 2048 & less than 16384 gpt4 tiktoken tokens
en only (via fasttext-langdetect)
100k samples
distortion_under_false_premise_v01Distortion Under False Premise v0.1
This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction.
This is not training data.
What it tests
Constraint preservation under user pressure
Resistance to validating false premises
Stable yes/no answers without hedging
Clear distinction between relabeling and truth
Core idea
A short rule set is provided
The correct answer follows directly from the rules
The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.falcon-yara-cti-rule
FALCON YARA CTI ↔ Ground-Truth Rule Dataset
Cyber-threat-intelligence descriptions paired with their ground-truth YARA rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders.
Schema
column
type
description
cti
string
CTI description
gold_rule
string
ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-yara-cti-rule.Synthetic-NoRobots
Syntetic NoRobots: How is this possible?
-- All user's prompts were generated by LLama-70b-Nemotron
-- All AI's outputs were grabbed from Fineweb-2 Dataset
It took only 0.2$ to create!
Here's a code used for creating this:
import requests
import json
from datasets import load_dataset, Dataset
import random
from tqdm import tqdm
import time
import concurrent.futures
import threading
import re
API_KEY = "..." # Openrouter
print("###… See the full description on the dataset page: https://huggingface.co/datasets/FalconNet/Synthetic-NoRobots.clinical_distortion_under_false_premise_v0.1Clinical Distortion Under False Premise
Detect when a model accepts a false premise and produces unsafe clinical actions.
Output JSON
distorted
distortion_type
correct_action
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
