datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indice-fallos-ia-espanol
Indice de Fallos IA en espanol
Snapshots mensuales del Indice de Fallos IA producido por el observatorio La AutopsIA (ApisDom Intelligence Group). Mide la fiabilidad de modelos LLM con benchmarks oficiales independientes, en formato citable y trazable.
Cifras del snapshot actual
765 mediciones en este snapshot.
Mes archivado: 2026-09.
Recomputado: 2026-09-01T21:34:26.058Z.
Frecuencia: sincronizacion mensual.
Para que sirve este dataset
Datos… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/indice-fallos-ia-espanol.medlineplus1FallAllD-WristhealifyLLM-QA-scraped-datasetopenassistant-falcon
Chat Fine-tuning Dataset - OpenAssistant Falcon
This dataset allows for fine-tuning chat models using '\Human:' AND '\nAssistant:' to wrap user messages.
It still uses <|endoftext|> as EOS and BOS token, as per Falcon.
Sample
Preparation:
The dataset is cloned from TimDettmers, which itself is a subset of the Open Assistant dataset, which you can find here. This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples.
The… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/openassistant-falcon.falsifyrl-adapted
FalsifyRL AutoScientist-Adapted Dataset
This is the exact audited Adaptive Data export used to train the FalsifyRL AutoScientist model.
train.csv is the immutable exported training artifact; its SHA-256 digest and Adaption dataset/run
identifiers are recorded in adaptation-audit.json and release-manifest.json.
FalsifyRL trains a critic to identify and repair proxy-reward failures in embodied multi-agent
reinforcement learning. Each input includes a task specification… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-adapted.arg_emo_fallacyThis dataset accompanies the paper Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detection. It includes annotations for logical fallacy labels, emotion categories, and argument convincingness ratings. Please refer to the paper and its repository for more details.
If you use this dataset, please include the following citation:
@inproceedings{chen-etal-2026-emotionally,
title = "Emotionally Charged, Logically Blurred: {AI}-driven Emotional Framing… See the full description on the dataset page: https://huggingface.co/datasets/nllg/arg_emo_fallacy.false-street-6418b6
false-street-6418b6
Synthetic sensors test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/kestrelWing/false-street-6418b6.fallacies-fallacy-base
Fallacies
This dataset was produced for the purpose of enabling more accurate detection and handling of logical and other fallacies in LLMs. Video Summary
Provenance
Seed data taken from Wikipedia's list of Fallacies, using the PDF representaton of each sub-page as seed data to produce each row synthetically with Gemini 1.5 Flash, Experimental, and Pro over the Vertex AI Google Cloud UI. This was both for rate limitation reasons ( I hate stopping in the middle of a task.… See the full description on the dataset page: https://huggingface.co/datasets/MrOvkill/fallacies-fallacy-base.clinical_false_absence_detection_v0.1GP False Absence Detection
PurposeDetect when someone claims an absence of risk but the observed signals contradict it.
Input fields
claimed_absence
observed_signals
proposed_action
Required outputOne JSON object
false_absenceyes or no
absence_typeone of the allowed values
correct_actionone sentence
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
humanoid-fall-recovery-dataset-v1Dataset for recovering balance after falls.
Description
Body orientation and ground contact signals mapped to recovery movements.
Task Description
Enables humanoid robots to safely stand up and regain balance after falling.
medcuad_datamedcuad_tr_datadistortion_under_false_premise_v01Distortion Under False Premise v0.1
This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction.
This is not training data.
What it tests
Constraint preservation under user pressure
Resistance to validating false premises
Stable yes/no answers without hedging
Clear distinction between relabeling and truth
Core idea
A short rule set is provided
The correct answer follows directly from the rules
The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.tidesTIDEs is a dataset for Transphobia Identification in Digital Environments. You can read our full paper here: https://dl.acm.org/doi/full/10.1145/3715275.3732095
Note that this dataset is fairly small and exploratory. TIDEs is subject to grow in future iterations of the dataset to improve its accuracy and transphobia detection capability. Please email flameiro at umich dot edu if you have any questions!
4_MPlus_Health_Topics_QAo1-medical-finetuning-trfallacy-datasetclinical_distortion_under_false_premise_v0.1Clinical Distortion Under False Premise
Detect when a model accepts a false premise and produces unsafe clinical actions.
Output JSON
distorted
distortion_type
correct_action
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
SC-ZS-test_AURORA-Gold-SDG_True-Positives-and-False-Positivesbhoomi50kemergency-medicineworld-cities-falseThis data is licensed under the Creative Common Attribution License as is the original data from https://www.geonames.org/. This means you have to credit geonames when using the data.
The data was taken from https://github.com/datasets/world-cities/ and modified to be entirely full of lies.
false_absence_detection_v01ClarusC64/false_absence_detection_v01
Dataset summary
This dataset tests whether models treat invisible entities as gone or still present.In each sequence, a target leaves the camera view.Some exits are real.Some are false absences with clear evidence that the target remains in the container.
Goal
check if models infer continued presence from indirect cues
avoid treating every disappearance as an exit
keep spatial grounding under occlusion and clutter
Key signals
absence_tag:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/false_absence_detection_v01.startup-investor-falsification-question-bank
Startup Investor Falsification Question Bank
A founder's materials are easier to review when each important claim has a
question that could disprove it. This open CSV resource gives founders 48
outside-reader questions across the same 12 dimensions used in a structured
private-company review.
Each row records the claim type, a falsification question, evidence to request,
a contradiction signal, the state to retain when the question is unresolved,
a founder repair action and a… See the full description on the dataset page: https://huggingface.co/datasets/mheilimo/startup-investor-falsification-question-bank.text_to_sql_FALCONfalcon-toc-generationreddit_falcon_summariesUS_Trump_2020_social_mediaalpaca_text
