datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vsi-bench-debiased-evalgen_debiased_nli
Overview
Original dataset available here.
@inproceedings{gen-debiased-nli-2022,
title = "Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets",
author = "Wu, Yuxiang and
Gardner, Matt and
Stenetorp, Pontus and
Dasigi, Pradeep",
booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics",
month = may,
year = "2022",
publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/gen_debiased_nli.epidemic_sound_effects_t5_debiasedlaion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs).
verifier-debias-v2
verifier-debias-v2 — de-biased SFT data for a generative verifier
This dataset was presented in the paper One Token to Fool LLM-as-a-Judge.
GitHub repository: yulaizhao/Master-RM
SFT data to train a generative verifier (GenRM, arXiv:2408.15240)
from Qwen/Qwen2.5-7B-Instruct. Each row is a conversational example
(messages = system + user + gold assistant) plus a verdict (PASS/FAIL) and
the underlying 1–5 score. The assistant target is a critique ending in
Verdict: PASS / Verdict:… See the full description on the dataset page: https://huggingface.co/datasets/narcolepticchicken/verifier-debias-v2.laion_text_debiased_100M
100M Text Debiased Subset from LAION 2B
Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images.
Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data.
CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.llm-debiasing-benchmark
Dataset Card for LLM-Debiasing-Benchmark
This dataset contains the various texts and LLM annotations used in the paper Benchmarking Debiasing Methods for LLM-based Parameter Estimates.
We used texts from four corpora:
Bias in Biographies: https://huggingface.co/datasets/LabHC/bias_in_bios
Misinfo-general: https://huggingface.co/datasets/ioverho/misinfo-general
Amazon reviews: https://aclanthology.org/P07-1056/
Germeval18:… See the full description on the dataset page: https://huggingface.co/datasets/nicaudinet/llm-debiasing-benchmark.dllm-attention-parents-dreamreasoner-finecode-pred-debias5-b64-v1
DART FineCode deterministic sample — Dream-org/DreamReasoner-8B-Base dependency sidecar
This is the sealed attention dependency release generated
with pred-debias5 corrected attention, aligned a5 release. It is compatible with block size 64 and pairs only
with zimplex/dllm-dreamreasoner-finecode-b64-v1,
whose manifest SHA-256 is 4a0aece1752c3cd1b3fcd09c9e1690a2e4814ea7853f9a3ec5299c963f2ab9e9. The Dream tokenizer alignment proof is retained as DREAM_ALIGNMENT_COMPLETE.json.
The… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-attention-parents-dreamreasoner-finecode-pred-debias5-b64-v1.dllm-attention-parents-dreamreasoner-finemath-half-pred-debias5-b64-v1
DART FineMath-4+ half — Dream-org/DreamReasoner-8B-Base dependency sidecar
This is the sealed attention dependency release generated
with pred-debias5 corrected attention, aligned a5 release. It is compatible with block size 64 and pairs only
with zimplex/dllm-dreamreasoner-finemath-half-b64-v1,
whose manifest SHA-256 is 1085241e1db60b1afa9f73c440c71e3c4f04decc2bd8c2d47edf5041fb2d6dc5. The Dream tokenizer alignment proof is retained as DREAM_ALIGNMENT_COMPLETE.json.
The sidecar… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-attention-parents-dreamreasoner-finemath-half-pred-debias5-b64-v1.dllm-attention-parents-llada2-finecode-pred-debias5-b64-v1
DART FineCode deterministic 1/8 sample — inclusionAI/LLaDA2.0-mini dependency sidecar
This is the sealed attention dependency release generated
with pred-debias5 corrected attention. It is compatible with block size 64 and pairs only
with zimplex/dllm-finecode-dagcover-full-v1,
whose manifest SHA-256 is 45e250f9f4a5302ef1bb04c16577eb8b6d85996368fe1dc359b86465cdb03176.
The sidecar manifest SHA-256 is f4dbab97f495e23682f26ac085b19a90e9e6a15cf664b789d9efb0f23566cd3e. The source… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-attention-parents-llada2-finecode-pred-debias5-b64-v1.dllm-attention-parents-llada2-finemath-half-pred-debias5-b64-v1
DART FineMath-4+ half — inclusionAI/LLaDA2.0-mini dependency sidecar
This is the sealed attention dependency release generated
with pred-debias5 corrected attention. It is compatible with block size 64 and pairs only
with zimplex/dllm-finemath-dagcover-half-v1,
whose manifest SHA-256 is e20f3ecdc7417e8f8141610d4e1c55eb2973d269d03117d47c4e27b3d1182a68.
The sidecar manifest SHA-256 is 3357a4171087f4124b62e0ceed4beee47390c55dedc3c985c261a44b95fcd512. The source
release contains 5… See the full description on the dataset page: https://huggingface.co/datasets/zimplex/dllm-attention-parents-llada2-finemath-half-pred-debias5-b64-v1.Winogrande_debiaseddebiased_embeddingwinogrande_debiased-telugu-romanized-nodictwinogrande_debiased-teluguUnsplash-Debiased_by_AI-18K
Unsplash-Debiased-by-AI-18K
17,812 high-quality image captions selected by a label-free, dual-modality, 6-judge
debiasing pipeline from 32,135 Unsplash-40K captions. Each caption passes three gates:
(a) high/mid tier in a debiased judge consensus (family-balanced × reliability-weighted,
rank-calibrated, self-preference-dropped), (b) image quality (NIQE/MUSIQ/LIQE), and
(c) image–caption alignment (CLIPScore/SigLIP).
Images are NOT included. This dataset ships captions +… See the full description on the dataset page: https://huggingface.co/datasets/henrywch2huggingface/Unsplash-Debiased_by_AI-18K.pick_place_tactile_debiasThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Inspire_2cam",
"total_episodes": 216,
"total_frames": 114402,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:216"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dishantpatel1207/pick_place_tactile_debias.winogrande_debiased-telugu_filteredwinogrande_debiased-telugu-romanizedBias-Debias-Alpaca
Responsible Media Content Matrix (RMCM): Overview
The RMCM is a strategic tool developed to address various forms of bias and unethical practices in media reporting. It encompasses several key categories, each focusing on a specific type of bias or ethical concern. The primary objective of the RMCM is to foster responsible journalism and content creation by providing clear guidelines on identifying and rectifying biased or harmful content.
Key Categories of RMCM:… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/Bias-Debias-Alpaca.gender_debias_disambiguate_oldgender_debias_disambiguatedebiased_dataset
Dataset Description
About the Dataset:
This dataset contains text data that has been processed to identify biased statements based on dimensions and aspects. Each entry has been processed using the GPT-4 language model and manually verified by 5 human annotators for quality assurance.
Purpose:
The dataset aims to help train and evaluate machine learning models in detecting, classifying, and correcting biases in text content, making it essential for NLP research related to fairness… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/debiased_dataset.audioset_t5_debiasedLLaMa_3B_NoCoT_DebiasingInstructionLLaMa_3B_Debiasing_Instruction_CoTtree-debiasing-stage3LLaMa_1B_Debiasing_Instruction_CoTLLaMa_3B_Debiasing_Instruction_NoCoTLLaMa_1B_Debiasing_Instruction_CoT_inferenced
