datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eagle
Eagle 🦅: Ethical Dataset Given from Real Interactions
Introduction
This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs).
If you use the Eagle dataset in your research, please cite the following:
@inproceedings{Eagle:arxiv:2024,
title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.AfriADRMASS-EX
MASS-EX: Expert-Annotated Dataset for Interpretable Sleep Staging
中文版
Associated Paper:Guifeng Deng, Pan Wang, Jiquan Wang, Shuying Rao, Junyi Xie, Wanjun Guo, Tao Li, Haiteng Jiang. "SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model." arXiv preprint, 2026. arXiv:2603.26738
Authors
Name
Affiliation
ORCID
Guifeng Deng
Zhejiang University
0009-0001-1940-7797
Pan Wang
Wenzhou Medical University
0009-0001-6664-6934
Wanjun… See the full description on the dataset page: https://huggingface.co/datasets/Feng613/MASS-EX.hls-ast-sagehls
SAGE-HLS Dataset: AST-Guided HLS-C Code Generation
The SAGE-HLS dataset is a large-scale, synthesis-friendly dataset for natural language to HLS-C code generation, enhanced by AST (Abstract Syntax Tree) representations. It supports training LLMs to generate high-quality, synthesis-ready high-level synthesis (HLS) code from functional descriptions, with structure-aware guidance.
📦 Dataset Structure
Each sample in the dataset contains the following fields:
Field… See the full description on the dataset page: https://huggingface.co/datasets/mashnoor/hls-ast-sagehls.odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.turkish-pii-masking-benchmark
Turkish PII Masking Benchmark (1,000 test cases)
A hand-built, fully synthetic benchmark for evaluating instruction-conditional PII masking
in Turkish. Designed to stress-test small LLMs (0.3B-1B) fine-tuned for privacy filtering in
banking/ERP user messages. All values are synthetic (no real persons); all sentence patterns
were written specifically for this benchmark (no training-set overlap).
Task
Given instruction (the masking policy) and input, the model must… See the full description on the dataset page: https://huggingface.co/datasets/cagrigungor/turkish-pii-masking-benchmark.cooking-master-boy-subtitle
Cooking Master Boy Chat Records
Chinese (trditional) subtitle of anime "Cooking Master Boy" (中華一番).
Introduction
This is a collection of subtitles from anime "Cooking Master Boy" (中華一番).
Dataset Description
The dataset is in CSV format, with the following columns:
episode: The episode index of subtitle belogs to.
caption_index: The autoincrement ID of subtitles.
time_start: The starting timecode, which subtitle supposed to appear.
time_end: The ending… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/cooking-master-boy-subtitle.chat-cooking-master-boy-100k
Cooking Master Boy Chat Records
Chat record dataset from Twitch channel "muse_tw" during the "Cooking Master Boy" (中華一番) marathon event.
Introduction
This is a chat dataset collected from Twitch channel "muse_tw", while the channel is hosting a marathon anime event featuring "Cooking Master Boy" (中華一番).
The featured anime "Cooking Master Boy" is a Japanese manga series written and illustrated by Etsushi Ogawa. And has a big impact on meme culture, and has a cult following… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/chat-cooking-master-boy-100k.persian_news_typos
Corrected News Articles Dataset
Overview:
This dataset contains pairs of raw text and their corrected versions, providing a valuable resource for tasks related to text correction, language modeling, and natural language processing (NLP). The dataset has been derived from various news articles and reports, focusing on correcting typographical errors, grammatical mistakes, and stylistic inconsistencies in the text. Each entry in the dataset consists of two fields: text (the… See the full description on the dataset page: https://huggingface.co/datasets/masoudkaviani/persian_news_typos.factver_master
Dataset Card for Dataset Name
FactVer_v2.0
Dataset Details
The dataset is curated for XAI research in Automated Fact Verification, to address the lack of explanation-focused datasets and the overemphasis on local explainability.
Dataset Description
It pairs each claim with multi ple annotated pieces of evidence within its thematic context (e.g., Climate change, COVID-19, Electric Vehicles). The dataset facilitates both local and global explainability by… See the full description on the dataset page: https://huggingface.co/datasets/manjuvallayil/factver_master.chat-cooking-master-boy-XL
Cooking Master Boy Chat Records
Chat record dataset from Twitch channel "muse_tw" during the "Cooking Master Boy" (中華一番) marathon event.
Introduction
This is a chat dataset collected from Twitch channel "muse_tw", while the channel is hosting a marathon anime event featuring "Cooking Master Boy" (中華一番).
The featured anime "Cooking Master Boy" is a Japanese manga series written and illustrated by Etsushi Ogawa. And has a big impact on meme culture, and has a cult following… See the full description on the dataset page: https://huggingface.co/datasets/h-alice/chat-cooking-master-boy-XL.
