datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
in-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.incremental-instruction-creative-writing
Incremental Instruction Creative Writing
Does delivering a writing brief over several conversation turns change what a
language model writes? This dataset supports that question with matched
creative-writing tasks evaluated under two delivery conditions:
FULL: the complete brief is supplied in one turn.
SHARDED: the same intended brief is introduced across five to nine turns.
The benchmark holds task content fixed while varying how the instructions are
delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.inconvenience-public-safety
inconvenience-public-safety
Three Korean public-safety registers converted to Korean braille under the 2017
revised rules (문화체육관광부 고시 제2017-15호). Every register is enumerated in
full, not sampled.
The registers are here because their documents are shaped differently, not
because three is more than one. A pesticide row is a filled-in form; a patient
leaflet is prose; an accident case is a paragraph an investigator wrote. Median
record length spans more than an order of magnitude… See the full description on the dataset page: https://huggingface.co/datasets/Yuyongkim/inconvenience-public-safety.countdown-arithmetic-training-pool
Countdown arithmetic training pool
Arithmetic puzzles of the Countdown kind: a handful of source numbers, a target, and the job of
writing an expression over the four operations that reaches the target, using each source number
at most once and not having to use them all. A set generated for this pool and three public
datasets read at the pinned revisions named below, laid out twice. Train on either layer or on
both.
pool.jsonl
Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/countdown-arithmetic-training-pool.opensre-incident-trajectories
OpenSRE Incident-Diagnosis Trajectories
Graded, multi-step SRE incident-diagnosis trajectories. A frozen LLM reads evidence through
diagnostic tools (describe_pod / get_events / get_logs / get_metrics / query_traces / …),
states a root cause + category + fix, and is scored on substance against ground truth. Built as a
HUD v6 RL environment with a deliberate model spanning set so difficulty is legible and the
within-group reward spread is real (the GRPO learning signal).
197… See the full description on the dataset page: https://huggingface.co/datasets/quantranger/opensre-incident-trajectories.inconvenience-msds
inconvenience-msds
Inconvenience #01 — Material Safety Data Sheets in Korean braille
Accessibility infrastructure for Korean chemical safety information.
48,966 chemicals × up to 16 MSDS sections, encoded as Korean braille
following the 2017 한국 점자 규정 (Korean Braille Standards).
What this is
This dataset is the first entry in the inconvenience series — a planned
sequence of accessibility-infrastructure datasets targeting domains where
the visually impaired user… See the full description on the dataset page: https://huggingface.co/datasets/Yuyongkim/inconvenience-msds.DuET-PD
DuET-PD: Dual Evaluation for Trust in Persuasive Dialogues
Dataset Summary
DuET-PD is a comprehensive framework and dataset designed to evaluate the robustness and adaptability of Large Language Models (LLMs) in multi-turn persuasive dialogues. The dataset probes an LLM's ability to navigate the critical tension between resisting misinformation (robustness) and accepting valid corrections (adaptability).
The "Dual" aspect of DuET-PD reflects its two core evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Incomple/DuET-PD.prometech_inc_basic_coder
Prometech Inc Basic Coder Dataset
Dataset Overview
Filename: prometech_inc_basic_coder.jsonlTotal Entries: 263,903File Size: ~854 MBProvider: Prometech Bilgisayar Bilimleri AŞ
This dataset is a unified collection of high-quality coding instruction-following records, designed for fine-tuning Large Language Models (LLMs) or for use in Retrieval-Augmented Generation (RAG) systems. It aggregates data from multiple open-source high-quality datasets, synthetic documentation… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/prometech_inc_basic_coder.incel-wiki
Incel Wiki
A full dump of Incel Wiki, a MediaWiki-based encyclopedia documenting incel culture, terminology, communities, etc. The dump includes all 3,138 pages (1,698 articles and 1,440 redirects) with raw wikitext markup preserved.
Columns
Column
Type
Description
title
string
Page title
page_id
int
MediaWiki page ID
revision_id
int
Revision ID of the exported version
timestamp
string
Last edit timestamp (ISO 8601)
contributor
string
Username of last… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/incel-wiki.Inclusive_Leadership_Belonging_Practical
Inclusive Leadership Belonging — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Inclusive_Leadership_Belonging_Practical.Inclusive_Leadership_Belonging_Theory
Inclusive Leadership Belonging — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Inclusive_Leadership_Belonging_Theory.
