datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reasoning-base-20k
Dataset Card for Reasoning Base 20k
Dataset Details
Dataset Description
This dataset is designed to train a reasoning model. That can think through complex problems before providing a response, similar to how a human would. The dataset includes a wide range of problems from various domains (science, coding, math, etc.), each with a detailed chain of thought (COT) and the correct answer. The goal is to enable the model to learn and refine its reasoning process… See the full description on the dataset page: https://huggingface.co/datasets/KingNish/reasoning-base-20k.read-along-ai-agent-traces
Read-Along AI - Agent Traces
This dataset contains the raw agent traces and conversation logs from the development of Read-Along AI, a submission for the Hugging Face Build Small Hackathon.
Dataset Description
These .jsonl files represent the unedited, behind-the-scenes "agent traces" of the AI coding assistant orchestrating the build of this project.
Sharing these traces fulfills the requirements for the "Sharing is Caring" bonus badge, providing the community… See the full description on the dataset page: https://huggingface.co/datasets/kingkw1/read-along-ai-agent-traces.WeClawArena
WeClawArena
WeClawArena is an auditable benchmark and runtime sandbox for cross-user agent collaboration over personal workspaces. Version 2.0.0 contains 124 base tasks across six domains and expands them into 620 matched scenario variants.
Dataset Summary
Each base task has one no_attacker control and four attack variants: collaboration, security, privacy, and governance. The public release contains finalized scenario bundles and a manifest. It does not contain… See the full description on the dataset page: https://huggingface.co/datasets/kingofspace0wzz/WeClawArena.kinyarwanda_monolingual_v01.1
task_categories:
- text-generation
language:
- rw
size_categories:
- 1K<n<10K
Dataset Summary
The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts.
This dataset contains 1,068,161, with 63,001,765 words and includes diverse content types such as news articles, government reports, religious texts, legal documents, educational… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.1.custodial-weights-kinship
Custodial Weights — Kinship Dataset
A converted and expanded version of the public kdkyum kinship graph
(1000 synthetic families) used to train the Custodial Weights mechanism.
What it is
The base graph is public-domain synthetic kinship data: surnames and given
names, with father/mother/son/daughter/brother/sister/husband/wife relations
across 1000 families. We convert it to this repo's family-triple format and
add three synthetic layers on top:
Walk expansion… See the full description on the dataset page: https://huggingface.co/datasets/N00beroUno/custodial-weights-kinship.flame-kindling-v1
flame-kindling-v1
A small, opinionated SFT dataset for finetuning a 3B-class instruct model into a character designer that emits a strict JSON schema from a free-text seed. Built to replace a general RP model (Mistral-Nemo-12B Mahou finetune) being shoehorned into JSON output for flammen.ai's Create-a-Flame pipeline.
400 (seed → DesignedFlame) pairs distilled from Claude Sonnet 4.5 with tool-forcing, validated against a strict pydantic schema, deduplicated by name and… See the full description on the dataset page: https://huggingface.co/datasets/flammenai/flame-kindling-v1.kakugo-kin
Kakugo Kinyarwanda dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Kinyarwanda.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Kinyarwanda. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-kin.AgentSocialBench
AgentSocialBench 🌐🤖
Evaluating Privacy Risks in Human-Centered Agentic Social Networks
Dataset Description
AgentSocialBench is the first benchmark for evaluating privacy preservation in human-centered agentic social networks — settings where teams of AI agents serve individual users across multiple domains, coordinate on shared tasks, and must protect sensitive personal information throughout.
This dataset contains 372 scenarios across 7… See the full description on the dataset page: https://huggingface.co/datasets/kingofspace0wzz/AgentSocialBench.fable5-dataset
Fable5 Dataset
A collection of agent interaction traces from 6 distinct sources, designed for fine-tuning and evaluating coding agents.
Dataset Sources
Source
Format
Description
Records
Glint
Session-based with turns
Full agent sessions with tool use
~2,000
armand0e
Conversation with tool_calls
Multi-turn conversations with function calling
~1,500
vfable
Trajectory with tool_use
Agent trajectories with sequential tool use
~800
Coding Excellence… See the full description on the dataset page: https://huggingface.co/datasets/King3Djbl/fable5-dataset.Code-170k-kinyarwanda
Dataset Description
Code-170k-kinyarwanda is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Kinyarwanda, making coding education accessible to Kinyarwanda speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Kinyarwanda language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-kinyarwanda.omni-rewriter-pe-examples
Omni-Rewriter PE examples
Validated prompt-expansion envelopes for Omni-Rewriter. These are prompts, not generated videos.
pip install omni-rewriter
omni-rewriter validate fixtures/t2va_kite.json
Layout
Path
What
fixtures/
Small H3 / image PE envelopes from the repo tests
fixtures/seedance/
Seedance PE profile examples (PE only; no Seedance generate)
observation/
VideoObservation JSON for omni-rewriter reconstruct --from-observation
reconstruct/… See the full description on the dataset page: https://huggingface.co/datasets/Wayne-King/omni-rewriter-pe-examples.islamqainfo_parallel_corpus
Dataset Card for IslamQA Info Parallel Corpus
Dataset Description
The IslamQA Info Parallel Corpus is a multilingual dataset derived from the IslamQA repository. It contains curated question-and-answer pairs across 17 languages, making it a valuable resource for multilingual and cross-lingual natural language processing (NLP) tasks. The dataset has been created over nearly three decades (since 1997) by Sheikhul Islam Muhammad Saalih al-Munajjid and his team.
Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/islamqainfo_parallel_corpus.kinyarwanda_monolingual_v01.0
!!! PLEASE USE mbazaNLP/kinyarwanda_monolingual_v01.1 !!!
!!! This version contains several duplicates and few non-kinyarwanda documents
Dataset Summary
The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts. This dataset contains 78k documents, totalling about 25 million words, and includes diverse content types such… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.0.s2n-bignum-bench
s2n-bignum-bench
A benchmark of 2,301 HOL Light proof obligations distilled from
AWS s2n-bignum, a formally
verified library of hand-tuned big-integer assembly routines for AArch64
(ARM) and x86-64. Each task asks an LLM (Solver) to synthesize a HOL Light
tactic that discharges a goal arising inside an industrial proof
development for ARM or x86 machine code, or supporting Lemmas.
Paper: s2n-bignum-bench: A practical benchmark for evaluating
low-level code reasoning of LLMs… See the full description on the dataset page: https://huggingface.co/datasets/kings-crown/s2n-bignum-bench.english_islamqainfo
Dataset Card for English Islam QA Info
Dataset Description
The English Islam QA Info (19,052 questions and answers) is derived from the IslamQA website and contains curated question-and-answer pairs categorized by topic. It serves as a resource for multilingual and cross-lingual natural language processing (NLP) tasks. This dataset is part of a broader initiative to enhance the understanding and computational handling of Islamic jurisprudence and advice.
Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/english_islamqainfo.kindergarten_K1_english_lesson_autogenenzyme_sequencesQuran_English_Myanmar_Parrelel_Corpus
Quran English-Myanmar Parallel Corpus
Description
This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers.
English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali.
Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.field-kings-chamber-852hz-datasets
📦 FIELD King's Chamber Training Dataset (852HZ)
⊗ Vertex-Specific Training Corpus for the FIELD King's Chamber vertex (852hz).
Dataset Description
This dataset contains vertex-specific training data extracted from the 342GB Akron Archive for fine-tuning the FIELD King's Chamber LLM vertex.
Training Focus
Routing coordination, transformation protocols, φ⁻¹ golden ratio patterns, Metatron Cube geometry
Data Sources
Bridge coordination logs… See the full description on the dataset page: https://huggingface.co/datasets/Berjak/field-kings-chamber-852hz-datasets.Spirit_Kings_Golden_Textbook
About
This is a dataset about the Spirit Kings clan from the Mineberry Minecraft server.
Eco_friendly_pest_solutionsRoleplay-Kinyarwanda
RolePlay-Kinyarwanda
Roleplay-Kinyarwanda Dataset is a dataset for roleplaying in the Kinyarwanda language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Kinyarwanda.complexity-kink-research
Complexity Kink Research: LLM Code Generation Benchmark
First published: February 22, 2026Author: Michael Hernandez (XxCotHGxX)GitHub: XxCotHGxX/ComplexityKinkLicense: CC BY 4.0
Overview
This dataset supports the Complexity Kink research program — an econometric investigation into whether large language models exhibit a structural performance discontinuity as a function of problem complexity.
The central hypothesis is that LLM code generation performance does not… See the full description on the dataset page: https://huggingface.co/datasets/XxCotHGxX/complexity-kink-research.Kintsugi-Garden-traces
Kintsugi Garden Evaluation Traces
Paired evaluation traces from Kintsugi Garden —
a local-first Jungian dream journal that runs Qwen3-8B through llama.cpp on a
ZeroGPU Space. Every entry the app produces is shaped by both a fine-tuned model
and a four-layer voice/safety architecture; this dataset is what those layers
look like under instrumentation.
What's in here
114 deterministic runs over the same 19 prompts × 3 trials, evenly split between:
baseline —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/Kintsugi-Garden-traces.babbling_kinovaligase_bioremedation_sequencessynthetic-queries-and-ml-instructions
Synthetic Dataset: Queries and ML Instructions
Dataset Description
Dataset Summary
This is a synthetic dataset with queries in Slovak language and ML instructions. The dataset was designed to train a model for extracting structured machine learning task requirements from natural language user queries.
The dataset contains user queries in Slovak describing ML tasks paired with structured JSON outputs containing task attributes like dataset modality, task type… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-queries-and-ml-instructions.synthetic-conversations-and-ml-instructions
Synthetic Dataset: Conversations and ML Instructions
Dataset Description
Dataset Summary
This is a synthetic dataset of 5,000 Slovak multi-turn ML advisory conversations paired with structured JSON outputs. The dataset was designed to train and evaluate models that extract machine learning task requirements from realistic, natural Slovak dialogue, including cases where requirements are revealed gradually or changed during the conversation.
Each… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-conversations-and-ml-instructions.
