datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KinshipKingJamesVersionBiblegest
GEST Dataset
This is a repository for the GEST dataset used to measure gender-stereotypical reasoning in language models and machine translation systems.
Paper: Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
Code and additional data (annotation details, translations) are available in our repository
Changelog
December 6th 2024 - gest_1.1.csv was added. This is a new version that has 244 typos and other errors… See the full description on the dataset page: https://huggingface.co/datasets/kinit/gest.NMT_Rwandan-Gazette_parallel_data_en_kin
Dataset Details
Dataset Description
This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix
Curated by: Digital Umuganda
Language(s) (NLP): Kinyarwanda and English
License: cc-by-4.0
Dataset Sources [optional]
The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.Kinship
Kinship KG–QA (Hinton)
A lightweight knowledge-graph + question answering (KGQA) resource adapted from the UCI Kinship dataset created by Geoff Hinton.
This dataset provides a small family-tree knowledge graph paired with templated multi-hop QA tasks designed for MultiHop KGQA.
Project: THESEUSPaper: Theseus in the GraphOriginal dataset: UCI Kinship
Key Features
Small knowledge graph with two family trees
Human-readable entities and relations
Templated 1--3 hop… See the full description on the dataset page: https://huggingface.co/datasets/HalcyonSolutions/Kinship.kids-product-recalls
KindlyKiddo Kids’ Product Recall Data
Machine-readable CPSC, NHTSA, and FDA recall records for products made for or primarily used by children. The dataset is normalized by KindlyKiddo and refreshed weekly from official U.S. federal sources.
Files
kids-product-recalls.csv — flat tabular dataset used by the Hugging Face viewer.
kids-product-recalls.json — records plus refresh time, coverage, source, license, and repository metadata.
data-dictionary.csv — field… See the full description on the dataset page: https://huggingface.co/datasets/kindlykiddo/kids-product-recalls.Dr.Sparse-OTF-test-set
Dr.Sparse OTF Test Set
100 sparse matrices from the SuiteSparse Matrix Collection,
converted to the flat binary format the Dr.Sparse
benchmark harness reads. This is the held-out evaluation set for LLM-generated
CUDA sparse kernels (SpMV / SpMM / SpGEMM), kept separate from the matrices the
models were developed against.
Layout
Matrices are grouped into size tiers by row count, the convention Dr.Sparse task
discovery scans for:
tier
rows
matrices
size… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-OTF-test-set.book_summary_datasetImage-Gen-or-Image-Editing
Image Gen or Image Editing
This dataset is designed for text classification of prompts provided by users. It determines whether a prompt is intended for image generation or image editing.
United-Kingdom-Stock-Symbols-and-Metadata
United Kingdom Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in United Kingdom.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/United-Kingdom-Stock-Symbols-and-Metadata.sma-upper-limb-kinect
SMA Upper-Limb Kinect Dataset and Reach-Intent Benchmark
This repository contains the four supplementary files from a longitudinal Kinect study and a derived reach-intent benchmark.
Publication files
The repository contains the four supplements listed by PLOS:
pone.0170472.s001.pdf: complete analysis report;
pone.0170472.s002.zip: prototype game, R analysis, and Python extraction source code;
pone.0170472.s003.zip: raw data, logs, extracted features, and… See the full description on the dataset page: https://huggingface.co/datasets/YannisTevissen/sma-upper-limb-kinect.Kinyarwanda_Engligh_Multilingual_ASRThis dataset was created from Mozilla's Common Voice dataset for the purposes of Multilingual ASR on Kinyarwanda and English.
The dataset contains 3000 hours of multilingual training samples, 300 hours of validation samples and 200 of testing samples.
kinyarwanda-tts-dataset
Kinyarwanda dataset for text to speech model
Kinyarwanda dataset for text to speech model holds data for ai modelling of Kinyarwanda chatbots or other use cases.
64-que-kinh-dich
64 quẻ Kinh Dịch
The 64 hexagrams of the I Ching
1. Mô tả · Description
Đủ 64 quẻ theo thứ tự Chu Dịch, kèm tên Việt, tên Hán, quẻ thượng, quẻ hạ và tượng quẻ.
All 64 hexagrams in King Wen order, with Vietnamese and Han names, upper and lower trigrams, and the image.
Số dòng · Rows: 64
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa · Meaning
id
string… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/64-que-kinh-dich.common-voice-kinyarwanda-english-dataset
Kinyarwanda-English Commonvoice dataset
A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR
Note: The audio dataset shall be added in the future
kinetics-400-splitsBGLdatasets-for-JCSEADeLe_battery_v1dot0
Dataset Card for ADeLe
Dataset Summary
ADeLe (Annotated-Demand-Levels) battery is a single, unified test set whose every item is labelled with the level (0-5+) it demands on 18 general ability dimensions (e.g. attention and scan, logical reasoning, various knowledge areas) plus an “unguessability” dimension. It is produced by applying the DeLeAn rubrics, via GPT-4o annotators, to AI benchmarks.
Version 1.0 contains 16 108 items drawn from 63 tasks spread across a diverse… See the full description on the dataset page: https://huggingface.co/datasets/CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0.ds_benchmark_edicom_edicom_3kin_20250421PIQA-kinNinjaMasker-PII-Redaction-DatasetBugzilla_Eclipse_Bug_Reports_Dataset
Special Thanks
Special thanks to Lamkanfi, Ahmed; Pérez, Javier; and Demeyer, Serge for their contributions. Please cite their paper, as this dataset is the processed part of their dataset.
Citation
@INPROCEEDINGS{6624028,
author={Lamkanfi, Ahmed and Pérez, Javier and Demeyer, Serge},
booktitle={2013 10th Working Conference on Mining Software Repositories (MSR)},
title={The Eclipse and Mozilla defect tracking dataset: A genuine dataset for mining bug… See the full description on the dataset page: https://huggingface.co/datasets/kinggdygi2/Bugzilla_Eclipse_Bug_Reports_Dataset.enzyme_sequencesHTML-CSS-UIenglish_islamqainfo
Dataset Card for English Islam QA Info
Dataset Description
The English Islam QA Info (19,052 questions and answers) is derived from the IslamQA website and contains curated question-and-answer pairs categorized by topic. It serves as a resource for multilingual and cross-lingual natural language processing (NLP) tasks. This dataset is part of a broader initiative to enhance the understanding and computational handling of Islamic jurisprudence and advice.
Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/english_islamqainfo.Kinyarwanda_English_parallel_dataset
Kinyarwanda-English parallel text
This dataset contains 55,000 Kinyarwanda-English sentence pairs, obtained by scraping web data from religious sources such as:
Bible
Quran
This dataset has not been curated only cleaned.
NMT_Health_parallel_data_en_kinRecommovie-TMDB-Dataset
Backend Data Artifacts
This directory stores the precomputed data used by the FastAPI recommender.
Files
merged_embeddings.npy — Float32 matrix of movie embeddings. Shape is recorded in merged_shape.txt (rows × dims).
merged_shape.txt — Two integers separated by a space or comma indicating the embedding matrix shape, e.g. 200000 384.
index.faiss — FAISS index built from merged_embeddings.npy for fast nearest‑neighbor queries.
Optional/auxiliary:
Any metadata/lookups… See the full description on the dataset page: https://huggingface.co/datasets/kinjalrk2k/Recommovie-TMDB-Dataset.
