datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xd-violence-rgb-videomae-chunked-testkinetics400
Kinetics-400 Video Dataset
This dataset is derived from the Kinetics-400 dataset, which is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Attribution
This dataset is derived from:
Original Dataset: Kinetics-400
Original Authors: Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, Andrew Zisserman
Original Paper: "The… See the full description on the dataset page: https://huggingface.co/datasets/liuhuanjim013/kinetics400.wikipedia_pos_tagged
POS tagged Wikipedia
This dataset is a POS-tagged version of the wikipedia dataset.
Different versions exist in this dataset:
nltk - these are tagged by the nltk pos tagger
spacy - these are tagged by the en_core_web_sm pos tagger
simple - these are from the simple English Wikipedia
prlang
Dataset Card for "prlang"
More Information needed
Kinetics-700
Damaged Videos List
The following rows are removed due to damaged video files:
NNazT7dDWxA_000130_000140
5d9mIpws4cg_000130_000140
SYTMgaqGhfg_000010_000020
ixQrfusr6k8_000001_000011
BSN_nDiTwBo_000004_000014
y7cYaYX4gdw_000047_000057
A-FCzUzEd4U_000000_000010
_dbw-EJqoMY_001023_001033
zLD_q2djrYs_000030_000040
FAqHwAPZfeE_000018_000028
kinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.bible_king_james_version_en
King James Version (1611)
Description
The most influential English Bible translation in history, commissioned by King James I of England and first published in 1611. The translation was prepared by 47 scholars organized into six committees, working from the original Hebrew, Aramaic, and Greek texts, as well as consulting earlier English translations (Tyndale, Coverdale, Geneva Bible) and the Latin Vulgate. The KJV is renowned for the majesty of its prose and its… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/bible_king_james_version_en.KSAFE-MM
KSAFE-MM
📑 Paper |
🛠️ Technical Blog
📢 News
⚡️ 2026/06/11: Released on Hugging Face 🤗
📑 2026/05/29: arXiv preprint released
📕 2026/05/20: Technical blog article published
⚠️ CONTENT WARNING
This dataset contains potentially harmful and sensitive visual and textual content across the following 11 safety risk categories:
Risk Domain
Categories
Content Safety Risks
Hate and Unfairness, Violence, Sexual, Self-harm
Socio-economic Risks
Political and… See the full description on the dataset page: https://huggingface.co/datasets/K-intelligence/KSAFE-MM.FineMath-3pluskinyarwanda_afrivoice_all_domains_v0.2
Kinyarwanda AfriVoice — All Domains (v0.2)
Cleaned version of ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.
Changes from v0.1
Removed rows with empty/null transcription values across all splits (train/validation/test)
Audio and domain labels unchanged; only null-transcription rows were dropped
Source
Original data from DigitalUmuganda/Afrivoice_Kinyarwanda (CC-BY-4.0),
extracted and concatenated across 5 domains (agriculture, education… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.2.serine_threonine_kinase_33_butkiewicz
Dataset Details
Dataset Description
The serine/threonine kinase, STK33, has been shown to
be relevant for proliferation of mutant KRAS-dependent cells involved
in cancer. Primary screen AID 2661. Counter screen AID 2821. AID504583
as validation screen. Actives in AID 2821 subtracted by the actives
from screen AID504583 resulted in the final set of 172 active
compounds.
Curated by: Adrian Mirza, Michael Pieler and Kevin Maik Jablonka
License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/serine_threonine_kinase_33_butkiewicz.kinova-pick-placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova",
"total_episodes": 50,
"total_frames": 8396,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/kinova-pick-place.kinetics-400zangei-dit-stage-1-250k-256px-dinov3
Zangei 256px DINOv3 Features — Repacked
Repacked from the known Colab extraction layout.
Source dataset: kingsidharth/zangei-dit-stage-1-250k-256px-img
Feature model: DINOv3 ViT-S/16
Image variants:
resized
square
Feature kinds:
cls
reg
patch
Files:
shards/dinov3_vits16_256_.safetensors
Matching row index:
shards/dinov3_vits16_256_.parquet
Tensor keys:
cls
reg
patch
Manifests:
manifests/files.parquet
manifests/files.csv
config.json
common_voice_1000_KINkinetics-700-2020LEPkinyarwanda-speech-500h
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.polymarket-updown-microstructure
Format. Three tables are published as parquet (under parquet/) for
the Hub viewer and pandas/polars/datasets users — pick a table from the
config dropdown above. The
honest-backtest loader
reads this parquet/ directory directly via its parquet adapter
(adapters.parquet_pm.load_corpus) for one-command reproduction of the
paper's results — see Reproduce the headline result below.
Dataset card — Polymarket crypto up/down microstructure (5m/15m)
Six weeks of real order-book… See the full description on the dataset page: https://huggingface.co/datasets/kinzikdza/polymarket-updown-microstructure.kinova-handover-0807This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "kinova",
"total_episodes": 44,
"total_frames": 24984,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:44"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/kinova-handover-0807.generative-papers-arxiv
arXiv Citation Embeddings Dataset
This dataset contains citation embeddings for arXiv papers, designed for training models that predict cited paper embeddings from citation context.
Dataset Structure
Core Files (Required for Training)
paper_embeddings.parquet: Paper-level sentence embeddings (shared across splits)
train/citations.jsonl: Training citation pairs with contexts
train/citation_embeddings_*.parquet: Citation context embeddings (sharded)… See the full description on the dataset page: https://huggingface.co/datasets/akalmbach-kinsol/generative-papers-arxiv.my_pushtThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 206,
"total_frames": 25650,
"total_tasks": 1,
"total_videos": 206,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:206"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kinam0252/my_pusht.peft-factorykinyarwanda-asr-track-akinova-handover-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "kinova",
"total_episodes": 22,
"total_frames": 13512,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:22"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/kinova-handover-v30.pick_saladThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 6423,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kinghanse/pick_salad.kinova-handover2-v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "kinova",
"total_episodes": 22,
"total_frames": 21164,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:22"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/kinova-handover2-v30.white_chess_king_pick_place_bboxes
white_chess_king_pick_place
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
grab_saladThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 31338,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kinghanse/grab_salad.Fnii-VLA-Kinova-1.0
