datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ruri-dataset-v2-ptWIP: 正式公開準備中
各データセットのライセンスは元データセットに従います。
ruri-dataset-reranker
Ruri-Dataset Reranker
Datasets used for training Ruri-Reranker.
Please refer to https://huggingface.co/datasets/hpprc/emb for individual datasets.
MDSTestEEG_records_raw_schizophrenia_bipolarnagekinoboureiwaintaishitai
Bangumi Image Base of Nageki No Bourei Wa Intai Shitai
This is the image base of bangumi Nageki no Bourei wa Intai shitai, we detected 78 characters, 5673 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/nagekinoboureiwaintaishitai.Knowledge_distilled_dataset_by_NAGI将棋AI用の知識蒸留済みのデータセットを公開します。およそ80億局面あります。 nodchip氏が公開しているtanuki-.nnue-pytorch-2024-07-30.1をhaoでqsearchシャッフルしたのち自作のNAGI(非公開)で評価値を書き換えました。Eval_Coef=600でDLモデルのvalueと評価値を変換しています。 データにバグがあるかもしれませんが、品質保証はしません。
https://huggingface.co/datasets/nodchip/tanuki-.nnue-pytorch-2024-07-30.1
UnCLIPImageInterpolationSamplesne-asr-dataset-nag-aug
NE ASR Augmented Dataset -- Nagamese (nag)
Augmented automatic speech recognition dataset for Nagamese (nag),
a Assamese-based creole language spoken in Nagaland, India.
Source
Augmented from sulabhkatiyar/ne-asr-nag
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Nagamese
ISO 639-3
nag
Family
Assamese-based creole
Region
Nagaland, India
Tonal
No
Tier
D (23.76h… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nag-aug.UnCLIPTextInterpolationSamplesTestingDataset
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/Naga1289/TestingDataset.Raw_dataset_by_NAGISA_V4
Raw dataset by NAGISA V4
Raw shogi teacher positions generated by attic-gensfen with the NAGISA V4
HalfKA-2304 evaluation function. This is the unprocessed source corpus behind
qleap/Knowledge_distilled_dataset_by_NAGISA_V4: no deduplication, no
shuffling, no filtering. 14,101,395 games, 1,419,944,833 positions.
Format
The corpus is one stream of fixed-width 60-byte records, split into chunks of
~10,000,000 positions. Chunks are cut at game boundaries, so each… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Raw_dataset_by_NAGISA_V4.ruri-v3-dataset-rerankerCreated from hpprc/reranker-scores.
We found that cleaning up noisy positives and negatives in our existing dataset using rerankers' scores had a massive impact on performance.
Concretely:
We averaged the scores from five off‑the‑shelf reranker models.
For "positive" examples (documents that contain the answer string for a given query), we only kept those with an average score ≥ 0.3.
For "negative" examples (documents that do not contain the answer string), we only kept those with an average… See the full description on the dataset page: https://huggingface.co/datasets/cl-nagoya/ruri-v3-dataset-reranker.MPK1_dataset_by_NAGISA_V4
MPK1 dataset by NAGISA V4
Shogi teacher positions generated by attic-gensfen with the NAGISA V4
HalfKA-2304 evaluation function, carried in MPK1 — one record per searched
position, before duplicate boards are folded into one. This is the input behind
qleap/Knowledge_distilled_dataset_by_NAGISA_V4: no folding, no deduplication,
no policy normalisation, no shuffling. 14,101,066 games, 1,372,612,150 positions.
This is not the whole corpus its generator wrote.… See the full description on the dataset page: https://huggingface.co/datasets/qleap/MPK1_dataset_by_NAGISA_V4.ruri-dataset-ft
Ruri-Dataset FT
Datasets used for fine-tuning Ruri.
Please refer to https://huggingface.co/datasets/hpprc/emb for individual datasets.
CheckpointMergerSamplesruri-dataset-v2-ftnaginoasukara
Bangumi Image Base of Nagi No Asukara
This is the image base of bangumi Nagi no Asukara, we detected 23 characters, 3162 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/naginoasukara.MPK1_dataset_by_NAGISA_V3
MPK1 dataset by NAGISA V3
Shogi teacher positions from self-play of the search engine attic reading the
NNUE weights NAGISA_V3, carried in MPK1 — the raw record stream, one
record per searched position, before duplicate boards are folded into one. This
is the unprocessed input behind
qleap/Knowledge_distilled_dataset_by_NAGISA_V3: no folding, no deduplication,
no policy normalisation, no shuffling beyond the one the games arrived in.
2,858,094 games, 300,045,733 positions.… See the full description on the dataset page: https://huggingface.co/datasets/qleap/MPK1_dataset_by_NAGISA_V3.latentsne-asr-dataset-nag
Nagamese (nag) — ASR dataset
A small Nagamese (nag) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
12,862
validation
1,532
test
1,717
Data fields
Each example has:… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nag.nagi_fireemblem
Dataset of nagi (Fire Emblem)
This is the dataset of nagi (Fire Emblem), containing 33 images and their tags.
The core tags of this character are green_hair, long_hair, green_eyes, pointy_ears, breasts, very_long_hair, large_breasts, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images
Size
Download… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/nagi_fireemblem.Knowledge_distilled_dataset_by_NAGISA_V3
NAGISA_V3 teacher shards
Training data distilled from self-play of the search engine attic reading
the NNUE weights NAGISA_V3, in the shape the trainers read directly.
Starting from a balanced-opening book, the engine played itself at MultiPV=5,
recording the root score and the candidate moves at every ply;
manaka-teacher turned that corpus into raw MPK1 streams, and manaka-pack
folded identical positions into one row each and wrote these parquet
shards.
Rows: 280,621,202… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Knowledge_distilled_dataset_by_NAGISA_V3.naga_fireemblem
Dataset of naga (Fire Emblem)
This is the dataset of naga (Fire Emblem), containing 30 images and their tags.
The core tags of this character are long_hair, pointy_ears, green_hair, breasts, green_eyes, very_long_hair, large_breasts, medium_breasts, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/naga_fireemblem.so-101_dataset02_20260827_130357This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/nagaenaga/so-101_dataset02_20260827_130357.idiom-vision-fooling
Idioms in Misleading Visual Context
A small, densely-annotated multimodal benchmark testing whether a misleading image can push a
vision-language model toward the wrong reading of a potentially idiomatic phrase, while human
annotators stay unaffected.
Each example pairs a sentence containing a potentially idiomatic expression with an image. The
image either matches the sentence's intended reading (aligned) or depicts the opposite
reading (misleading). Annotators label how… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/idiom-vision-fooling.SO-101_dataset03_20260827_145444This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/nagaenaga/SO-101_dataset03_20260827_145444.nagao_kagetora_fgo
Dataset of nagao_kagetora/長尾景虎/长尾景虎 (Fate/Grand Order)
This is the dataset of nagao_kagetora/長尾景虎/长尾景虎 (Fate/Grand Order), containing 350 images and their tags.
The core tags of this character are white_hair, multicolored_hair, black_hair, two-tone_hair, long_hair, hair_between_eyes, breasts, very_long_hair, streaked_hair, yellow_eyes, medium_breasts, green_eyes, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/nagao_kagetora_fgo.saccade-egomotion-bench
Saccade ego-motion benchmark
The stream, the raw decision signals, and the per-frame measurements behind
Saccade — an always-on edge VLM that re-encodes only
the image patches whose change ego-motion cannot explain.
This dataset exists so the central claim can be checked without running our code.
💻 Code: https://github.com/NagaYu/saccade
🤖 Model: https://huggingface.co/NagaYu/saccade-predictor
🚀 Demo: https://huggingface.co/spaces/NagaYu/saccade
The claim, in… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/saccade-egomotion-bench.Training_dataset_by_NAGISA_V4
NAGISA_V4 ply-37 teacher shards, games played to the end
Self-play of attic-gensfen reading the NNUE weights NAGISA_V4
(HalfKA-2304), in the shape the trainers read directly. Every game starts from a
balanced ply-37 position, makes no random moves, and runs until it actually
ends. Identical positions are folded into one row each.
67,108,864 rows — exactly 2^26
16 shards of 4,194,304 rows, 512 row groups each, zstd, 5,208,620,326 B total
Against
Opening_dataset_by_NAGISA_V4… See the full description on the dataset page: https://huggingface.co/datasets/qleap/Training_dataset_by_NAGISA_V4.brain-memory
🧠 NIFTY AI Agent: Memory OS Cloud Snapshot
Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS.
• Repository: nagarhimanshu37/brain-memory• Total Stored Records: 235• Last Synchronized: 2026-09-25 17:46:58 UTC
📊 Partition Statistics
Partition
Records
Description
conversation_memory
83
Multi-turn trader dialogues & intent logs
episodic_memory
50
Trading day episodes (facts vs interpretations)
experience_memory
50
Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.
