datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlux-researchinception-v1-microscope-data
Inception V1 Microscope Data
This dataset powers the Inception V1
Microscope,
an interactive interface for exploring visual features learned by individual
neurons in Inception V1.
It combines two complementary interpretability views:
Activation maximization: one synthesized visualization optimized to
strongly activate each neuron.
Top dataset examples: the ten ImageNet examples producing the strongest
recorded activations for each neuron, paired with crops associated with the… See the full description on the dataset page: https://huggingface.co/datasets/akankshanc/inception-v1-microscope-data.TACO-Benchmark
TACO-Benchmark
TACO (Text-to-SQL with Ambiguous and Cross-database Open-domain queries) is a benchmark for real-world data-lake Text-to-SQL.
📢 News (2026): TACO has been accepted to VLDB 2026! 🎉📄 Paper: arXiv:2606.14201
GitHub (code & evaluation): Akanezora0/TACO-Benchmark
Google Drive mirror: TACO-Benchmark.zip
Overview
Unlike Spider or BIRD — where the target database is known and schemas are clean — TACO evaluates systems on three challenges common in… See the full description on the dataset page: https://huggingface.co/datasets/Akanezora/TACO-Benchmark.offsec-400just change the dataset and use it as u want I just created it for training and checking things
cuda-error-resolution-analysisDrivingVQA-conflictcounterfactual-pendulum-multilingual
📌 Dataset Summary
When a Vision-Language Model (VLM) is given an image along with a text prompt containing contradictory or misleading information, how does it react? Does it rely on the visual evidence, succumb to textual bias, or honestly abstain when faced with unresolvable conflict?
This dataset adapts the Counterfactual Pendulum scenario across two visual conflict dimensions:
Angular (Angle): Conflict in the pendulum's angle of inclination.
Light: Conflict in the light… See the full description on the dataset page: https://huggingface.co/datasets/akanshjain37/counterfactual-pendulum-multilingual.nonohara_akane_theidolmstermillionlive
Dataset of nonohara_akane/野々原茜/노노하라아카네 (THE iDOLM@STER: Million Live!)
This is the dataset of nonohara_akane/野々原茜/노노하라아카네 (THE iDOLM@STER: Million Live!), containing 127 images and their tags.
The core tags of this character are short_hair, brown_hair, brown_eyes, bangs, red_hair, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/nonohara_akane_theidolmstermillionlive.saas-erp-support-ru-train-v13
Capstone support: F training shards
Train-only synthetic Russian SaaS / fictional ERP / telecom scenarios:
430 dialogues, 1113 target turns, 13 local JSONL shards. No real company or customer data.
No validation/test split is published here. Files preserve project-relative paths
under data/quality90_v1/train; restore those paths in a checkout to reuse them.
Each dialogue contains context and turn-level target labels/replies. The exact
file list and SHA-256 hashes are in… See the full description on the dataset page: https://huggingface.co/datasets/AkanaYB/saas-erp-support-ru-train-v13.akane_bluearchive
Dataset of akane/室笠アカネ/朱音 (Blue Archive)
This is the dataset of akane/室笠アカネ/朱音 (Blue Archive), containing 500 images and their tags.
The core tags of this character are long_hair, breasts, large_breasts, halo, glasses, hair_between_eyes, brown_eyes, light_brown_hair, animal_ears, bow, fake_animal_ears, rabbit_ears, black-framed_eyewear, blue_bow, brown_hair, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/akane_bluearchive.oasst1-akankurokawa_akane_oshinoko
Dataset of Kurokawa Akane
This is the dataset of Kurokawa Akane, containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
200
Download
Raw data with meta information.
raw-stage3
408
Download
3-stage cropped raw data with meta information.
384x512
200
Download
384x512 aligned dataset.
512x512
200… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/kurokawa_akane_oshinoko.shinoda_akane_nonnonbiyori
Dataset of Shinoda Akane
This is the dataset of Shinoda Akane, containing 82 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
82
Download
Raw data with meta information.
raw-stage3
205
Download
3-stage cropped raw data with meta information.
raw-stage3-eyes
233
Download
3-stage cropped (with eye-focus) raw… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/shinoda_akane_nonnonbiyori.Akan_non_standardspeechmath-visual-multihop-qaakan-speech-collect
Akan Speech Collection
Read speech in Akan with transcripts, for training speech recognition.
akan_audioThis work is comprised of audio data of Twi, a low resourced language spoken by the Akan people in Ghana.
This has been adapted by NLPGhana.math-visual-qa-cuakan_tts_datasetAkananuru
Akananuru (அகநானூறு)
Summary
Akananuru (அகநானூறு) is one of the classical Tamil Sangam literature anthologies consisting of 400 poems. It belongs to the Ettuthokai (Eight Anthologies) corpus and focuses on the inner (அகம்) themes of life such as love, emotions, and human experiences.
This dataset provides each poem along with metadata such as title, note (explanation/context), and poet information. It is useful for NLP research, classical text analysis, poetry generation… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/Akananuru.ugspeech-akan-clean-100hrspeople_daily_news
人民日报(1946-2025)数据集
The dataset is part of CialloCorpus, available at https://github.com/prnake/CialloCorpus
akan-umbundu_sentence-pairs
Akan-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Umbundu_Sentence-Pairs
Number of Rows: 22651
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-umbundu_sentence-pairs.eimura_akane_ahogirl
Dataset of Eimura Akane
This is the dataset of Eimura Akane, containing 89 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
89
Download
Raw data with meta information.
raw-stage3
185
Download
3-stage cropped raw data with meta information.
384x512
89
Download
384x512 aligned dataset.
512x512
89
Download… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/eimura_akane_ahogirl.akan-english-emotions-corpus
Akan-english Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Akan-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-english-emotions-corpus.bank-customer-churnakane_bluearchive
Dataset of akane (Blue Archive)
This is the dataset of akane (Blue Archive), containing 498 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
This is a WebUI contains crawlers and other thing: (LittleAppleWebUI)
Name
Images
Download
Description
raw
498
Download
Raw data with meta information.
raw-stage3
1352
Download
3-stage cropped raw data with… See the full description on the dataset page: https://huggingface.co/datasets/AppleHarem/akane_bluearchive.exploitdataakan-tts-wavtokenizer-combined-v2Akan_Llama_Threads
