datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images
HSTLI: A Dataset of Human Semen Time Lapse Images
Dataset Details
HSTLI contains 3,266 time-lapse microscopy videos of human sperm.Clips were recorded from two imaging modalities:
CASA system (Sperm Class Analyzer)
Optical microscope (Swift M10DB-MP + Fujifilm X-T30)
A subset of videos was manually annotated with bounding boxes around each visible sperm head.
The dataset supports detection, tracking and motility computation.
Total contents:
34… See the full description on the dataset page: https://huggingface.co/datasets/DFL-KamLab/HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images.transformers_image_dockaminotoukoubousen
Bangumi Image Base of Kami No Tou: Koubou-sen
This is the image base of bangumi Kami no Tou: Koubou-sen, we detected 110 characters, 9146 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminotoukoubousen.Kamba-ASR-Data-Subset-484H
Kamba ASR Data Subset 484H
Kamba speech dataset for automatic speech recognition.
TinyStories-Algerian-DarijaMA_Query_Expansion_MLT26kamuicode-i2i-images
AI画像編集モデル 総合ベンチマーク v5.4
概要
本ディレクトリは、AI画像編集モデルの性能を総合的に評価するためのベンチマークスイートです。
47種類のテストを4つの難度レベルに分類し、65種類のモデル(旧45モデル + 新20モデル)を比較評価しています。
評価カバレッジ
項目
旧モデル群 (35)
新モデル群 (20)
合計
モデル数
45 (I2I/R2I 22 + T2I 13 + その他 10)
20
65
テスト数
34
47
47
生成画像
3,446
1,301 / 2,403
4,747+
VLM評価済み
完了
1,301 (54%)
進行中
評価方式
VLM-as-a-Judge: Gemini 2.5 Flash による5軸評価(Structure / Identity / Reasoning / Instruction / Quality)
スタイル多様性: photo, anime_flat, anime_cg… See the full description on the dataset page: https://huggingface.co/datasets/yumenojmd/kamuicode-i2i-images.Algerian-Youtube-Commentskamitachinihirowaretaotoko2ndseason
Bangumi Image Base of Kami-tachi Ni Hirowareta Otoko 2nd Season
This is the image base of bangumi Kami-tachi ni Hirowareta Otoko 2nd Season, we detected 75 characters, 4264 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamitachinihirowaretaotoko2ndseason.kaminomizoshirusekai
Bangumi Image Base of Kami Nomi Zo Shiru Sekai
This is the image base of bangumi Kami Nomi zo Shiru Sekai, we detected 60 characters, 5684 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminomizoshirusekai.kaminotou
Bangumi Image Base of Kami No Tou
This is the image base of bangumi Kami no Tou, we detected 49 characters, 3102 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminotou.kamisamakiss
Bangumi Image Base of Kamisama Kiss
This is the image base of bangumi Kamisama Kiss, we detected 50 characters, 2686 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamisamakiss.kamonohashironnokindansuiri
Bangumi Image Base of Kamonohashi Ron No Kindan Suiri
This is the image base of bangumi Kamonohashi Ron no Kindan Suiri, we detected 36 characters, 4169 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamonohashironnokindansuiri.pcod_graph_idskamitachinihirowaretaotoko
Bangumi Image Base of Kami-tachi Ni Hirowareta Otoko
This is the image base of bangumi Kami-tachi ni Hirowareta Otoko, we detected 75 characters, 4446 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamitachinihirowaretaotoko.kamiwagameniueteiru
Bangumi Image Base of Kami Wa Game Ni Ueteiru
This is the image base of bangumi Kami wa Game ni Ueteiru, we detected 68 characters, 4275 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamiwagameniueteiru.postcrossing-daily-growthseta-env-final-filteredkamikazekaitoujeanne
Bangumi Image Base of Kamikaze Kaitou Jeanne
This is the image base of bangumi Kamikaze Kaitou Jeanne, we detected 43 characters, 3600 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamikazekaitoujeanne.aes_enem_dataset
Automated Essay Score (AES) ENEM Dataset
Use Case and Creators
Intended Use: Estimate Essay Score
Creators: Igor Cataneo Silveira, André Barbosa and Denis Deratani Mauá
Contact Information: igorcs@ime.usp.br; andre.barbosa@ime.usp.br
Licensing Information
License: MIT License
Citation Details
Preferred Citation:
@proceedings{DBLP:conf/propor/2024,
editor = {Igor Cataneo Silveira, André Barbosa and Denis Deratani Mauá},
title =… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/aes_enem_dataset.kamierabi
Bangumi Image Base of Kamierabi
This is the image base of bangumi Kamierabi, we detected 55 characters, 5241 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamierabi.2wikimultihopqa\ai-chaos-theory-heatdeutsche-bahn-data
Deutsche Bahn Train Data
This dataset contains public historical data from Deutsche Bahn, the largest German train company. It includes train schedules, delays, and cancellations from stations across Germany.
For more info visit the project page at GitHub: https://github.com/piebro/deutsche-bahn-data
Dataset Structure
Monthly Processed Data
The monthly processed data is located in monthly_processed_data/ and contains files named data-YYYY-MM.parquet.
Schema:… See the full description on the dataset page: https://huggingface.co/datasets/kamalnsr123456/deutsche-bahn-data.cv_for_spd_fr_processedaflow_graph_idsalgerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.cv_for_spd_fr_syntheticami_spd_augmented_test2_processedami
