datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images
HSTLI: A Dataset of Human Semen Time Lapse Images
Dataset Details
HSTLI contains 3,266 time-lapse microscopy videos of human sperm.Clips were recorded from two imaging modalities:
CASA system (Sperm Class Analyzer)
Optical microscope (Swift M10DB-MP + Fujifilm X-T30)
A subset of videos was manually annotated with bounding boxes around each visible sperm head.
The dataset supports detection, tracking and motility computation.
Total contents:
34… See the full description on the dataset page: https://huggingface.co/datasets/DFL-KamLab/HSTLI_A-Dataset-of-Human-Semen-Time-Lapse-Images.Kamba-ASR-Data-Subset-484H
Kamba ASR Data Subset 484H
Kamba speech dataset for automatic speech recognition.
TinyStories-Algerian-Darijakamuicode-i2i-images
AI画像編集モデル 総合ベンチマーク v5.4
概要
本ディレクトリは、AI画像編集モデルの性能を総合的に評価するためのベンチマークスイートです。
47種類のテストを4つの難度レベルに分類し、65種類のモデル(旧45モデル + 新20モデル)を比較評価しています。
評価カバレッジ
項目
旧モデル群 (35)
新モデル群 (20)
合計
モデル数
45 (I2I/R2I 22 + T2I 13 + その他 10)
20
65
テスト数
34
47
47
生成画像
3,446
1,301 / 2,403
4,747+
VLM評価済み
完了
1,301 (54%)
進行中
評価方式
VLM-as-a-Judge: Gemini 2.5 Flash による5軸評価(Structure / Identity / Reasoning / Instruction / Quality)
スタイル多様性: photo, anime_flat, anime_cg… See the full description on the dataset page: https://huggingface.co/datasets/yumenojmd/kamuicode-i2i-images.Algerian-Youtube-Commentskamitachinihirowaretaotoko2ndseason
Bangumi Image Base of Kami-tachi Ni Hirowareta Otoko 2nd Season
This is the image base of bangumi Kami-tachi ni Hirowareta Otoko 2nd Season, we detected 75 characters, 4264 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamitachinihirowaretaotoko2ndseason.kaminotou
Bangumi Image Base of Kami No Tou
This is the image base of bangumi Kami no Tou, we detected 49 characters, 3102 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminotou.kamisamakiss
Bangumi Image Base of Kamisama Kiss
This is the image base of bangumi Kamisama Kiss, we detected 50 characters, 2686 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamisamakiss.kamonohashironnokindansuiri
Bangumi Image Base of Kamonohashi Ron No Kindan Suiri
This is the image base of bangumi Kamonohashi Ron no Kindan Suiri, we detected 36 characters, 4169 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamonohashironnokindansuiri.kamitachinihirowaretaotoko
Bangumi Image Base of Kami-tachi Ni Hirowareta Otoko
This is the image base of bangumi Kami-tachi ni Hirowareta Otoko, we detected 75 characters, 4446 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamitachinihirowaretaotoko.kamiwagameniueteiru
Bangumi Image Base of Kami Wa Game Ni Ueteiru
This is the image base of bangumi Kami wa Game ni Ueteiru, we detected 68 characters, 4275 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamiwagameniueteiru.kamikazekaitoujeanne
Bangumi Image Base of Kamikaze Kaitou Jeanne
This is the image base of bangumi Kamikaze Kaitou Jeanne, we detected 43 characters, 3600 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kamikazekaitoujeanne.aes_enem_dataset
Automated Essay Score (AES) ENEM Dataset
Use Case and Creators
Intended Use: Estimate Essay Score
Creators: Igor Cataneo Silveira, André Barbosa and Denis Deratani Mauá
Contact Information: igorcs@ime.usp.br; andre.barbosa@ime.usp.br
Licensing Information
License: MIT License
Citation Details
Preferred Citation:
@proceedings{DBLP:conf/propor/2024,
editor = {Igor Cataneo Silveira, André Barbosa and Denis Deratani Mauá},
title =… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/aes_enem_dataset.deutsche-bahn-data
Deutsche Bahn Train Data
This dataset contains public historical data from Deutsche Bahn, the largest German train company. It includes train schedules, delays, and cancellations from stations across Germany.
For more info visit the project page at GitHub: https://github.com/piebro/deutsche-bahn-data
Dataset Structure
Monthly Processed Data
The monthly processed data is located in monthly_processed_data/ and contains files named data-YYYY-MM.parquet.
Schema:… See the full description on the dataset page: https://huggingface.co/datasets/kamalnsr123456/deutsche-bahn-data.algerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.cv_for_spd_fr_syntheticamiagentvidbench
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline)
by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI)
Layout
.
├── README.md
├── questions.jsonl # 100 rows — one per question
├── videos.jsonl # 71 rows — one per unique video
├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.jbcs2025_experiments_report
JBCS 2025: Experimental Artefacts for AES in Brazilian Portuguese
This repository contains all experimental artefacts (logs, configurations, predictions, and evaluation results) described in the paper:
Exploring the Usage of LLMs for Automatic Essay Scoring in Brazilian Portuguese EssaysAndré Barbosa, Igor Cataneo Silveira, Denis Deratani MauáTODO
📦 What's in this dataset repo?
This dataset is not a training dataset. Instead, it provides comprehensive logs and… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/jbcs2025_experiments_report.forest-fire-annotations
Forest Fire Detection Dataset — Auto-Annotated
Bounding-box annotated version of touati-kamel/forest-fire-dataset,
built for training forest-fire / smoke / fog object detection models.
Overview
This dataset contains video frames auto-labeled with bounding boxes for fire and
smoke-related visual phenomena, using a zero-shot open-vocabulary object detector
(Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image
classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.bangla-instruction-dataset
🧠 Bangla Instruction Dataset
This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models.
📚 Dataset Splits
The dataset is organized into the following splits:
Split Name
Source Dataset
Description
OdiaGenAI
OdiaGenAI/all_combined_bengali_252k
A large-scale collection of diverse Bangla instructions and responses.
chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.reddit-mental-health-classificationmixture_ami_synthetic_bigkaminakisekainokamisamakatsudou
Bangumi Image Base of Kaminaki Sekai No Kamisama Katsudou
This is the image base of bangumi Kaminaki Sekai no Kamisama Katsudou, we detected 72 characters, 5020 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kaminakisekainokamisamakatsudou.medium_articles
Dataset Card for "medium_articles"
More Information needed
Persian-conversational-datasetpersian-conversational-datasetStock-AAPLcommit-messages-datasetDiseases_Dataset
🩺 Diseases Dataset
A consolidated medical dataset combining disease names, symptoms, and treatments collected from multiple public datasets across Hugging Face and Kaggle. This dataset can be used for building disease prediction, symptom clustering, and medical assistant models.
📦 Dataset Summary
Field
Type
Description
Disease
string
Name of the disease or condition
Symptoms
string
List of symptoms or Description of symptomps
Treatments
string
(Optional)… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/Diseases_Dataset.Killer
Getting Started
RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text
documents coming from 84 CommonCrawl snapshots and processed using
the CCNet pipeline. Out of these, there are 30B documents in the corpus
that additionally come with quality signals. In addition, we also provide the ids of duplicated documents which can be
used to create a dataset with 20B deduplicated documents.
Check out our blog post for more details on the… See the full description on the dataset page: https://huggingface.co/datasets/KamDickGoon/Killer.
