datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mierukochan
Bangumi Image Base of Mieruko-chan
This is the image base of bangumi Mieruko-chan, we detected 52 characters, 3778 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mierukochan.mieb-train
mieb-train
CLIP training data built from the MIEB
zero-shot image-classification datasets, with the text side re-written so the
class labels are usable as captions.
4,434,196 training rows across 23 datasets.
column
notes
image
passed through from the source repo, not re-encoded
text
the caption — this is what CLIP trains on
label
source class id (-1 where the source has no class list)
class_name
normalized class name
dataset_name
source dataset key — filter… See the full description on the dataset page: https://huggingface.co/datasets/PumeTu/mieb-train.github-repos-metadata-ge3
GitHub All Repositories Metadata Dataset (>= 3 Stars)
A comprehensive metadata dataset covering 5,289,726 public GitHub repositories with 3 or more stars (>= 3) spanning the history of GitHub from 2008 to 2026.
Data Recency & Snapshot Notice
[!NOTE]
Snapshot Methodology: This dataset combines a comprehensive historical base archive (up to mid-2022) with continuous periodic crawler snapshots (2023 through 2026).
Star Counts & Metrics: Star counts and repository… See the full description on the dataset page: https://huggingface.co/datasets/Mieaz/github-repos-metadata-ge3.MieDB-100k
MieDB-100k: A Comprehensive Dataset for Medical Image Editing
📄 Introduction
MieDB-100k is a large-scale, high-quality and diverse dataset for text-guided medical image editing,
which includes 104,267 editing data, covering 63 distinct editing targets and 10 diverse medical image modalities.
We categorize editing tasks into three types: Perception, Modification and Transformation, which consider both model's intrinsic understanding and generation abilities on medical… See the full description on the dataset page: https://huggingface.co/datasets/Laiyf/MieDB-100k.Mielikki-Erebus_87k-ShareGPT
I wouldnt use this set as is or the original set until it is updated further.
Original Dataset: https://huggingface.co/datasets/Mielikki/Erebus-87k
Converted, fixed punctuation, using: https://github.com/The-Chaotic-Neutrals/ShareGPT-Formaxxing
To do:
(add instructions) Deslop, deduplicate, rejection filter, grammar correct.
Dataset contains:
87k human story submissions, ranging across these categories-
[adventure
western
urban-fantasy
thriller
suspense
speculative… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Mielikki-Erebus_87k-ShareGPT.github-issuesErebus-87k87k human story submissions, ranging across these categories:
adventure
western
urban-fantasy
thriller
suspense
speculative
science-fiction
sad
romance
mystery
middle-school
inspirational
horror
holiday
historical-fiction
high-school
happy
funny
friendship
fiction
fantasy
drama
crime
creative-nonfiction
contemporary
coming-of-age
christmas
christian
bedtime
american
mie-curriculum-sft-grade9
MIE Curriculum SFT Grade 9
This dataset contains answer-only supervised fine-tuning rows for Grade 9 curriculum tutoring.
The data is provided as a single JSONL file:
all_stamped.jsonl
Dataset Summary
Grade: 9
Language: English
Rows: 42,771
Subjects: 19
KB-backed subunits covered: 509
Format: instruction/input/output JSONL
Row Format
Each training row contains:
id
grade
subject
unit_id
subunit_id
instruction
input
output
Use these column… See the full description on the dataset page: https://huggingface.co/datasets/roshans89/mie-curriculum-sft-grade9.MI-errorsMielikki_Erebus-87k-axoMIETICretrogames-forum-polishhumanoid-mieayam-datatest13aug16thaugmietrecht-qa-dataset16augv216thaugv5PeruConsul
