datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagefolder_with_metadataaudiofolder_two_configs_in_metadata_with_defaultimagefolder_with_metadata_no_splitsc4-en-html-with-metadatacivitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/civitai-top-nsfw-images-with-metadata.5000-podcast-conversations-with-metadata-and-embedding-dataset
🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset
ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications.
This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network.
AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.c4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.c4-en-html-with-training_metadata_alldataset_cards_with_metadatamodel_cards_with_metadataswe_gym_with_metadata_etashswebench_test_with_pr_metadataworldcuisines_format_sea_country_only_with_metadatacivitai-top-sfw-images-with-metadata
CivitAI Top SFW Images Dataset
This dataset contains 12k+ top SFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-SFW-Images-Dataset-12k-images
Dataset collection date
July 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpeg
│ ├── 2.jpeg
│ ├── 3.jpeg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/civitai-top-sfw-images-with-metadata.civitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/pasindu29/civitai-top-nsfw-images-with-metadata.real-cultural-relic-images-with-metadata
真实文物图像数据集
本数据集包含 10 万张真实文物及艺术藏品图片,本次公开1000张,以及与图片一一对应的 JSON 元信息文件。数据覆盖武器与盔甲、绘画、雕塑、装饰艺术等多种藏品类型,每张图片均提供藏品题名、类别、年代、创作者、文化背景、材质、馆藏来源及许可证等信息,并使用 SHA-256 哈希值辅助文件校验与去重。
当前数据集中图片短边尺寸为 200~2574 像素。全部 JSON 文件均已通过格式解析检查。
数据集用途
本数据集可用于文物与艺术品图像分类、藏品类型识别、年代与文化背景研究、图文检索、多模态模型训练与评估、数字博物馆应用及计算机视觉教学等场景。
文件结构
Cultural Relic Images and Metadata/
├── cma_100001.jpg # 文物或艺术藏品图片
├── cma_100001.json # 与图片同名的 JSON 元信息
├── cma_100022.jpg
├── cma_100022.json
└── ...
图片与… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-cultural-relic-images-with-metadata.audiofolder_two_configs_in_metadata_with_defaultcivitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/Avanish11/civitai-top-nsfw-images-with-metadata.rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
real-dog-images-with-breed-and-metadata
真实犬类图像测试数据集
本数据集包含10w张真实犬类图片,本次已上传1000张,以及与这1000张图片一一对应的 JSON 元信息文件。所有图片像素短边均不低于 512,并经过哈希去重和完整性校验。
数据集可用于犬类图像分类、犬种识别、狗数量统计、图像质量研究、计算机视觉教学及模型测试。
文件结构
dog_dataset/
├── images/ # 图片文件
├── labels/ # 每张图片对应的 JSON 元信息
├── manifest.jsonl # 数据汇总清单
├── audit_report.json # 完整性检查结果
└── README.md # 数据集说明
元信息字段
每个 JSON 文件包含:
id:唯一编号;
image_file:图片文件名;
sha256:图片哈希;
contains_dog:是否包含狗;
dog_count:狗的数量;
environment:室内、室外或未知;… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-dog-images-with-breed-and-metadata.pmc_derma_VQA_with_metadata_processeddatasets_with_metadata_and_summariesreal-vehicle-images-with-visual-metadata-100k
真实车辆图片与视觉元信息测试数据集
数据集简介
本数据集可扩展至 10 万张真实车辆图片,本次上传人工筛选后的 1,000 张测试数据,覆盖轿车、公交车、客车和卡车,可用于车辆图像分类、目标检测、视觉识别、多模态训练及相关算法测试。
当前版本为测试数据集。如需更多图片、指定车辆类别或定制元信息,可扩展量级,欢迎通过邮件联系:xiehan@mykj.club。
数据概况
图片数量:1,000 张
图片格式:JPEG
最小短边:513px,全部不低于 512px
类别分布:109 张轿车、442 张公交车或客车、449 张卡车
文件命名:每张图片使用唯一文件名
内容去重:使用 SHA-256 进行内容级去重
标注形式:每张图片对应一个同名 JSON 文件
目录结构
vehicle_dataset_1000/
├── README.md
├── images/ # 车辆图片
│ └── vehicle_xxx.jpg
├──… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-vehicle-images-with-visual-metadata-100k.real-plant-images-with-species-and-metadata
真实植物图像数据集
本数据集包含 10 万张真实植物图片,本次公开1000张,以及与图片一一对应的 JSON 元信息文件。数据覆盖多种植物物种,每张图片均提供物种名称、图像尺寸、质量信息、来源链接、许可证及署名信息,并使用 SHA-256 哈希值辅助文件校验与去重。
当前数据集中图片短边尺寸为 419~1024 像素。全部 JSON 文件均已通过格式解析检查。
数据集用途
本数据集可用于植物图像分类、植物物种识别、生物多样性研究、植物图像检索、计算机视觉模型训练与评估、多模态学习、数据清洗及教学演示等场景。
文件结构
Plant images and metadata/
├── 0001_394891155.jpg # 植物图片
├── 0001_394891155.json # 与图片同名的 JSON 元信息
├── 0002_394892202.jpg
├── 0002_394892202.json
└── ...
图片与 JSON 文件使用相同的文件名主体,可通过… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-plant-images-with-species-and-metadata.meta_data_annual_reports_tokenized_llama3_8b_with_logged_return_matrixreal-animal-images-with-species-and-metadata
真实动物图像数据集
本数据集包含 10 万张真实动物图片,本次公开1000张,以及与图片一一对应的 JSON 元信息文件。数据覆盖多种动物物种,每张图片均提供物种名称、图像尺寸、质量信息、来源链接、许可证及署名信息,并使用 SHA-256 哈希值辅助文件校验与去重。
当前数据集中图片短边尺寸为 213~1024 像素。全部 JSON 文件均已通过格式解析检查。
数据集用途
本数据集可用于动物图像分类、物种识别、生物多样性研究、图像检索、计算机视觉模型训练与评估、多模态学习、数据清洗及教学演示等场景。
文件结构
Animal images and metadata/
├── 0001_394885493.jpg # 动物图片
├── 0001_394885493.json # 与图片同名的 JSON 元信息
├── 0002_394885471.jpg
├── 0002_394885471.json
└── ...
图片与 JSON 文件使用相同的文件名主体,可通过… See the full description on the dataset page: https://huggingface.co/datasets/MYtechnology/real-animal-images-with-species-and-metadata.dhivehi-english-pairs-with-metadata
Dhivehi–English Parallel Pairs with Metadata
The dhivehi-english-pairs-with-metadata dataset contains 56,029 aligned sentence pairs between Dhivehi (written in Thaana script) and English, automatically generated using a gemma-3 language model. Each sentence pair includes additional metadata such as the opinion expressed in the sentence, a general sentiment classification (Positive, Negative, or Neutral), the intent behind the statement (e.g., Inform, Request, Accuse), and its… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-english-pairs-with-metadata.dataset_cards_with_metadata_with_sample_rowspmc_derma_VQA_with_metadatadataset_cards_with_metadata_with_embeddings
