datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
suno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.FullBenchhonesty-index
The Kerne Honesty Index
What each synthetic dollar advertises, next to what it actually paid.
Advertised APY versus realized APY for 21 synthetic dollar vaults, recomputed hourly from
ERC-4626 share price growth on chain, and signed.
The realized column is not taken from anybody's dashboard. It is measured directly from the vault
contract: convertToAssets(10**decimals) read at two block heights, divided by 10**asset_decimals,
annualized over the real elapsed time between those… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/honesty-index.food17-flaggedkhondo
Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms
Real Bangla and English government forms assembled into document packets for the
packet-splitting task: given a packet of concatenated form pages, recover which pages
belong to each source document and restore each document's original page order.
This is the dataset and its schema. The full benchmark pipeline (inference,
evaluation, and analysis) lives in the code repository:… See the full description on the dataset page: https://huggingface.co/datasets/Mausul/khondo.MolPuzzle_data
MolPuzzle: A Multimodal Benchmark for Molecular Structure Elucidation
Dataset Description:
The MolPuzzle dataset is a newly developed resource designed to challenge Large Language Models with multi-modal, multi-step reasoning tasks (molecular structure elucidation). This dataset consists of 217 diverse and intricate structure elucidation challenges that require LLMs to demonstrate advanced reasoning capabilities, integrating multimodal data and deep chemical understanding… See the full description on the dataset page: https://huggingface.co/datasets/kguo2/MolPuzzle_data.LLM4SGG
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
This directory is designed for convenient dataset downloads to facilitate code implementation.
Since LLM4SGG utilizes language supervision in the scene graph generation (SGG) task, we learn models from distinct caption (i.e., image-text pair) datasets and perform inference on the SGG datasets.
We provide enhanced scene graph datasets made by LLM (ChatGPT).
Furthermore, we provide the necessary meta-data for… See the full description on the dataset page: https://huggingface.co/datasets/kb-kim/LLM4SGG.tmp-DL_Keyslicestgk-ai-image-generators-2026
We Tested 10 AI Image Generators on Faces, Text and Ads
Most AI image-generator comparisons reduce the models to a score. We wanted to see the mistakes.
These Guys Know gave ten current models the same three practical briefs in August 2026: a close-up face, exact medical text inside a photographed hospital monitor, and a luxury fragrance advertisement where the person, bottle, label and location needed to look believable together.
We kept the first valid output for every… See the full description on the dataset page: https://huggingface.co/datasets/These-Guys-Know/tgk-ai-image-generators-2026.Anime_UserRatingscyber_drug_dataset
Digital Forensic Investigation Scenario Dataset: Online Drug Trafficking
This dataset is a comprehensive collection of digital artifacts and investigative reports designed for forensic research and education. It simulates a sophisticated Online Drug Trafficking scenario, covering the entire investigation lifecycle from initial intelligence gathering to suspect arrest and financial analysis.
Dataset Structure
The dataset is indexed via a standardized 7-column metadata… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/cyber_drug_dataset.shipwall-products
ShipWall: launched products and the badge check behind each one
One row per product that has launched on the board: what it is, where it lives, what it was built with, and whether the embed badge was actually found on the site it points at.
Rows in this cut
43
One row is
one launched product
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ShipWall
Method
https://toolproof.thecompound.tech/methodology
Licence
Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/shipwall-products.kitgrade-kits
KitGrade: SaaS starter kits and what is measurably in the box
One row per graded SaaS starter kit: its stack, licence and pricing, the release and commit activity behind it, what its own documentation says ships in the box, and the evidence level the grade rests on.
Rows in this cut
36
One row is
one starter kit
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
KitGrade
Method
https://toolproof.thecompound.tech/methodology
Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/kitgrade-kits.Aperdata-SimReady-Kitchen-01
Aperdata-SimReady-Kitchen-01
A curated collection of SimReady 3D assets for World Models and Physical AI research, especially for synthetic data generation, scene composition, embodied AI simulation, and vision-language/robot-learning experiments. Free for non-commercial use.
Dataset Version: 1.0.0
Assets list
Scene:
Kitchen_Indoor
props:
biscuit
coffee_bean
coffee_machine
container_storage
cooking_clay_pot
cooking_pot
cup_coffee
cup_general… See the full description on the dataset page: https://huggingface.co/datasets/Aperdata-tech/Aperdata-SimReady-Kitchen-01.Malayalam-word-freq
Word Frequency Profile of Malayalam
The repo contains Malayalam words and their frequencies as obtained from AI4Bharat Indic NLP corpus.
There is an associated python script to plot the word frequnecy profile.
gmpo_mas3k-compressedCivitAI-As-CharactersDeduplicated set of CivitAI images as searched by SD XL-derived models that have been described by Llava1.6-34b as Characters.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by a Llava1.6-34b. Here is an example:
{
"age": "22",
"eyes": "Bright blue, striking",
"face": "Smooth, elegant, with a gentle expression",
"hair": "Long, straight, brown"… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/CivitAI-As-Characters.character-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes.
Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by Claude Opus. Here is an example:
{… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.Sap_Kush_Med_Deepfake
Sap_Kush_Med_Deepfake Dataset
Paired medical-image forgery lineages across six modalities. Every lineage is
one source image, one mask, one seed: the arms differ only in what was done
inside the mask, so a comparison between arms isolates the manipulation rather
than an encoding artefact.
3956 lineages, 30385 files, 8.49 GiB.
v2 adds a removal arm grounded in human annotation for three more modalities
(endoscopy, ultrasound, MRI). v1 had removal for CT only.
What the… See the full description on the dataset page: https://huggingface.co/datasets/Kanhaiyya/Sap_Kush_Med_Deepfake.keywording2OccDet-Calib
OccDet-Calib
OccDet-Calib is a benchmark and analysis framework for studying detector calibration under partial visibility.
This repo contains
benchmark metadata CSVs
processed evaluation tables
paper figures and montage images
sample public assets
reproducibility-facing artifacts
This repo does not contain
raw COCO images
raw BDD100K images
checkpoints
virtual environments or cache folders
Main conditions
overlap occlusion
distractor control… See the full description on the dataset page: https://huggingface.co/datasets/karthik789338/OccDet-Calib.aiijc-rostelecomasl_signsLearningChat_accounting_ai_questions
공개용 회계입문 AI질문 이미지 데이터셋
이 데이터셋은 회계입문 수업의 AI질문 1회 과제 제출 이미지들을 공개용으로 문서화하기 위해 정리한 메타데이터 패키지다. 현재 폴더에 존재하는 6개 수집 배치 전체를 통합했으며, 공개 버전에서는 학생 실명과 원본 파일명을 직접 노출하지 않도록 비식별 규칙을 적용했다.
본 문서와 함께 제공되는 metadata.csv는 이미지 파일 1개당 1행을 가지는 인벤토리다. 실제 공개 배포 시 이미지 파일은 data/images/AIQ-XXXXXX.ext 형식으로 익명 재배치하는 것을 전제로 한다.
1. 데이터셋 범위
대상 과목: 회계입문
대상 과제: AI질문 1회
포함 범위: 현재 작업 폴더에 있는 6개 수집 배치 전체
레코드 단위: 이미지 파일 1개 = metadata.csv 1행
공개 버전 기준: 완전 비식별 전제
분반별 구성
section_id
익명 제출자 수
이미지 수… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_accounting_ai_questions.gdgood
dataset-openmoji
Dataset OpenMoji
Creator: https://www.kaggle.com/krayc81This is base on https://openmoji.org/License https://creativecommons.org/licenses/by-sa/4.0
Files:
README.md this :)
data.csv containing all data see bellow description
openmoji folder containing the image files
The data.csv contains:
idx the character as int
character text representation
bytes representation
hex representation (replace Ox with U+ for unicode)
description of the emoji
path_black path to the bw image… See the full description on the dataset page: https://huggingface.co/datasets/Kray-C/dataset-openmoji.Amazon-ML-Challenge-2024filtered_k12_resample_no_chineseCrop_Disease_Images
Crop Disease Expert Annotations
1,092 crop images annotated by agricultural experts with ground truth diagnoses covering pests, diseases, and nutrient deficiencies across 74 crop types.
Dataset
File: annotations.csv (4 columns)
Column
Description
image_url
URL to the crop image
crop
Crop or plant identified by the expert
diagnosis
Pest, disease, or nutrient deficiency name(s); Healthy if none
details
Expert's observations on visible symptoms… See the full description on the dataset page: https://huggingface.co/datasets/kirutheen/Crop_Disease_Images.ko-naver_review_medi
