datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hand_gesture_dataThe study is conducted on a total of 7 participants. The participants were instructed to perform three hand gestures (Hold, Single Tap and Double Tap) under different light conditions (low(100-200 lux, medium (600-750 lux) and high (1500-1600 lux)) and at different distances from the light sensor (low(2-4 cm) and high(8-10 cm))
RDB2G-Bench
RDB2G-Bench
This is an offical dataset of the paper RDB2G-Bench: A Comprehensive Benchmark for Automatic Graph Modeling of Relational Databases.
RDB2G-Bench is a toolkit for benchmarking graph-based analysis and prediction tasks by converting relational database data into graphs.
Our code is available at GitHub.
Overview
RDB2G-Bench provides comprehensive performance evaluation data for graph neural network models applied to relational database tasks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaistdata/RDB2G-Bench.kaiser_raman_ecoli_fermentation
Dataset Card for Kaiser Raman Spectra from E. coli Cultivation
Dataset Summary
This dataset contains offline Raman spectra collected from the samples containing cells during the cultivation of Escherichia coli W3110. Measurements were performed using Kaiser RXN1 from samples of a bioreactor cultivation involving a 10-hour batch phase followed by a pulse-based fed-batch phase to induce metabolic transients.
Dataset Structure
Data Points
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/chlange/kaiser_raman_ecoli_fermentation.streptococcus_thermophilus_fermentation_kaiser
Dataset Card for Raman Spectra from Fermentations of Streptococcus thermophilus Kaiser
Dataset Summary
This dataset contains offline Raman spectra collected during batch cultivations of Streptococcus thermophilus. The data is intended for researchers developing chemometric models and machine learning pipelines for bioprocess monitoring.
The dataset includes two distinct fermentation runs conducted in shake flasks over a 24-hour period, paired with reference analytical… See the full description on the dataset page: https://huggingface.co/datasets/chlange/streptococcus_thermophilus_fermentation_kaiser.kaiser_raman_ecoli_fermentation_supernatant
Dataset Card for Kaiser Raman Spectra from E. coli Cultivation Supernatants
Dataset Summary
This dataset contains offline Raman spectra collected from cell-free supernatant samples during the cultivation of Escherichia coli W3110. Measurements were performed using Kaiser RXN1 from the complex cultivation media. The dataset corresponds to a bioreactor cultivation involving a 10-hour batch phase followed by a pulse-based fed-batch phase to induce metabolic transients.… See the full description on the dataset page: https://huggingface.co/datasets/chlange/kaiser_raman_ecoli_fermentation_supernatant.LearningChat_reflective_writing_vaults
AI활용성찰적글쓰기(2025-2) 학생별 옵시디언 볼트 공개용 데이터셋
한 줄 요약
2025-2학기 한림대학교 AI활용성찰적글쓰기 수업의 기말과제 제출물인 학생별 개인 Obsidian 볼트 묶음을 공개용 기준으로 문서화한 데이터셋입니다.
데이터셋 개요
샘플 단위: 학생별 옵시디언 볼트 묶음 1개
총 샘플 수: 46
메타데이터 파일: metadata.csv
공개용 식별 방식: student_001부터 student_046까지의 익명 샘플 ID
데이터 성격: 학생별 개인 지식관리 볼트 제출물 요약 메타데이터
이 데이터셋은 개별 노트를 독립 샘플로 다루지 않습니다. 각 샘플은 하나의 학생 제출 묶음이며, 개별 Markdown 노트, 이미지, PDF, Canvas 파일은 해당 샘플의 하위 구성요소로 취급합니다.
생성 배경
본 데이터셋은 한림대학교 2025-2학기 AI활용성찰적글쓰기 수업의 기말과제 제출물을… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_reflective_writing_vaults.Hate-Speech-TweetsMAG188
MAG188
MAG188 is a DFT benchmark of 188 experimentally characterized collinear magnetic materials from the
MAGNDATA database. All structures were recomputed with spin-polarized DFT and relaxed. MAG188 contains
two benchmarks:
MAG188-EXP: the relaxed experimentally reported magnetic state of each material (188 endpoints).
It measures accuracy on experimentally established magnetic states.
MAG188-SAMPLE: 1,195 relaxed endpoints that cover several self-consistent collinear spin… See the full description on the dataset page: https://huggingface.co/datasets/kairosmaterial/MAG188.hanzi-sino-vietnamese
HSK × Sino-Vietnamese (Hán-Việt) character dataset
768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese.
Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.LearningChat_ai_video_production
Hallym AI Video Production Practice 2025-2 Public Dataset
1. 데이터셋 개요
데이터셋명: 한림대학교 AI영상제작실습 2025-2 공개용 데이터셋
교과목명: AI영상제작실습
학기: 2025-2
생성 배경: 2025학년도 2학기 AI영상제작실습 수업에서 조별로 제작·제출한 AI 기반 영상 결과물을 공개용 데이터셋 형태로 정리한 것이다.
목적: 수업 기반 AI 영상 창작 결과물을 공개 아카이브 형태로 정리하고, 작품 단위 메타데이터를 함께 제공하기 위함이다.
2. 데이터셋 범위
총 작품 수: 20편
데이터 단위: 조별 제출 영상 1편 = metadata.csv 1행
포함 대상: 1조부터 20조까지 각 팀 폴더의 원본 MP4 1개
제외 대상:
보고서 파일(.pdf, .docx, .hwp)
라이선스 동의서 파일
AI제작콘텐츠 발표회 2025 출품작 폴더에 따로 복사된 중복 MP4… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_ai_video_production.new-mom-5ef1df
new-mom-5ef1df
Synthetic sensors test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Silver-Kai/new-mom-5ef1df.international-celebration-7366b6
international-celebration-7366b6
Synthetic products test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number… See the full description on the dataset page: https://huggingface.co/datasets/Kestrel-Kai/international-celebration-7366b6.noamundi-rockfall-simulation-dataset
🪨 Noamundi Rockfall Simulation Dataset
📌 1. Dataset Overview
This dataset provides a physics-informed simulation of rockfall precursor conditions and trajectory runouts for 500 monitoring locations in the Noamundi mining region in Jharkhand, India. It is designed specifically for training and benchmarking Machine Learning models on early-warning systems, anomaly detection, rare-event classification, and spatio-temporal (ST-GNN) trajectory modeling in… See the full description on the dataset page: https://huggingface.co/datasets/Kaizen696/noamundi-rockfall-simulation-dataset.Fertilizer-Predictionnetwork-QnA-datasetLearningChat_accounting_ai_questions
공개용 회계입문 AI질문 이미지 데이터셋
이 데이터셋은 회계입문 수업의 AI질문 1회 과제 제출 이미지들을 공개용으로 문서화하기 위해 정리한 메타데이터 패키지다. 현재 폴더에 존재하는 6개 수집 배치 전체를 통합했으며, 공개 버전에서는 학생 실명과 원본 파일명을 직접 노출하지 않도록 비식별 규칙을 적용했다.
본 문서와 함께 제공되는 metadata.csv는 이미지 파일 1개당 1행을 가지는 인벤토리다. 실제 공개 배포 시 이미지 파일은 data/images/AIQ-XXXXXX.ext 형식으로 익명 재배치하는 것을 전제로 한다.
1. 데이터셋 범위
대상 과목: 회계입문
대상 과제: AI질문 1회
포함 범위: 현재 작업 폴더에 있는 6개 수집 배치 전체
레코드 단위: 이미지 파일 1개 = metadata.csv 1행
공개 버전 기준: 완전 비식별 전제
분반별 구성
section_id
익명 제출자 수
이미지 수… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_accounting_ai_questions.hallym_AI_OpenDataset
Hallym Adult and Child Speech Dataset
This dataset contains speech recordings and transcriptions collected from adult and child speakers for AI-based speech and language research.
Dataset Overview
Total Records: 2,714
Speakers: 49 (adult: 25, child: 24)
Groups: adult, child
File Format: WAV (audio) + TXT (transcription)
Speaker Statistics
Group
Count
Gender
Age Range
Adult
25명
남/여
50~78세
Child
24명
남/여
3~8세
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_AI_OpenDataset.charades-sta-test
Charades-STA test set
This is the test set of the Charades-STA dataset.
Website: https://prior.allenai.org/projects/charades
Github (code for evaluation, training, etc.): https://github.com/jiyanggao/TALL
Original paper: https://arxiv.org/abs/1705.02101
Files:
charades_sta_test.txt : Original test set answers
videos.zip : test set videos (1334 in total) zipped
Small test set
There is a small test set that can be used for small scale testing (150 queries)
Files:… See the full description on the dataset page: https://huggingface.co/datasets/Kainuo/charades-sta-test.triage-ai-dataset
Triage AI Dataset
Du lieu phan loai muc do uu tien kham benh (KTAS-based).
Files
File
Records
Description
ktas_cleaned.csv
561
Du lieu thuc sau lam sach tu KTAS goc
augmented_ktas.csv
8,361
561 real + 7,800 synthetic (Rule-based Clinical Generator)
Features (16)
age, gender (Nam/Nu), heart_rate, respiratory_rate, temperature, spo2,
systolic_bp, diastolic_bp, pulse_pressure, shock_index, map, tachycardia,
bradycardia, hypotension… See the full description on the dataset page: https://huggingface.co/datasets/Kaiyow0/triage-ai-dataset.test5190-news-source
