datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AiEdit
🎧 AiEdit Dataset
📖 Introduction
AiEdit is a large-scale, cross-lingual speech editing dataset designed to advance research and evaluation in Speech Editing tasks. We have constructed an automated data generation pipeline comprising the following core modules:
Text Engine: Powered by Large Language Models (LLMs), this engine intelligently processes raw text to execute three types of editing operations: Addition, Deletion, and Modification.
Speech Synthesis & Editing:… See the full description on the dataset page: https://huggingface.co/datasets/nccm2p2/AiEdit.AiEdit
🎧 AiEdit Dataset
📖 Introduction
AiEdit is a large-scale, cross-lingual speech editing dataset designed to advance research and evaluation in Speech Editing tasks. We have constructed an automated data generation pipeline comprising the following core modules:
Text Engine: Powered by Large Language Models (LLMs), this engine intelligently processes raw text to execute three types of editing operations: Addition, Deletion, and Modification.
Speech Synthesis & Editing:… See the full description on the dataset page: https://huggingface.co/datasets/JunXueTech/AiEdit.LearningChat_reflective_writing_vaults
AI활용성찰적글쓰기(2025-2) 학생별 옵시디언 볼트 공개용 데이터셋
한 줄 요약
2025-2학기 한림대학교 AI활용성찰적글쓰기 수업의 기말과제 제출물인 학생별 개인 Obsidian 볼트 묶음을 공개용 기준으로 문서화한 데이터셋입니다.
데이터셋 개요
샘플 단위: 학생별 옵시디언 볼트 묶음 1개
총 샘플 수: 46
메타데이터 파일: metadata.csv
공개용 식별 방식: student_001부터 student_046까지의 익명 샘플 ID
데이터 성격: 학생별 개인 지식관리 볼트 제출물 요약 메타데이터
이 데이터셋은 개별 노트를 독립 샘플로 다루지 않습니다. 각 샘플은 하나의 학생 제출 묶음이며, 개별 Markdown 노트, 이미지, PDF, Canvas 파일은 해당 샘플의 하위 구성요소로 취급합니다.
생성 배경
본 데이터셋은 한림대학교 2025-2학기 AI활용성찰적글쓰기 수업의 기말과제 제출물을… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_reflective_writing_vaults.cyber_drug_dataset
Digital Forensic Investigation Scenario Dataset: Online Drug Trafficking
This dataset is a comprehensive collection of digital artifacts and investigative reports designed for forensic research and education. It simulates a sophisticated Online Drug Trafficking scenario, covering the entire investigation lifecycle from initial intelligence gathering to suspect arrest and financial analysis.
Dataset Structure
The dataset is indexed via a standardized 7-column metadata… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/cyber_drug_dataset.Pathological-child-voice
Speech Dataset for AI-Based Language Assessment in Children
The "Speech Database of Typically Developing and Speech-Impaired Children" is an open speech dataset designed to support the development of AI-based language assessment systems. It contains speech samples from children aged 2 to 9 who are either typically developing or have reduced consonant articulation accuracy.
This dataset is based on standardized Korean articulation tools:
APAC (Articulation and Phonology… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/Pathological-child-voice.iPCL-R
This is dataset for project iPCL-R.
A Pre-training foundation model for Chip Layout Routing
Demo of iPCL-R
A chip layout generation demo by iPCL-R
Project Overview
The framework of iPCL-R. (a) Dataset. (b) Symbolic system. (c) Pre-training. (d) Inference.
iPCL-R addresses the challenge of automated routing pattern generation in chip design by treating routing patterns as sequences that can be learned and generated by large language… See the full description on the dataset page: https://huggingface.co/datasets/AiEDA/iPCL-R.LearningChat_ai_video_production
Hallym AI Video Production Practice 2025-2 Public Dataset
1. 데이터셋 개요
데이터셋명: 한림대학교 AI영상제작실습 2025-2 공개용 데이터셋
교과목명: AI영상제작실습
학기: 2025-2
생성 배경: 2025학년도 2학기 AI영상제작실습 수업에서 조별로 제작·제출한 AI 기반 영상 결과물을 공개용 데이터셋 형태로 정리한 것이다.
목적: 수업 기반 AI 영상 창작 결과물을 공개 아카이브 형태로 정리하고, 작품 단위 메타데이터를 함께 제공하기 위함이다.
2. 데이터셋 범위
총 작품 수: 20편
데이터 단위: 조별 제출 영상 1편 = metadata.csv 1행
포함 대상: 1조부터 20조까지 각 팀 폴더의 원본 MP4 1개
제외 대상:
보고서 파일(.pdf, .docx, .hwp)
라이선스 동의서 파일
AI제작콘텐츠 발표회 2025 출품작 폴더에 따로 복사된 중복 MP4… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_ai_video_production.finger_print_dataset
Latent Fingerprint Detection Dataset
The "Latent Fingerprint Detection Dataset" is an open image dataset designed to support the development of AI-based fingerprint detection and analysis systems. It contains fingerprint images collected using various detection methods on different surface types.
This dataset includes fingerprint samples detected using chemical and powder-based methods:
Ninhydrin - Chemical detection method
1,2-Indandione - Chemical detection method
Black Magnetic… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/finger_print_dataset.3D_capacitance
3D Capacitance Dataset
This project serves as the submodule of PC-Cap.
The dataset generation is briefly shown in Fig. 1.
Dataset Description
Directory Descriptions
The complete directory structure is as follows:
dataset
├── pc
├── pc_annotation_coupling_test_clean.csv
├── pc_annotation_coupling_test.csv
├── pc_annotation_coupling_train_clean.csv
├── pc_annotation_coupling_train.csv
├── pc_annotation.csv
├── pc_annotation_total_test_clean.csv
├──… See the full description on the dataset page: https://huggingface.co/datasets/AiEDA/3D_capacitance.hallym_coding_easy
NOTI Coding Dataset - Easy Difficulty
This is a subset of the NOTI Coding Education Dataset filtered by difficulty = Easy.
Dataset Overview
Total Records: 834
Filter Criteria: difficulty = Easy
Data Structure
Each row represents a single grading record:
problem_title: Problem identifier
student_id: Student identifier (anonymized)
code: Submitted code
grader_id: Grader identifier
score: Human grading score (0-10)
grading_details: Detailed grading feedback… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_coding_easy.Ai_education_datasetlevel-aware-qa-dataset
Construction of LLM-based Level-Aware QA Dataset for AI Tutor Development
This dataset is a resource-based QA dataset of approximately 5,000 items generated using GPT-4o based on deep learning major lecture materials and foundational papers. It is a multi-modal dataset containing not only text but also visual information such as formulas and charts. The questions and answers are systematically organized according to the learner's comprehension level (High/Mid/Low). In particular… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/level-aware-qa-dataset.ai-editor-training-dataonline_fraud_dataset
Digital Forensic Investigation Scenario Dataset (Online Fraud)
This dataset is a structured collection of digital evidence for research in digital forensics and criminal investigation. It simulates a multi-stage Online Fraud (Voice Phishing) scenario. All metadata and evidence linkages are defined in the accompanying metadata.csv file.
Dataset Structure
The dataset features follow the exact structure of the provided metadata file, ensuring data integrity and consistency… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/online_fraud_dataset.LearningChat_accounting_ai_questions
공개용 회계입문 AI질문 이미지 데이터셋
이 데이터셋은 회계입문 수업의 AI질문 1회 과제 제출 이미지들을 공개용으로 문서화하기 위해 정리한 메타데이터 패키지다. 현재 폴더에 존재하는 6개 수집 배치 전체를 통합했으며, 공개 버전에서는 학생 실명과 원본 파일명을 직접 노출하지 않도록 비식별 규칙을 적용했다.
본 문서와 함께 제공되는 metadata.csv는 이미지 파일 1개당 1행을 가지는 인벤토리다. 실제 공개 배포 시 이미지 파일은 data/images/AIQ-XXXXXX.ext 형식으로 익명 재배치하는 것을 전제로 한다.
1. 데이터셋 범위
대상 과목: 회계입문
대상 과제: AI질문 1회
포함 범위: 현재 작업 폴더에 있는 6개 수집 배치 전체
레코드 단위: 이미지 파일 1개 = metadata.csv 1행
공개 버전 기준: 완전 비식별 전제
분반별 구성
section_id
익명 제출자 수
이미지 수… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_accounting_ai_questions.hallym_AI_OpenDataset
Hallym Adult and Child Speech Dataset
This dataset contains speech recordings and transcriptions collected from adult and child speakers for AI-based speech and language research.
Dataset Overview
Total Records: 2,714
Speakers: 49 (adult: 25, child: 24)
Groups: adult, child
File Format: WAV (audio) + TXT (transcription)
Speaker Statistics
Group
Count
Gender
Age Range
Adult
25명
남/여
50~78세
Child
24명
남/여
3~8세
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_AI_OpenDataset.korean_monosyllabic_speech
Korean Monosyllabic Speech Perception Test Dataset
「Korean Monosyllabic Speech Perception Test Database」set is an open speech dataset created for evaluating monosyllables (meaningless, meaningful) and researching error patterns in elderly individuals with mild to moderate hearing loss. This dataset selected only monosyllables with a correct response rate of 80~100% out of 3192 possible Korean consonant-vowel combination sounds.
The dataset is distributed under the CC BY NC ND 4.0… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/korean_monosyllabic_speech.hallym_coding_javascript
NOTI Coding Dataset - JavaScript Language Submissions
This is a subset of the NOTI Coding Education Dataset filtered by programming_language = JavaScript.
Dataset Overview
Total Records: 691
Filter Criteria: programming_language = JavaScript
Data Structure
Each row represents a single grading record:
problem_title: Problem identifier
student_id: Student identifier (anonymized)
code: Submitted code
grader_id: Grader identifier
score: Human grading score… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_coding_javascript.pig_butchering_dataset
Financial Transaction Dataset for Online Fraud Investigation
This dataset is a structured collection of digital financial evidence designed for research in digital forensics and anti-money laundering (AML). It simulates various bank account activities, including internet banking logs and personal profiles linked to an Online Fraud scenario.
Dataset Structure
The dataset features follow the exact 7-column structure of the provided metadata, ensuring data integrity and… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/pig_butchering_dataset.hallym_coding_c
NOTI Coding Dataset - C Language Submissions
This is a subset of the NOTI Coding Education Dataset filtered by programming_language = C.
Dataset Overview
Total Records: 1963
Filter Criteria: programming_language = C
Data Structure
Each row represents a single grading record:
problem_title: Problem identifier
student_id: Student identifier (anonymized)
code: Submitted code
grader_id: Grader identifier
score: Human grading score (0-10)
grading_details:… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_coding_c.hallym_coding_python
NOTI Coding Dataset - Python Language Submissions
This is a subset of the NOTI Coding Education Dataset filtered by programming_language = Python.
Dataset Overview
Total Records: 2964
Filter Criteria: programming_language = Python
Data Structure
Each row represents a single grading record:
problem_title: Problem identifier
student_id: Student identifier (anonymized)
code: Submitted code
grader_id: Grader identifier
score: Human grading score (0-10)… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_coding_python.hallym_coding_medium
NOTI Coding Dataset - Medium Difficulty
This is a subset of the NOTI Coding Education Dataset filtered by difficulty = Medium.
Dataset Overview
Total Records: 3650
Filter Criteria: difficulty = Medium
Data Structure
Each row represents a single grading record:
problem_title: Problem identifier
student_id: Student identifier (anonymized)
code: Submitted code
grader_id: Grader identifier
score: Human grading score (0-10)
grading_details: Detailed grading… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_coding_medium.hallym_coding_hard
NOTI Coding Dataset - Hard Difficulty
This is a subset of the NOTI Coding Education Dataset filtered by difficulty = Hard.
Dataset Overview
Total Records: 954
Filter Criteria: difficulty = Hard
Data Structure
Each row represents a single grading record:
problem_title: Problem identifier
student_id: Student identifier (anonymized)
code: Submitted code
grader_id: Grader identifier
score: Human grading score (0-10)
grading_details: Detailed grading feedback… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_coding_hard.aiedu0407aiedu0406
