datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
university-1652
University-1652: Drone-based Geo-localization Benchmark 🚁
University-1652 is a multi-view dataset for drone-based geo-localization, annotating 1652 buildings across 72 universities (ACM Multimedia 2020, paper). Cited in 500+ papers, it supports Drone → Satellite localization and Satellite → Drone navigation.
🔗 Official code & baseline: layumi/University1652-Baseline
· Leaderboard: State-of-the-art results
🔐 Access
This dataset is gated: click "Request… See the full description on the dataset page: https://huggingface.co/datasets/layumi/university-1652.aozorabunko-clean
Overview
This dataset provides a convenient and user-friendly format of data from Aozora Bunko (青空文庫), a website that compiles public-domain books in Japan, ideal for Machine Learning applications.
[For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f
Methodology
The code to reproduce this dataset is made available on GitHub: globis-org/aozorabunko-exctractor.
1. Data collection
We firstly downloaded the CSV file that… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-clean.University-HMER-RealClassroom
University HMER RealClassroom
University HMER RealClassroom is a curated dataset of 1,636 handwritten
university-calculus expressions captured under practical classroom-style conditions.
Each image is paired with a manually reviewed tokenized LaTeX transcription and
structured metadata describing expression family and sequence-length difficulty.
This initial Hub release is private while ownership, privacy, annotation quality, and
licensing are reviewed. The current license:… See the full description on the dataset page: https://huggingface.co/datasets/tuan3110/University-HMER-RealClassroom.Rustins_Super_Mega_Awesome_VEDU_Model
Rustin's Super Mega Awesome VEDU Model
A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive
winter-annual grass) across Montana from satellite + environmental data.
Science reference: docs/VEDU_48_predictors_detailed.md
Data decisions & gotchas: docs/CONTRADICTIONS.md
Parity with the Earth Engine build: docs/GEE_PARITY.md
Continue-the-build guide: docs/HANDOFF.md
Label inventory: docs/DATA_SOURCES.md
What it produces
57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.ENADE_Brazilian_national_university_examination_MCQ_483alitaqi000_world-university-rankings-2023
World University Rankings 2023
World University Rankings 2023 include 1,799 universities across 104 countries.
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 8,164
Files: 1
Files
World University Rankings 2023.csv
Mirrored from Kaggle
marburg-university-buildings
Marburg University Buildings
The Marburg University Buildings dataset contains high-resolution photographs of 29 university buildings in Marburg, Germany. Each image is labeled by the name of the building and intended for use in multi-class image classification tasks.
Dataset Summary
Total classes: 29
Image format: JPG
Label format: Folder name as string
Collection: Manually photographed under consistent daylight and angle conditions (SAMSUNG galaxy A54)
No compression… See the full description on the dataset page: https://huggingface.co/datasets/soroushdft/marburg-university-buildings.cukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.turkish-university-entrance-exam-questionsus-university-college-school-district-education-layoffs-warn-act-notices-daily
US university, college, school and education layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-23. 1,165 layoff and closure notices filed by
universities and colleges, for-profit career schools and their campuses, private and charter schools, school districts where a state chose to publish them, Head Start and childcare providers, school-bus and campus contractors, and education publishers and online-learning companies with US state labor departments… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-university-college-school-district-education-layoffs-warn-act-notices-daily.vn-provinces-university-lecturers
Vietnam provinces university lecturers
Vietnam provinces university lecturers. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (409 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (54 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-university-lecturers.vn-provinces-university-students
Vietnam provinces university students
Vietnam provinces university students. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (410 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (54 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-university-students.maya-dataset-v1PerMed-MM
PerMed-MM: A Multimodal, Multi-Specialty Persian Medical Benchmark
🤗 Dataset | 📖 Paper | 📄 PDF
Dataset Description
PerMed-MM is a multimodal, multi-specialty benchmark designed to evaluate Vision Language Models (VLMs) on Persian medical question answering.
The dataset consists of 733 multiple-choice questions sourced from the Iranian National Medical Board Exams (years 2021 and 2023). Each question is paired with 1 to 5 clinically relevant images, totaling… See the full description on the dataset page: https://huggingface.co/datasets/universitytehran/PerMed-MM.Bank-ReviewNigerian Banks - Bank Reviews Dataset Collection
A comprehensive collection of customer reviews from Google Play Store (from app launch to 2024)
Access Bank - CSV File: access_reviews.csv
EcoBank - CSV File: ecoBank_reviews.csv
First Bank of Nigeria (FBN) - CSV File: fbn_reviews.csv
FCMB - CSV File: fcmb_reviews.csv
Fidelity Bank -… See the full description on the dataset page: https://huggingface.co/datasets/Federal-University-Lokoja/Bank-Review.LearningChat_reflective_writing_vaults
AI활용성찰적글쓰기(2025-2) 학생별 옵시디언 볼트 공개용 데이터셋
한 줄 요약
2025-2학기 한림대학교 AI활용성찰적글쓰기 수업의 기말과제 제출물인 학생별 개인 Obsidian 볼트 묶음을 공개용 기준으로 문서화한 데이터셋입니다.
데이터셋 개요
샘플 단위: 학생별 옵시디언 볼트 묶음 1개
총 샘플 수: 46
메타데이터 파일: metadata.csv
공개용 식별 방식: student_001부터 student_046까지의 익명 샘플 ID
데이터 성격: 학생별 개인 지식관리 볼트 제출물 요약 메타데이터
이 데이터셋은 개별 노트를 독립 샘플로 다루지 않습니다. 각 샘플은 하나의 학생 제출 묶음이며, 개별 Markdown 노트, 이미지, PDF, Canvas 파일은 해당 샘플의 하위 구성요소로 취급합니다.
생성 배경
본 데이터셋은 한림대학교 2025-2학기 AI활용성찰적글쓰기 수업의 기말과제 제출물을… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_reflective_writing_vaults.university_disciplines_45kUniversity discipline dataset, including Discrete Mathematics, Introduction to Artificial Intelligence, Principles and Applications of Databases, and Computer Networks, etc.
university-1652
University-1652: Drone-based Geo-localization Benchmark 🚁
University-1652 is a multi-view dataset for drone-based geo-localization, annotating 1652 buildings across 72 universities (ACM Multimedia 2020, paper). Cited in 50+ papers, it supports Drone → Satellite localization and Satellite → Drone navigation.
Dataset Structure
Splits:
Train: 50,218 images (drone, satellite, street, google; 33 universities)
Test:
query_drone: 37,855 images
gallery_drone: 51,355… See the full description on the dataset page: https://huggingface.co/datasets/shanmuga12nivetha/university-1652.cardiff-university-tm-en-cy
Dataset Card for Cardiff University Translation Memories
Dataset Summary
This dataset consists of English-Welsh sentence pairs extracted from Cardiff University translation memories.
Supported Tasks and Leaderboards
language-modeling
parsing
semantic-segmentation
semantic-similarity-classification
semantic-similarity-scoring
sentiment-analysis
sentiment-classification
sentiment-scoring
Languages
English
Welsh
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/cardiff-university-tm-en-cy.LearningChat_ai_video_production
Hallym AI Video Production Practice 2025-2 Public Dataset
1. 데이터셋 개요
데이터셋명: 한림대학교 AI영상제작실습 2025-2 공개용 데이터셋
교과목명: AI영상제작실습
학기: 2025-2
생성 배경: 2025학년도 2학기 AI영상제작실습 수업에서 조별로 제작·제출한 AI 기반 영상 결과물을 공개용 데이터셋 형태로 정리한 것이다.
목적: 수업 기반 AI 영상 창작 결과물을 공개 아카이브 형태로 정리하고, 작품 단위 메타데이터를 함께 제공하기 위함이다.
2. 데이터셋 범위
총 작품 수: 20편
데이터 단위: 조별 제출 영상 1편 = metadata.csv 1행
포함 대상: 1조부터 20조까지 각 팀 폴더의 원본 MP4 1개
제외 대상:
보고서 파일(.pdf, .docx, .hwp)
라이선스 동의서 파일
AI제작콘텐츠 발표회 2025 출품작 폴더에 따로 복사된 중복 MP4… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_ai_video_production.Pathological-child-voice
Speech Dataset for AI-Based Language Assessment in Children
The "Speech Database of Typically Developing and Speech-Impaired Children" is an open speech dataset designed to support the development of AI-based language assessment systems. It contains speech samples from children aged 2 to 9 who are either typically developing or have reduced consonant articulation accuracy.
This dataset is based on standardized Korean articulation tools:
APAC (Articulation and Phonology… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/Pathological-child-voice.University-News-Instruction-Zh一些高校校园新闻,约 65k * 3(类任务) 条,稍微做了一点点脱敏,尽可能地遮盖了作者名等。数据已经整理成了指令的形式,格式如下:
{
"id": <id>,
"category": "(title_summarize|news_classify|news_generate)",
"instruction": <对应的具体指令>,
"input": <空>,
"output": <指令对应的输出>
}
总共三类任务:标题总结、栏目分类、新闻生成,本质上是利用新闻元数据中的标题、栏目、内容排列组合生成的,所以可以保证数据完全准确。每个字段内容已经整理成了单行的格式。下面是三类任务的样例:
// 标题总结
{
"id": 22106,
"category": "title_summarize",
"instruction": "请你给下面的新闻取一则标题:\n点击图片观看视频… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/University-News-Instruction-Zh.cyber_drug_dataset
Digital Forensic Investigation Scenario Dataset: Online Drug Trafficking
This dataset is a comprehensive collection of digital artifacts and investigative reports designed for forensic research and education. It simulates a sophisticated Online Drug Trafficking scenario, covering the entire investigation lifecycle from initial intelligence gathering to suspect arrest and financial analysis.
Dataset Structure
The dataset is indexed via a standardized 7-column metadata… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/cyber_drug_dataset.university-students-complaints
University Students Complaints Dataset
A labeled dataset of 332 real-world complaints collected from university students across multiple departments. Each complaint is annotated with category, severity, responsible departments, and fine-grained aspect labels — making it suitable for multi-label text classification, severity prediction, and department routing tasks.
Dataset Summary
University students regularly raise concerns about infrastructure, academics, technical… See the full description on the dataset page: https://huggingface.co/datasets/alaminxpro/university-students-complaints.Indiana_University_Chest_X-ray_Collectionap-exam-credit-by-university
AP exam score to college credit, compared across universities
Canonical, always-current version: https://referencesource.org/ap-exam-credit-by-university/
Machine-readable: https://referencesource.org/ap-exam-credit-by-university/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2027-08-05 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 221
What each university actually grants for a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ap-exam-credit-by-university.clep-credit-by-university
CLEP exam credit policies by university
Canonical, always-current version: https://referencesource.org/clep-credit-by-university/
Machine-readable: https://referencesource.org/clep-credit-by-university/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-18
Stale after: 2027-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 130
Which CLEP exams each university accepts for credit, the minimum… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/clep-credit-by-university.University_Mevzuat_QA_v2
University Mevzuat QA v2 Dataset
Dataset
This dataset comprises question-and-answer pairs derived from the regulations of universities in Turkey. The initial version, available at University_Mevzuat_QA has been enhanced.
For each regulation from every university, three question-and-answer pairs have been created. The questions are based on the issues that students may encounter in their daily academic lives, and the answers include references to the respective… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/University_Mevzuat_QA_v2.aozorabunko-chats
Overview
This dataset is of conversations extracted from Aozora Bunko (青空文庫), which collects public-domain books in Japan, using a simple heuristic approach.
[For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f
Method
First, lines surrounded by quotation mark pairs (「」) are extracted as utterances from the text field of globis-university/aozorabunko-clean.
Then, consecutive utterances are collected and grouped together.
The code… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-chats.turkish-university-mevzuat
Turkey University Regulation Data Collection
This dataset provides a comprehensive collection of regulatory documents of Turkish universities obtained from mevzuat.gov.tr.
It includes full texts of regulations with detailed publication information and unique identifiers.
Overview
Data Sources: mevzuat.gov.tr website
Technologies Used: Selenium, BeautifulSoup, Python
Data Formats: CSV
CSV Data Structure
Column
Description
Üniversite
Name of the… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/turkish-university-mevzuat.
