datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aozorabunko-clean
Overview
This dataset provides a convenient and user-friendly format of data from Aozora Bunko (青空文庫), a website that compiles public-domain books in Japan, ideal for Machine Learning applications.
[For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f
Methodology
The code to reproduce this dataset is made available on GitHub: globis-org/aozorabunko-exctractor.
1. Data collection
We firstly downloaded the CSV file that… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-clean.Rustins_Super_Mega_Awesome_VEDU_Model
Rustin's Super Mega Awesome VEDU Model
A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive
winter-annual grass) across Montana from satellite + environmental data.
Science reference: docs/VEDU_48_predictors_detailed.md
Data decisions & gotchas: docs/CONTRADICTIONS.md
Parity with the Earth Engine build: docs/GEE_PARITY.md
Continue-the-build guide: docs/HANDOFF.md
Label inventory: docs/DATA_SOURCES.md
What it produces
57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.ENADE_Brazilian_national_university_examination_MCQ_483cukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.PerMed-MM
PerMed-MM: A Multimodal, Multi-Specialty Persian Medical Benchmark
🤗 Dataset | 📖 Paper | 📄 PDF
Dataset Description
PerMed-MM is a multimodal, multi-specialty benchmark designed to evaluate Vision Language Models (VLMs) on Persian medical question answering.
The dataset consists of 733 multiple-choice questions sourced from the Iranian National Medical Board Exams (years 2021 and 2023). Each question is paired with 1 to 5 clinically relevant images, totaling… See the full description on the dataset page: https://huggingface.co/datasets/universitytehran/PerMed-MM.university_disciplines_45kUniversity discipline dataset, including Discrete Mathematics, Introduction to Artificial Intelligence, Principles and Applications of Databases, and Computer Networks, etc.
University-News-Instruction-Zh一些高校校园新闻,约 65k * 3(类任务) 条,稍微做了一点点脱敏,尽可能地遮盖了作者名等。数据已经整理成了指令的形式,格式如下:
{
"id": <id>,
"category": "(title_summarize|news_classify|news_generate)",
"instruction": <对应的具体指令>,
"input": <空>,
"output": <指令对应的输出>
}
总共三类任务:标题总结、栏目分类、新闻生成,本质上是利用新闻元数据中的标题、栏目、内容排列组合生成的,所以可以保证数据完全准确。每个字段内容已经整理成了单行的格式。下面是三类任务的样例:
// 标题总结
{
"id": 22106,
"category": "title_summarize",
"instruction": "请你给下面的新闻取一则标题:\n点击图片观看视频… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/University-News-Instruction-Zh.aozorabunko-chats
Overview
This dataset is of conversations extracted from Aozora Bunko (青空文庫), which collects public-domain books in Japan, using a simple heuristic approach.
[For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f
Method
First, lines surrounded by quotation mark pairs (「」) are extracted as utterances from the text field of globis-university/aozorabunko-clean.
Then, consecutive utterances are collected and grouped together.
The code… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-chats.aalen_university_faculty_computer_science
Dataset Card
This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created.
It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.ap-exam-credit-by-university
AP exam score to college credit, compared across universities
Canonical, always-current version: https://referencesource.org/ap-exam-credit-by-university/
Machine-readable: https://referencesource.org/ap-exam-credit-by-university/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2027-08-05 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 221
What each university actually grants for a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ap-exam-credit-by-university.clep-credit-by-university
CLEP exam credit policies by university
Canonical, always-current version: https://referencesource.org/clep-credit-by-university/
Machine-readable: https://referencesource.org/clep-credit-by-university/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-18
Stale after: 2027-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 130
Which CLEP exams each university accepts for credit, the minimum… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/clep-credit-by-university.indian_university_guidance_for_bangladeshi_students
Indian University Guidance for Bangladeshi Students Dataset
Dataset Description
This dataset contains 7,044 high-quality, instruction-formatted Question-Answer pairs designed for fine-tuning Large Language Models (LLMs). The primary goal of this dataset is to create a specialized AI counselor that provides accurate, culturally relevant, and comprehensive guidance on Indian universities for Bangladeshi students.
The dataset was generated through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/millat/indian_university_guidance_for_bangladeshi_students.arabic-legal-text
UAE Arabic Legal Text Corpus
Dataset Description
This repository contains a curated corpus of United Arab Emirates (UAE) legal texts, structured specifically for Natural Language Processing (NLP) tasks.
The dataset is maintained by University of Dubai Research to support research in:
Arabic legal intelligence
Automated summarization
Retrieval-Augmented Generation (RAG)
Legal information systems
Key Information
Curated by: Mohamed Asath… See the full description on the dataset page: https://huggingface.co/datasets/University-of-Dubai/arabic-legal-text.All_university
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/xzitao/All_university.university-ai-policies
GradPilot University AI Policies
GradPilot University AI Policies is a structured dataset of current university admissions policies on generative AI use in application essays and related written materials.
Each row represents one institution-level policy record and includes:
GradPilot's current L/D/E classification
supporting trigger quotes
source URLs from official university pages
scope and audience notes
any program-specific overrides captured in the current record
The… See the full description on the dataset page: https://huggingface.co/datasets/gradpilot/university-ai-policies.dc_elements_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology elements.
dc_terms_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology terms.
dc_terms_promptsfordham-university
ibleducation/fordham-university
This dataset contains a set of query and response pairs about Fordham university
Data for the dataset was scrapped from fordham.edu using GptCrawler.
The resulting pages were then converted to query response pairs using GPT-3.5
A total of 2707 data points exist in this dataset.
UniversityKarachi-University-Prospectus
About Dataset:
The dataset is created using an amazing library called Augmentoolkit. You can use this dataset to fine-tune the llms.
License: MIT
Task_categories: text2text-generation
Language: English
Tags: Education , KarachiUniversity , Prospectus
Size_categories: n<1K
smart-university-kz-ragdc_elements_promptsuniversity_que_ans_fordham-universityeon-university-knowledge-baseuk-university-advisorBisha_University_QA
