CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01globis-university /aozorabunko-clean Overview This dataset provides a convenient and user-friendly format of data from Aozora Bunko (青空文庫), a website that compiles public-domain books in Japan, ideal for Machine Learning applications. [For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f Methodology The code to reproduce this dataset is made available on GitHub: globis-org/aozorabunko-exctractor. 1. Data collection We firstly downloaded the CSV file that… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-clean.texttext-generation10K<n<100K48 likes1.2k downloads3y agoHugging Face02UniversityOfMontanaSAL /Rustins_Super_Mega_Awesome_VEDU_Model Rustin's Super Mega Awesome VEDU Model A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive winter-annual grass) across Montana from satellite + environmental data. Science reference: docs/VEDU_48_predictors_detailed.md Data decisions & gotchas: docs/CONTRADICTIONS.md Parity with the Earth Engine build: docs/GEE_PARITY.md Continue-the-build guide: docs/HANDOFF.md Label inventory: docs/DATA_SOURCES.md What it produces 57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.imagen<1K0 likes446 downloads12d agoHugging Face03northern-64bit /ENADE_Brazilian_national_university_examination_MCQ_483textquestion-answeringn<1K0 likes297 downloads2y agoHugging Face04Naholav /cukurova_university_chatbot Çukurova University Computer Engineering Chatbot Dataset 📊 Dataset Overview This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information. 🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.textquestion-answering10K<n<100K0 likes138 downloads1y agoHugging Face05universitytehran /PerMed-MM PerMed-MM: A Multimodal, Multi-Specialty Persian Medical Benchmark 🤗 Dataset | 📖 Paper | 📄 PDF Dataset Description PerMed-MM is a multimodal, multi-specialty benchmark designed to evaluate Vision Language Models (VLMs) on Persian medical question answering. The dataset consists of 733 multiple-choice questions sourced from the Iranian National Medical Board Exams (years 2021 and 2023). Each question is paired with 1 to 5 clinically relevant images, totaling… See the full description on the dataset page: https://huggingface.co/datasets/universitytehran/PerMed-MM.imagevisual-question-answeringn<1K1 likes82 downloads3mo agoHugging Face06GGGgrey /university_disciplines_45kUniversity discipline dataset, including Discrete Mathematics, Introduction to Artificial Intelligence, Principles and Applications of Databases, and Computer Networks, etc. tabular10K<n<100K2 likes64 downloads9mo agoHugging Face07Mxode /University-News-Instruction-Zh一些高校校园新闻,约 65k * 3(类任务) 条,稍微做了一点点脱敏,尽可能地遮盖了作者名等。数据已经整理成了指令的形式,格式如下: { "id": <id>, "category": "(title_summarize|news_classify|news_generate)", "instruction": <对应的具体指令>, "input": <空>, "output": <指令对应的输出> } 总共三类任务:标题总结、栏目分类、新闻生成,本质上是利用新闻元数据中的标题、栏目、内容排列组合生成的,所以可以保证数据完全准确。每个字段内容已经整理成了单行的格式。下面是三类任务的样例: // 标题总结 { "id": 22106, "category": "title_summarize", "instruction": "请你给下面的新闻取一则标题:\n点击图片观看视频… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/University-News-Instruction-Zh.textzero-shot-classification100K<n<1M4 likes40 downloads1y agoHugging Face08globis-university /aozorabunko-chats Overview This dataset is of conversations extracted from Aozora Bunko (青空文庫), which collects public-domain books in Japan, using a simple heuristic approach. [For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f Method First, lines surrounded by quotation mark pairs (「」) are extracted as utterances from the text field of globis-university/aozorabunko-clean. Then, consecutive utterances are collected and grouped together. The code… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-chats.texttext-generation1K<n<10K12 likes34 downloads3y agoHugging Face09Puidii /aalen_university_faculty_computer_science Dataset Card This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created. It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.textquestion-answering1K<n<10K0 likes31 downloads2y agoHugging Face10referencesource /ap-exam-credit-by-university AP exam score to college credit, compared across universities Canonical, always-current version: https://referencesource.org/ap-exam-credit-by-university/ Machine-readable: https://referencesource.org/ap-exam-credit-by-university/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-05 Stale after: 2027-08-05 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 221 What each university actually grants for a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ap-exam-credit-by-university.textn<1K0 likes31 downloads1mo agoHugging Face11referencesource /clep-credit-by-university CLEP exam credit policies by university Canonical, always-current version: https://referencesource.org/clep-credit-by-university/ Machine-readable: https://referencesource.org/clep-credit-by-university/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-18 Stale after: 2027-08-18 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 130 Which CLEP exams each university accepts for credit, the minimum… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/clep-credit-by-university.textn<1K0 likes31 downloads1mo agoHugging Face12millat /indian_university_guidance_for_bangladeshi_students Indian University Guidance for Bangladeshi Students Dataset Dataset Description This dataset contains 7,044 high-quality, instruction-formatted Question-Answer pairs designed for fine-tuning Large Language Models (LLMs). The primary goal of this dataset is to create a specialized AI counselor that provides accurate, culturally relevant, and comprehensive guidance on Indian universities for Bangladeshi students. The dataset was generated through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/millat/indian_university_guidance_for_bangladeshi_students.text1K<n<10K1 likes24 downloads1y agoHugging Face13University-of-Dubai /arabic-legal-text UAE Arabic Legal Text Corpus Dataset Description This repository contains a curated corpus of United Arab Emirates (UAE) legal texts, structured specifically for Natural Language Processing (NLP) tasks. The dataset is maintained by University of Dubai Research to support research in: Arabic legal intelligence Automated summarization Retrieval-Augmented Generation (RAG) Legal information systems Key Information Curated by: Mohamed Asath… See the full description on the dataset page: https://huggingface.co/datasets/University-of-Dubai/arabic-legal-text.texttext-classificationn<1K1 likes23 downloads8mo agoHugging Face14xzitao /All_university Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/xzitao/All_university.textquestion-answering100K<n<1M0 likes20 downloads1y agoHugging Face15gradpilot /university-ai-policies GradPilot University AI Policies GradPilot University AI Policies is a structured dataset of current university admissions policies on generative AI use in application essays and related written materials. Each row represents one institution-level policy record and includes: GradPilot's current L/D/E classification supporting trigger quotes source URLs from official university pages scope and audience notes any program-specific overrides captured in the current record The… See the full description on the dataset page: https://huggingface.co/datasets/gradpilot/university-ai-policies.texttext-classificationn<1K1 likes16 downloads7mo agoHugging Face16BOP-Berlin-University-Alliance /dc_elements_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology elements. texttext-classificationn<1K0 likes15 downloads3y agoHugging Face17BOP-Berlin-University-Alliance /dc_terms_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology terms. texttext-classificationn<1K0 likes15 downloads3y agoHugging Face18BOP-Berlin-University-Alliance /dc_terms_promptstextn<1K1 likes12 downloads3y agoHugging Face19iblai /fordham-university ibleducation/fordham-university This dataset contains a set of query and response pairs about Fordham university Data for the dataset was scrapped from fordham.edu using GptCrawler. The resulting pages were then converted to query response pairs using GPT-3.5 A total of 2707 data points exist in this dataset. textquestion-answering1K<n<10K0 likes11 downloads3y agoHugging Face20OpenXPlus123 /Universitytabular10K<n<100K0 likes11 downloads2y agoHugging Face21BilalHaneef /Karachi-University-Prospectus About Dataset: The dataset is created using an amazing library called Augmentoolkit. You can use this dataset to fine-tune the llms. License: MIT Task_categories: text2text-generation Language: English Tags: Education , KarachiUniversity , Prospectus Size_categories: n<1K textn<1K0 likes11 downloads2y agoHugging Face22Nucelios /smart-university-kz-ragdocumentn<1K0 likes10 downloads3mo agoHugging Face23BOP-Berlin-University-Alliance /dc_elements_promptstextn<1K0 likes8 downloads3y agoHugging Face24Prathamesh25 /university_que_ans_textn<1K0 likes5 downloads3y agoHugging Face25yoongjerung /fordham-universitytext1K<n<10K0 likes2 downloads2y agoHugging Face26Eonuniversity /eon-university-knowledge-basetextn<1K0 likes2 downloads10mo agoHugging Face27kishanmadhesiya /uk-university-advisortext1K<n<10K0 likes2 downloads6mo agoHugging Face28SarahKH13 /Bisha_University_QAtextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.