datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.noun-courseware-sft
NOUN Courseware SFT Dataset
Dataset Description
This dataset contains high-quality Supervised Fine-Tuning (SFT) question-and-answer pairs generated from the National Open University of Nigeria (NOUN) course materials. It is designed to train language models on diverse academic concepts ranging from agriculture to law.
Dataset Structure
The dataset is divided into specific Subsets (Configs) based on the academic Faculty. You can select different… See the full description on the dataset page: https://huggingface.co/datasets/Mavies5526/noun-courseware-sft.openedu-courses
1.2k OpenEdu.ru courses (2025)
Overview
The dataset contains structured information about courses from the OpenEdu.ru platform: course cards with basic metadata, extended information extracted from each individual course page, and a mapping between course directions (groups). It is useful for analysing educational programmes, searching via metadata, and training text-classification models, including RuBERT and other transformers, because it includes direction codes that… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/openedu-courses.task1193_food_course_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1193_food_course_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1193_food_course_classification.smolified-course-selector
🤏 smolified-course-selector
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Rudraksh2004/smolified-course-selector.
📦 Asset Details
Origin: Smolify Foundry (Job ID: b279efb3)
Records: 1228
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Rudraksh2004.
Generated via Smolify.ai.
smolified-course-selector
🤏 smolified-course-selector
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-course-selector.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 65ce4eaf)
Records: 376
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
