datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-course-final-assignment
Agent Course Final Assignment - Unified Dataset
Author: Arte(r)m Sedov
GitHub: https://github.com/arterm-sedov/
Project link: https://huggingface.co/spaces/arterm-sedov/agent-course-final-assignment
Dataset Description
This dataset is produced by the GAIA Unit 4 Agent for the Hugging Face Agents Course final assignment as part of an experimental multi-LLM agent system that demonstrates advanced AI agent capabilities. It demonstrates advanced AI agent capabilities for… See the full description on the dataset page: https://huggingface.co/datasets/arterm-sedov/agent-course-final-assignment.NLP_Assignment_1Assignment3spoken_norm_assignment
VietAI assignment: Vietnamese Inverse Text Normalization dataset
Dataset Description
Inverse text normalization (ITN) is the task that transforms spoken to written styles. It is particularly useful in automatic speech recognition (ASR) systems where proper names are often miss-recognized by their pronunciations instead of the written forms. By applying ITN, we can improve the readability of the ASR system’s output significantly. This dataset provides data for doing ITN… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/spoken_norm_assignment.deductive_logical_reasoning-room_assignmentPKU_NLPDL_Assignment1edan20-assignment2-selma-ngramsAAI_assignment2dom-formula-assignment-data
DOM Formula Assignment Dataset
Training and Testing Data for A Machine Learning and Benchmarking Approach for Molecular Formula Assignment of Ultra High-Resolution Mass Spectrometry Data from Complex Mixtures
Paper: Under review
Abstract
A machine learning approach to molecular formula assignment is crucial for unlocking the full potential of ultra-high resolution mass spectrometry (UHRMS) when analyzing complex mixtures. By combining data-driven models with… See the full description on the dataset page: https://huggingface.co/datasets/SaeedLab/dom-formula-assignment-data.slm-assignment-data
Grade-Level Vocabulary-Locked Writing Tutor — Dataset
Training and evaluation data for fine-tuning a small open model (Qwen3-0.6B)
into a grade 7–8 writing/grammar tutor whose vocabulary and sentence
complexity stay locked to the band — it introduces at most one word above
grade level per reply (always immediately defined) and never escalates, even
under pressure ("use bigger words", "give me the college version") or
jailbreak-style attacks.
The dataset is the deliverable. ~80%… See the full description on the dataset page: https://huggingface.co/datasets/blackbird0831/slm-assignment-data.selma-ngrams-assignment2
Selma Lagerlof n-gram counts
This dataset was produced for Language Technology Assignment 2 from the
course's Swedish Selma Lagerlof corpus. Each JSONL row has an ngram string
array and its integer count. The three configurations expose unigram,
bigram, and trigram counts respectively.
The files were generated by 2-language_models.ipynb using Unicode-aware
tokenization and explicit <s> / </s> sentence markers.
Assignment-1-IndexerAssignment_3
Green Patent Detection: Multi-Agent HITL + PatentSBERTa
This repository contains an advanced green patent detection workflow built for binary classification of patent claims into:
1 = Green / climate mitigation related
0 = Non-green
The project extends a baseline PatentSBERTa workflow by adding a Human-in-the-Loop (HITL) review stage and a multi-agent debate system before final fine-tuning.
Project overview
The goal of this project is to improve green patent… See the full description on the dataset page: https://huggingface.co/datasets/Sristtee/Assignment_3.Assignment1-Sprakteknologiaudio_assignmentAssignment1edan20-assignment1EDAN20_Assignment1Assignment_3_Dataset
Assignment 3 Dataset — QLoRA/HITL Artifacts for Green Patent Detection
Dataset Summary
This repository contains the Assignment 3 data artifacts used and produced in the advanced QLoRA workflow for green patent detection, including:
top-100 uncertainty-selected claims
QLoRA reviewed outputs
final gold labels
Part C logs/summaries required by the assignment
Transparency Note on HITL Agreement Reporting
In this Assignment 3 run, i did not manually review and… See the full description on the dataset page: https://huggingface.co/datasets/CTB2001/Assignment_3_Dataset.edan20-assignment1assignment-2-datasetsAssignment2Assignment2assignment-2aEDAN20_Assignment2EDAN20_Assignment_2edan20-assignment-1assignment2LTassignment_dataset_llama_0221aviation-gate-assignment-arrival-bank-coherence-risk-v0.1What this repo is for
Detect when arrivals land
but cannot park.
Flags
dense arrival bank with low gate availability
slow gate turns creating holding
remote stand use rising without bank pressure
gate holds as a downstream delay trigger
