datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-course-final-assignment
Agent Course Final Assignment - Unified Dataset
Author: Arte(r)m Sedov
GitHub: https://github.com/arterm-sedov/
Project link: https://huggingface.co/spaces/arterm-sedov/agent-course-final-assignment
Dataset Description
This dataset is produced by the GAIA Unit 4 Agent for the Hugging Face Agents Course final assignment as part of an experimental multi-LLM agent system that demonstrates advanced AI agent capabilities. It demonstrates advanced AI agent capabilities for… See the full description on the dataset page: https://huggingface.co/datasets/arterm-sedov/agent-course-final-assignment.NLP_Assignment_1Assignment3DISC-Assignment本仓库包含如下内容:
1.报告 2.训练代码 3.测试结果 4.指令构造数据集 5.任务书
报告是ACL2023格式的PDF文件Reading_Comprehension_LLM
训练代码在Finetune文件中,详细信息可看文件夹中的README。
指令数据集构造代码与最终数据集在Dataset文件中,详细信息可看文件夹中的README。
实验代码与实验结果在Experiment文件中,详细信息可看文件夹中的README。
任务书即DISC-Assignment,包含实验的指导与具体要求。
adaption-resource-assignment-configs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-resource_assignment_configs
This dataset contains pairs of natural language resource allocation problems and their corresponding JSON configuration outputs. Each sample defines constraints such as categorical matching, spatial proximity, capacity limits, or priority scoring based on provided data schemas. The configurations specify goals with weights, award types, and logic… See the full description on the dataset page: https://huggingface.co/datasets/intellign/adaption-resource-assignment-configs.spoken_norm_assignment
VietAI assignment: Vietnamese Inverse Text Normalization dataset
Dataset Description
Inverse text normalization (ITN) is the task that transforms spoken to written styles. It is particularly useful in automatic speech recognition (ASR) systems where proper names are often miss-recognized by their pronunciations instead of the written forms. By applying ITN, we can improve the readability of the ASR system’s output significantly. This dataset provides data for doing ITN… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/spoken_norm_assignment.deductive_logical_reasoning-room_assignmentPKU_NLPDL_Assignment1edan20-assignment2-selma-ngramsDMT_Assignment_2AAI_assignment2dom-formula-assignment-data
DOM Formula Assignment Dataset
Training and Testing Data for A Machine Learning and Benchmarking Approach for Molecular Formula Assignment of Ultra High-Resolution Mass Spectrometry Data from Complex Mixtures
Paper: Under review
Abstract
A machine learning approach to molecular formula assignment is crucial for unlocking the full potential of ultra-high resolution mass spectrometry (UHRMS) when analyzing complex mixtures. By combining data-driven models with… See the full description on the dataset page: https://huggingface.co/datasets/SaeedLab/dom-formula-assignment-data.maigurski-customer-personality-assignment1
Customer Personality Analysis – EDA Results
1. Project Goal
The goal of this project is to use numeric-focused Exploratory Data Analysis (EDA) on the Customer Personality Analysis dataset to understand:
Which customer characteristics are associated with higher spending.
How these characteristics differ between customers who responded to the last marketing campaign and those who did not.
The main outcome variable is:
Response (0 = no, 1 = yes) – did the customer respond… See the full description on the dataset page: https://huggingface.co/datasets/maigurski/maigurski-customer-personality-assignment1.task-assignment-rolloutsmaigurski-customer-personality-assignment1
Customer Personality Analysis – EDA Results
1. Project Goal
The goal of this project is to use numeric-focused Exploratory Data Analysis (EDA) on the Customer Personality Analysis dataset to understand:
Which customer characteristics are associated with higher spending.
How these characteristics differ between customers who responded to the last marketing campaign and those who did not.
The main outcome variable is:
Response (0 = no, 1 = yes) – did the customer… See the full description on the dataset page: https://huggingface.co/datasets/advsp/maigurski-customer-personality-assignment1.slm-assignment-data
Grade-Level Vocabulary-Locked Writing Tutor — Dataset
Training and evaluation data for fine-tuning a small open model (Qwen3-0.6B)
into a grade 7–8 writing/grammar tutor whose vocabulary and sentence
complexity stay locked to the band — it introduces at most one word above
grade level per reply (always immediately defined) and never escalates, even
under pressure ("use bigger words", "give me the college version") or
jailbreak-style attacks.
The dataset is the deliverable. ~80%… See the full description on the dataset page: https://huggingface.co/datasets/blackbird0831/slm-assignment-data.selma-ngrams-assignment2
Selma Lagerlof n-gram counts
This dataset was produced for Language Technology Assignment 2 from the
course's Swedish Selma Lagerlof corpus. Each JSONL row has an ngram string
array and its integer count. The three configurations expose unigram,
bigram, and trigram counts respectively.
The files were generated by 2-language_models.ipynb using Unicode-aware
tokenization and explicit <s> / </s> sentence markers.
Assignment-1-IndexerAssignment_3
Green Patent Detection: Multi-Agent HITL + PatentSBERTa
This repository contains an advanced green patent detection workflow built for binary classification of patent claims into:
1 = Green / climate mitigation related
0 = Non-green
The project extends a baseline PatentSBERTa workflow by adding a Human-in-the-Loop (HITL) review stage and a multi-agent debate system before final fine-tuning.
Project overview
The goal of this project is to improve green patent… See the full description on the dataset page: https://huggingface.co/datasets/Sristtee/Assignment_3.edan20-assignment1-index-selma-novelsassignment-movie-posters
Movie Posters Audio Text Data Notes
Dataset summary
A documented Movie Posters data-preparation workflow for Audio Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/joh-benn/assignment-movie-posters.Assignment1-Sprakteknologiaudio_assignmentassignment-document-ocr40
Document OCR Sensor Fusion Data Notes
Dataset summary
This data card accompanies a lightweight Document OCR loader for Sensor Fusion metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/darrenten/assignment-document-ocr40.Assignment1edan20-assignment1EDAN20_Assignment1Assignment_3_Dataset
Assignment 3 Dataset — QLoRA/HITL Artifacts for Green Patent Detection
Dataset Summary
This repository contains the Assignment 3 data artifacts used and produced in the advanced QLoRA workflow for green patent detection, including:
top-100 uncertainty-selected claims
QLoRA reviewed outputs
final gold labels
Part C logs/summaries required by the assignment
Transparency Note on HITL Agreement Reporting
In this Assignment 3 run, i did not manually review and… See the full description on the dataset page: https://huggingface.co/datasets/CTB2001/Assignment_3_Dataset.assignment-comics
Comics Pointcloud Text Data Notes
Dataset summary
This data card accompanies a lightweight Comics loader for Pointcloud Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/ashishnan/assignment-comics.assignment-geology
Geology Audio Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Geology work with Audio Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/ffernandezalejandro/assignment-geology.
