datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
BACHI_Chord_Recognition
BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
Paper | Project Page | Code | POP909-CL Dataset
This repository contains trained model weights and classical datasets for the paper:
Mingyang Yao, Ke Chen, Shlomo Dubnov and Taylor Berg-Kirkpatrick
"BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music."ICASSP 2026, 2025
Abstract
Automatic chord… See the full description on the dataset page: https://huggingface.co/datasets/Itsuki-music/BACHI_Chord_Recognition.pxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.Food121
Dataset Details
Dataset Description
This dataset is the combination of the Food101, Indian Food Classification and The-massive-Indian-Food-Dataset datasets.
This Dataset aims to be a viable dataset for Image Classification of Foods with an added Indian context. This dataset has 121 classes with each class having 800 images in the train split and 200 images in the test split. Maximum resolution of images is 512*512.
The Food121-224 dataset has all images downscaled to a… See the full description on the dataset page: https://huggingface.co/datasets/ItsNotRohit/Food121.wikireading
Dataset Card for Wikireading
This is a dataset of book chapters scraped from a Russian website called Wikireading.
Dataset Details
Dataset Description
Wikireading is a collection of non-fiction educational books in various domains: Biology, Art, History, Religion and much more. The books are highly educational and provide vast knowledge in different domains, making this dataset a good choice for pretraining.
The resulting dataset contains ~26M rows, which in… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/wikireading.FinSearchCompThis repository contains the FinSearchComp dataset, a benchmark for evaluating financial search and reasoning capabilities of LLM-based agents, as presented in the paper FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning.
Project Page: https://randomtutu.github.io/FinSearchComp/
FinSearchComp is the first fully open-source agent benchmark designed for realistic, open-domain financial search and reasoning. It comprises three tasks that closely… See the full description on the dataset page: https://huggingface.co/datasets/itsakhilyou/FinSearchComp.NotAllCodeIsEqual
NotAllCodeIsEqual
This dataset was created for the paper Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning.
It contains code fine-tuning datasets split by complexity metrics for studying the relationship between code complexity and reasoning capabilities.
We provide 2 types of dataset, that cover complementary settings:
CodeNet (solution-driven complexity):
The CodeNet splits contain the same programming problems across all complexity levels, but with… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/NotAllCodeIsEqual.smishing-syntheticIT_Support_V2
Mack: IT Support & Admin Dataset
📋 Dataset Description
This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent.
The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging.
Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.synthetic-logs
Synthetic Logs (Wild)
Just messy, realistic-looking logs paired with their parsed version. Each row has a raw log line and what you'd want a parser to pull out of it.
100,000 rows total, split across 3 files in data/:
data/logs-0001.parquet
data/logs-0002.parquet
data/logs-0003.parquet
2 columns: raw_log (the messy string) and parsed_json (the answer, as JSON string)
Covers 130+ services — nginx, postgres, k8s, lambda, python tracebacks, etc. — in 16 formats like syslog, JSON… See the full description on the dataset page: https://huggingface.co/datasets/itsrishub/synthetic-logs.task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task755_find_longest_substring_and_replace_its_sorted_lowercase_version_in_both_lists.sepediMedAgentSim-datasets
MedAgentSim Datasets
GitHub: https://github.com/MAXNORM8650/MedAgentSimWebsite: https://medagentsim.netlify.app
This repository contains various datasets used in the MedAgentSim project for simulating medical agent interactions.
Datasets Included
Dataset
Rows
Description
medqa_v1.parquet
107
General medical question-answering OSCE examinations
medqa_extended_v1.parquet
214
Extended medical QA with comprehensive coverage
mimiciv_v1.parquet
288
Patient… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/MedAgentSim-datasets.itsm_ticketsFood121-224
Dataset Details
Dataset Description
This dataset is the downscaled version of the Food121 dataset. All images are downscaled to a maximum of 224*224.
This dataset is the combination of the Food101, Indian Food Classification and The-massive-Indian-Food-Dataset datasets.
This Dataset aims to be a viable dataset for Image Classification of Foods with an added Indian context. This dataset has 121 classes with each class having 800 images in the train split and 200 images in… See the full description on the dataset page: https://huggingface.co/datasets/ItsNotRohit/Food121-224.pcbdraft_nbaarabic-itsm-dataset
Arabic ITSM Dataset
A synthetic dataset of 10,000 Arabic IT support tickets, labeled with a structured 3-level ITSM taxonomy, generated using LLMs, and validated programmatically before release.
Tickets are written in Egyptian Arabic (عامية مصرية) and cover the full range of helpdesk scenarios: access issues, network problems, hardware faults, software errors, security incidents, and service requests. Arabic technical vocabulary is mixed with English terms as they naturally… See the full description on the dataset page: https://huggingface.co/datasets/albaz2000/arabic-itsm-dataset.bluesky-10m-posts-15-languages
Dataset Card: Bluesky 10M Multilingual
📊 Overview
Total Posts: 10,099,990
Languages: 15 (en, tr, es, pt, de, fr, ja, it, nl, pl, ru, ko, zh, ar, hi)
Collection Period: August 9-12, 2026
Source: Bluesky Jetstream API (public firehose)
Format: JSONL
Size: ~3 GB
🌍 Language Distribution
Language
Code
Posts
%
English
en
6,843,995
67.8%
Japanese
ja
1,547,179
15.3%
German
de
373,626
3.7%
Portuguese
pt
331,093
3.3%
Spanish
es
325,865… See the full description on the dataset page: https://huggingface.co/datasets/itsmebatuhan/bluesky-10m-posts-15-languages.proper-agents-data
ProPer Agents — data
Data for ProPer Agents: Proactivity Driven Personalized Agents for Advancing
Knowledge Gap Navigation (ACL 2026).
Paper ·
Adapters
Three domains: code, medical, pwab (product recommendation).
Layout
{domain}/
raw/train.jsonl source examples
raw/test.jsonl
raw/{domain}_rga_{train,test}.jsonl RGA SFT data (Alpaca format)
raw/{domain}_dga_{train,test}.jsonl DGA SFT data (Alpaca format)… See the full description on the dataset page: https://huggingface.co/datasets/itsgupta/proper-agents-data.ultralivecell
Cell Segmentation Masks Dataset
This repository contains NumPy mask files generated using AutoSeg-SAM2 for cell segmentation experiments.
Dataset Structure
The dataset is organized by experiment type and image number:
{experiment_type}/
└── image_{number}/
└── cached_masks.npy
Experiments Included
Cdx2_Gata6_Oct4
Images: image_2, image_3, image_4
Sox2_Sox17
Images: image_1, image_2, image_3
Usage
To load the NumPy mask… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/ultralivecell.ITSC
ITSC
Park-vector current loci of a three-phase induction motor, as a four-class stator
winding task: normal / phase_A / phase_B / phase_C. 183 records.
reasoning is empty here; the twin repo ITSC-annotated is identical except that field is filled.
The reading
Both coordinates are divided by the radius of the equal-area circle, so that circle is
the 1.0 ring on every image and the tick numbers carry no current amplitude at all. Two
steps, and both are drawn on… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/ITSC.ITSC-perception
Roles
Roles: perception view of ITSC — annot is the source label (normal / phase_A / phase_B / phase_C), kept machine-parseable as the gold for verification and reward parsing; the model reads query + image, where the repo ships an induction motor's stator-current magnitude in four image encodings as four equal-sized configs — reshaped (consecutive samples arranged as the rows of a grayscale square), scalogram (a continuous-wavelet time-scale view), spectrogram (a short-time… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/ITSC-perception.TACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual inspections.… See the full description on the dataset page: https://huggingface.co/datasets/ITS23/TACK_Tunnel_Data.nepali-asr-processedit-support-llmepfl-enterprise-osai-adoption-research-data
EPFL Enterprise Open-Source AI Adoption Research Dataset
Dataset Summary
This dataset contains mixed-methods research data from 100 organizations regarding their strategic adoption of open-source AI through the Hugging Face ecosystem. The research was conducted at EPFL (École Polytechnique Fédérale de Lausanne) and supports the development of the Gate-Lever framework for enterprise open-source AI adoption.
Dataset Structure
This dataset is organized into 4… See the full description on the dataset page: https://huggingface.co/datasets/itseffi/epfl-enterprise-osai-adoption-research-data.ConvES30K
ConvES10K-HQ-LLM — High-Quality Spanish Conversations (LLM-Generated)
Recommended for AI training. This version replaces the template-based builds and is fully LLM-generated for true independence.
Why this version
Previous builds (30K and 10K-template) were template-based (gen.py + 178 situations):
13,156 distinct messages from 74,526 total → 82.3% duplicate messages, max x109 on narrative blocks (Mesa seis pegada a la mesa siete...)
Coherence failures from… See the full description on the dataset page: https://huggingface.co/datasets/itsZyn/ConvES30K.privasis-reasoning-qa
Privasis Reasoning-QA
Open-ended reasoning question–answer pairs derived from the
NVIDIA Privasis-Zero dataset.
Two configs are provided:
qa50k — 50,000 pairs sampled from the Privasis-Zero corpus split (record field). Main set.
qa500 — 500 pairs from the hard_test split (original_record field). Original pilot.
from datasets import load_dataset
ds = load_dataset("ItsMaxNorm/privasis-reasoning-qa", "qa50k", split="train")
Each item presents one question that requires… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/privasis-reasoning-qa.bigger-ru-book
