datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
huggingface-spaces-codes
📊 Dataset Description
This dataset comprises code files of Huggingface Spaces that have more than 0 likes as of November 10, 2023. This dataset contains various programming languages totaling in 672 MB of compressed and 2.05 GB of uncompressed data.
📝 Data Fields
Field
Type
Description
repository
string
Huggingface Spaces repository names.
sdk
string
Software Development Kit of the space.
license
string
License type of the space.… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/huggingface-spaces-codes.sms-spam-collection
SMS Spam Collection v.1
DESCRIPTION
The SMS Spam Collection v.1 (hereafter the corpus) is a set of SMS tagged messages that have been collected for SMS Spam research. It contains one set of SMS messages in English of 5,574 messages, tagged acording being ham (legitimate) or spam.
1.1. Compilation
This corpus has been collected from free or free for research sources at the Web:
A collection of between 425 SMS spam messages extracted manually from the Grumbletext Web… See the full description on the dataset page: https://huggingface.co/datasets/codesignal/sms-spam-collection.tsla-historic-pricescode_search_net_python_10000_examplesharbor-benchwarn-act-notice-type-codes-crosswalk
WARN Act notice-type codes — the crosswalk
Every US state publishes WARN Act layoff notices with a free-text column saying
what kind of event it is. The statute recognises two: a plant closing and a
mass layoff. Across 48 states that column contains
552 distinct exact strings (531
once you fold case).
This dataset is the crosswalk: one row per raw string, how many notices carry
it, which states emit it, and what it normalizes to.
The finding that matters
521 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.CodeSearchNet-Pythonlanguage_codes_marianMTappliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.python_codes_samplequestion-answernaturalnesshvac-error-codes
HVAC Bench Error Code Dataset
Error codes from HVAC equipment sold in the United States, United Kingdom, and Europe, as published by HVAC Bench. Each record gives the manufacturer, the code, the product family the definition applies to, a plain-language meaning, the checks an owner can safely make, the point at which a technician is needed, and the date the definition was last checked against manufacturer documentation. Codes are specific to a product family and are not… See the full description on the dataset page: https://huggingface.co/datasets/mukarram360/hvac-error-codes.neuromoyo-sahara-codeswitch-benchmark
NEUROMOYO — Sahara CodeSwitch Africa Benchmark
🔗 Live Benchmark Results
Interactive benchmark:
https://www.neuromoyo.app/benchmark
This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech.
🚀 Live NEUROMOYO Demo
Live application:
https://www.neuromoyo.app
The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.Code-Syntax-Expanded
Code-Syntax-Expanded
A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows.
📊 Dataset Overview
Property
Value
Total rows
5,000,000+
File size
~1.1 GB (uncompressed CSV)
Languages
33
Unique templates
160+ error patterns
Format
CSV (4 columns)
License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.code-switching-codesaviours-si26-hamzaMoroccan-Codeswitching
Moroccan Darija Code-Switched Corpus (Sentence-level TSV)
Dataset Summary
This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text.
Languages
The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.code-switching-codesaviours-si26-muhammadahmad
Code-Switching Codesaviours SI26 — Muhammad Ahmad
Dataset Description
This dataset contains 155 naturally occurring Roman Urdu–English
code-switched sentences (1,400+ word-level entries), reflecting how
Roman Urdu and English are mixed in everyday informal communication by
Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit).
Code-switching — alternating between two or more languages within a single
sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.ghs-pictograms-h-p-codes-2026
Canonical landing page: https://www.smartqhse.com/datasets/ghs-pictograms-h-p-codes-2026
GHS Hazard Pictograms + H-Statement + P-Statement Codes 2026
Consolidated UN GHS Rev. 10 reference — 9 GHS hazard pictograms (GHS01–GHS09), ~80 H-codes (hazard statements) and ~90 P-codes (precautionary statements) with category tags (physical / health / environmental / general / prevention / response / storage / disposal). Used for SDS authoring, label compliance, REACH/CLP/HazCom 2012… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/ghs-pictograms-h-p-codes-2026.code-switching-codesaviours-si26-SanaCodeSearchNetsentinelng-data-grid
SentinelNG Data Grid Dataset
This repository contains a CSV-based data-grid resource for SentinelNG. The primary artifact is a tabular dataset intended to support grid, location, or structured alert workflows in the SentinelNG ecosystem.
Working with the data
Load the CSV with a tool that preserves column names and types, then inspect missing values, coordinate or identifier semantics, and duplicate rows before use. Do not infer geographic or security meaning from… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/sentinelng-data-grid.anime-fighting-simulator-codes
Anime Fighting Simulator codes, as data
A small, CC0 dataset of reward codes for the Roblox experience Anime Fighting Simulator, published as a flat CSV so it can be diffed, joined, or charted without scraping a web page first.
The point of this repository is not the list. Plenty of sites publish the list. The point is that this one records what is known and what is not, as separate facts.
What is in here
File
What it is
codes.csv
The data. One row per… See the full description on the dataset page: https://huggingface.co/datasets/x15987856/anime-fighting-simulator-codes.code-switching-codesaviours-si26-warood
Code-Switching Codesaviours SI-26 Dataset
Dataset Description
This dataset contains 150 naturally occurring Roman Urdu–English code-switched sentences, commonly spoken by Pakistani speakers in casual, everyday communication. Each sentence has been broken down into individual words, and every word is labeled by language.
Roman Urdu–English code-switching is extremely common in Pakistan (spoken/written by an estimated 230+ million people) but is poorly handled by… See the full description on the dataset page: https://huggingface.co/datasets/waroodzkhan/code-switching-codesaviours-si26-warood.ETF-CodeSumEval
CodeSumEval Dataset
An annotated dataset for studying hallucination in code summarization. Each sample consists of a Java code snippet, a generated summary from a large language model, and detailed annotations marking entity-level correctness and hallucination causes.
📖 Overview
CodeSumEval is a first-of-its-kind dataset designed to evaluate and analyze hallucinations in code summarization. It comprises:
411 generated summaries of Java methods, produced by 7 different… See the full description on the dataset page: https://huggingface.co/datasets/kishanmaharaj/ETF-CodeSumEval.windows-event-codes-qandaI used my notebook here to generate CSV files for Windows event codes:
https://github.com/whit3rabbit/Windows-Event-Codes-CSV
CSV is here: https://github.com/whit3rabbit/Windows-Event-Codes-CSV/blob/main/updated_detailed_events.csv
Converted each line to markdown and used it to generate Questions and Answers. These have not been vetted for accuracy so use with caution.
code-switching-codesaviours-si26-bilal
Roman Urdu-English Code Switching Dataset
Dataset Description
A manually labeled dataset of 160+ code-switching sentences where Roman Urdu and English are naturally mixed — reflecting how 230 million Pakistanis actually communicate online.
Each word in every sentence is tagged with a language label, making this dataset suitable for token-level language identification and code-switching NLP research.
Label Meanings
Label
Description
Examples… See the full description on the dataset page: https://huggingface.co/datasets/Noisy77/code-switching-codesaviours-si26-bilal.ground-truth-ob
Ground Truth OB
This repository contains ground_truth_kb.csv, a tabular ground-truth or knowledge-base resource. The current repository is deliberately small and contains no executable training or evaluation script.
Recommended use
Load the CSV, inspect its column names and encoding, validate identifiers and labels, and record the provenance of every ground-truth field before joining it with model outputs. Keep an immutable copy of the raw file and create derived… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/ground-truth-ob.id-en-codeswitch-dataset-alternative
Indonesian–English Code-Switching Synthetic Speech Dataset
Synthetic speech generated for the undergraduate final project "Handling
Code-Switching in Automatic Speech Recognition for Low-Resource Language
Pairs: An Indonesian–English Case Study", School of Electrical Engineering
and Informatics, Institut Teknologi Bandung.
This dataset contains synthetic audio produced from the Indonesian–English
code-switching text corpora released in the companion repository below. It
was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.bs-mapsThe data about the built in maps in Beat Saber. Contains all OST, Camellia, and Extra songs. A couple of songpacks are added.
