datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmmluplus
TMMLU+ : Large scale traditional chinese massive multitask language understanding
We present TMMLU+, a traditional Chinese massive multitask language understanding dataset. TMMLU+ is a multiple-choice question-answering dataset featuring 66 subjects, ranging from elementary to professional level.
The TMMLU+ dataset is six times larger and contains more balanced subjects compared to its predecessor, TMMLU. We have included benchmark results in TMMLU+ from closed-source models and 20… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/tmmluplus.Linguistic-Diagnostics-Syntax
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
LINDSEA Syntax only has an Indonesian (id) split, with additional splits containing fewshot examples. Below… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax.Linguistic-Diagnostics-Syntax-Judge
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
Data Sources
Data Source
License
Language/s
Split/s
CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.macula-hebrew-syntax
NuBerea MACULA Hebrew Syntax Trees (OT)
Full syntactic tree annotation of the Hebrew Bible from the MACULA Hebrew Linguistic Dataset, packaged as relational tables for computational biblical studies. The dataset covers word-level linguistic annotation (morphology, glosses, lexical semantics), sentence segmentation, and hierarchical syntactic structure (clauses and phrases with their roles and containment relations) over the Westminster Leningrad Codex base text.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-hebrew-syntax.macula-sblgnt-syntax
NuBerea MACULA Greek (SBLGNT) Syntax Trees (NT)
Full syntactic tree annotation of the Greek New Testament from the MACULA Greek SBLGNT edition. Relational tables cover word-level tokens with morphological, semantic, and cross-language features; sentence boundaries; word groups (clauses and phrases) with syntactic rules and roles; the word-group hierarchy; and word-group membership. Together they let researchers traverse the full syntax tree of every sentence in the New Testament… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-sblgnt-syntax.syntaxgymaction-roleplay-data
Action Roleplay Data
Data package for the Action SA-MP Android client.
The client connects to 92.119.165.177:5636. The files/ directory contains the extracted game data, cache.zip is the archive consumed by the initial installer, files.json is the file-by-file manifest, and client_config.json contains the public endpoints. Runtime logs were excluded from the distributable package.
The APK included here is a debug build for testing and is signed with a debug key.
swe-bench-opus-logs
Claude 3 inference SWE-Bench results
Contains prompting responses from SWE-bench on these 2 settings:
Oracle retrieval
BM25 retrieval
Each of the subsets contains an additional log_last_line attributes which is the last line from log files generated during evaluation step.
Results:
Model
BM25 Retrieval Resolved (%)
Oracle Retrieval Resolved (%)
GPT-4*
0
1.74
Claude-2
1.96
4.80
Claude-3 Opus (20240229)
3.24
6.42
Claude-2 and GPT-4 results from SWE-bench… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/swe-bench-opus-logs.syntaxgym_sentencesML-1M-Syntax-Validated-Python-Code
ML-1M Syntax-Validated Python Code
Dataset Summary
ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code.
The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.Code-Syntax-Expanded
Code-Syntax-Expanded
A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows.
📊 Dataset Overview
Property
Value
Total rows
5,000,000+
File size
~1.1 GB (uncompressed CSV)
Languages
33
Unique templates
160+ error patterns
Format
CSV (4 columns)
License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.Ultra-FineWeb-L3-zh-hant-translated
Ultra-FineWeb-L3 (Traditional Chinese Translation)
Traditional Chinese translation of the English openbmb/Ultra-FineWeb-L3 corpus, targeting Taiwan-standard Traditional Chinese (臺灣正體中文).
Background
Ultra-FineWeb-L3 is the L3-refined tier of the UltraData data management framework developed by OpenBMB. Starting from Ultra-FineWeb (the high-quality web corpus behind MiniCPM4), L3 refinement transforms raw web documents into two structured synthesis formats:
Q&A… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/Ultra-FineWeb-L3-zh-hant-translated.medical-billing-icd10-qammevol-zh-hant
MMEvol - Translated Chinese Traditional
A subset of Tongyi-ConvAI/MMEvol translated using yentinglin/Llama-3-Taiwan-70B-Instruct from english to traditional chinese.
Read the Note below before use.
Image source distribution:
Dataset
Count
Percentage
coco
6598
29.8%
Q-Instruct-DB
5856
26.4%
clevr
2383
10.8%
chartqa
1733
7.8%
hfdata
1296
5.9%
geo170k
706
3.2%
data_engine
6983.2%
mathvision
644
2.9%
docvqa
600
2.7%
alfworld
401
1.8%
arxivqa
337
1.5%… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/mmevol-zh-hant.syntax-gossip
Syntax Gossip (candidate)
Syntax IR records over Crownless event-to-speech gossip: exact source bytes,
UTF-8 byte spans, UDPipe UD dependencies, NP/VP chunks, clause spans with
hearsay attribution links, and audit lineage per record.
Status: candidate (braid-syntax-gossip-v0.1.0-candidate.1). Not
frozen, authorizes no training. Cloud verification and gold review are
outstanding.
Splits: train.jsonl (4,000) / validation.jsonl (500) / test.jsonl (500).
Backend:… See the full description on the dataset page: https://huggingface.co/datasets/ratimics/syntax-gossip.weather-prediction-prototype-aws
Weather prediction prototype database.
This database was made using data provided by KMI.
This database will only be used to train a prototype.
Dataset Details
Dataset Description
Dataset Sources [optional]
KMI
Dataset Structure
Normalized columns:
timestamp
air_pressure
relative_humidity
precipitation
wind_speed
wind_direction
More information about these columns can be found in the information_10min.txt file.
reasoning-conversations
Multilingual Reasoning Dataset
Include languages from German, Korean, Spanish, Japanese, French, Simplified Chinese, Traditional Chinese
Reasoning traces from Deepseek-v3-R1, Deepseek-v3-R1-Zero
Credits sponsored by Currents API
czech-punctuation-pos-syntax
Czech Punctuation, POS and Syntactic Dataset 🇨🇿
A High-Quality Dataset for Punctuation Restoration and Neuro-Symbolic LLM Grounding
This dataset is a structured, linguistically annotated corpus of the Czech language, specifically designed for Punctuation Restoration tasks, Part-of-Speech (POS) tagging, and token-level syntax embedding (such as nanoGPT custom metadata training).
Unlike pure raw text corpora, this dataset provides a deterministic 1:1 token-level mapping… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/czech-punctuation-pos-syntax.medical-billing-icd10-qa-sample
Medical Billing & ICD-10 Synthetic Dataset (Sample
🚀 NEED THE FULL ENTERPRISE COMMERCIAL DATASET?
Get instant access to the full 50,000+ cleaned JSONL dataset for fine-tuning production models:
🏥 50,000+ verified ICD-10 / CPT billing scenarios
📄 Clean JSONL format (instruction, input, output)
🔒 Safe for HIPAA/GDPR—100% synthetic, zero real patient data
💼 Full commercial license for SaaS and Enterprise applications
👉 Buy Full Enterprise Dataset ($249) - Instant Download… See the full description on the dataset page: https://huggingface.co/datasets/Builder-syntaxlabs/medical-billing-icd10-qa-sample.syntaxgym-hexatagged
Dataset Card for "syntaxgym-hexatagged"
More Information needed
Code-Syntax-Expanded
Code-Syntax-Expanded
A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows.
📊 Dataset Overview
Property
Value
Total rows
5,000,000+
File size
~1.1 GB (uncompressed CSV)
Languages
33
Unique templates
160+ error patterns
Format
CSV (4 columns)
License… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax-Expanded.reprompts-20k-sample
Reprompts of conversations using Opus-20240229 20k samples
Sauce:
lmsys/lmsys-chat-1m - en only
allenai/WildChat-1M - en only
teknium/OpenHermes-2.5
teknium/OpenHermes-2.5
ShareGPT
Code-Syntax
Code Syntax Dataset (S)
A large-scale, high‑quality dataset for teaching large language models to identify and correct common syntax errors across 30+ programming languages.Contains 500,000+ unique examples (≈110 MB) with English explanations – no artificial padding.
📊 Dataset Format
The dataset is provided as a single CSV file with the following columns:
Column
Type
Description
wrong_code
string
Code snippet containing a syntax error
correct_code… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax.Essay-Syntax-Instructsyntaxgym-hexataggedinstruct_code_cleaning
SFT code dataset building
Contain a list of tasks useful when building a iniitial dataset source:
reverse_translation
Given a history of conversations, what would the human ask next?
reverse_translation_first_round
Suppose you already have a response, the LLM must predict what question does the human asked
clean_code
Given a code snippet, it determines whether its useful and atomic enough to be use for a response by LLM
gen_code_question
Generates a question given a… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/instruct_code_cleaning.rag_pipelinedeveloper-portfolio-ragenhanced-ud-syntax
Enhanced Universal Dependencies (syntax only) dataset
This repo aggregates syntactic markup from UD_English-EWT and UD_English-GUM datasets.
Changes made to the source datasets:
Only syntactic tags (head, deprel, deps) are preserved
Short sentences (fewer than 3 tokens) are removed
Source datasets:
UD_English-EWT: https://github.com/UniversalDependencies/UD_English-EWT
UD_English-GUM: https://github.com/UniversalDependencies/UD_English-GUM
syntax_analysis
