datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmmluplus
TMMLU+ : Large scale traditional chinese massive multitask language understanding
We present TMMLU+, a traditional Chinese massive multitask language understanding dataset. TMMLU+ is a multiple-choice question-answering dataset featuring 66 subjects, ranging from elementary to professional level.
The TMMLU+ dataset is six times larger and contains more balanced subjects compared to its predecessor, TMMLU. We have included benchmark results in TMMLU+ from closed-source models and 20… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/tmmluplus.Linguistic-Diagnostics-Syntax
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
LINDSEA Syntax only has an Indonesian (id) split, with additional splits containing fewshot examples. Below… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax.Linguistic-Diagnostics-Syntax-Judge
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
Data Sources
Data Source
License
Language/s
Split/s
CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.macula-hebrew-syntax
NuBerea MACULA Hebrew Syntax Trees (OT)
Full syntactic tree annotation of the Hebrew Bible from the MACULA Hebrew Linguistic Dataset, packaged as relational tables for computational biblical studies. The dataset covers word-level linguistic annotation (morphology, glosses, lexical semantics), sentence segmentation, and hierarchical syntactic structure (clauses and phrases with their roles and containment relations) over the Westminster Leningrad Codex base text.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-hebrew-syntax.macula-sblgnt-syntax
NuBerea MACULA Greek (SBLGNT) Syntax Trees (NT)
Full syntactic tree annotation of the Greek New Testament from the MACULA Greek SBLGNT edition. Relational tables cover word-level tokens with morphological, semantic, and cross-language features; sentence boundaries; word groups (clauses and phrases) with syntactic rules and roles; the word-group hierarchy; and word-group membership. Together they let researchers traverse the full syntax tree of every sentence in the New Testament… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-sblgnt-syntax.maandagtestThis dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 49,
"total_frames": 23136,
"total_tasks": 1,
"total_videos": 147,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:49"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/maandagtest.ML-1M-Syntax-Validated-Python-Code
ML-1M Syntax-Validated Python Code
Dataset Summary
ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code.
The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.DisndagTestDoezo67This dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 25,
"total_frames": 14789,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/DisndagTestDoezo67.Ultra-FineWeb-L3-zh-hant-translated
Ultra-FineWeb-L3 (Traditional Chinese Translation)
Traditional Chinese translation of the English openbmb/Ultra-FineWeb-L3 corpus, targeting Taiwan-standard Traditional Chinese (臺灣正體中文).
Background
Ultra-FineWeb-L3 is the L3-refined tier of the UltraData data management framework developed by OpenBMB. Starting from Ultra-FineWeb (the high-quality web corpus behind MiniCPM4), L3 refinement transforms raw web documents into two structured synthesis formats:
Q&A… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/Ultra-FineWeb-L3-zh-hant-translated.reasoning-conversations
Multilingual Reasoning Dataset
Include languages from German, Korean, Spanish, Japanese, French, Simplified Chinese, Traditional Chinese
Reasoning traces from Deepseek-v3-R1, Deepseek-v3-R1-Zero
Credits sponsored by Currents API
medical-billing-icd10-qammevol-zh-hant
MMEvol - Translated Chinese Traditional
A subset of Tongyi-ConvAI/MMEvol translated using yentinglin/Llama-3-Taiwan-70B-Instruct from english to traditional chinese.
Read the Note below before use.
Image source distribution:
Dataset
Count
Percentage
coco
6598
29.8%
Q-Instruct-DB
5856
26.4%
clevr
2383
10.8%
chartqa
1733
7.8%
hfdata
1296
5.9%
geo170k
706
3.2%
data_engine
6983.2%
mathvision
644
2.9%
docvqa
600
2.7%
alfworld
401
1.8%
arxivqa
337
1.5%… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/mmevol-zh-hant.czech-punctuation-pos-syntax
Czech Punctuation, POS and Syntactic Dataset 🇨🇿
A High-Quality Dataset for Punctuation Restoration and Neuro-Symbolic LLM Grounding
This dataset is a structured, linguistically annotated corpus of the Czech language, specifically designed for Punctuation Restoration tasks, Part-of-Speech (POS) tagging, and token-level syntax embedding (such as nanoGPT custom metadata training).
Unlike pure raw text corpora, this dataset provides a deterministic 1:1 token-level mapping… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/czech-punctuation-pos-syntax.instruct_code_cleaning
SFT code dataset building
Contain a list of tasks useful when building a iniitial dataset source:
reverse_translation
Given a history of conversations, what would the human ask next?
reverse_translation_first_round
Suppose you already have a response, the LLM must predict what question does the human asked
clean_code
Given a code snippet, it determines whether its useful and atomic enough to be use for a response by LLM
gen_code_question
Generates a question given a… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/instruct_code_cleaning.reprompts-20k-sample
Reprompts of conversations using Opus-20240229 20k samples
Sauce:
lmsys/lmsys-chat-1m - en only
allenai/WildChat-1M - en only
teknium/OpenHermes-2.5
teknium/OpenHermes-2.5
ShareGPT
omy_f3m_TestWoensdagSingle1453This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 25,
"total_frames": 19149,
"total_tasks": 1,
"total_videos": 75,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_TestWoensdagSingle1453.syntaxgym-hexatagged
Dataset Card for "syntaxgym-hexatagged"
More Information needed
omy_f3m_Donderdag10cmModel1206This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 99,
"total_frames": 46759,
"total_tasks": 1,
"total_videos": 198,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:99"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_Donderdag10cmModel1206.omy_f3m_PickNPlaceDisndag1420This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 50,
"total_frames": 27949,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_PickNPlaceDisndag1420.omy_f3m_BlokjegroenpakkenMaandag1333This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 20,
"total_frames": 9641,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_BlokjegroenpakkenMaandag1333.Essay-Syntax-InstructDinsdagV2This dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 50,
"total_frames": 27949,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/DinsdagV2.rag_pipelineMove_Cube_v1.3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 25,
"total_frames": 13927,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SyntaxError-v2/Move_Cube_v1.3.100episodesDonderdag1This dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 99,
"total_frames": 46759,
"total_tasks": 1,
"total_videos": 198,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:99"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/100episodesDonderdag1.syntaxgym-hexataggedomy_f3m_WoensdagTestMulti1540This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 31,
"total_frames": 7357,
"total_tasks": 3,
"total_videos": 93,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_WoensdagTestMulti1540.omy_f3m_PickNPlaceDisndag1411This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 1,
"total_frames": 1500,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_PickNPlaceDisndag1411.WoensdagV1This dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 49,
"total_frames": 30967,
"total_tasks": 1,
"total_videos": 98,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:49"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/WoensdagV1.testThis dataset was created using Physical AI Tools and LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m",
"total_episodes": 5,
"total_frames": 3719,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/test.
