CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01syntaxsynth /tmmluplus TMMLU+ : Large scale traditional chinese massive multitask language understanding We present TMMLU+, a traditional Chinese massive multitask language understanding dataset. TMMLU+ is a multiple-choice question-answering dataset featuring 66 subjects, ranging from elementary to professional level. The TMMLU+ dataset is six times larger and contains more balanced subjects compared to its predecessor, TMMLU. We have included benchmark results in TMMLU+ from closed-source models and 20… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/tmmluplus.text10K<n<100K0 likes1.6k downloads1y agoHugging Face02aisingapore /Linguistic-Diagnostics-Syntaxgated LINDSEA Syntax LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian. Supported Tasks and Leaderboards LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs). Languages Indonesian (id) Dataset Details LINDSEA Syntax only has an Indonesian (id) split, with additional splits containing fewshot examples. Below… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax.texttext-generationn<1K0 likes1.3k downloads9mo agoHugging Face03aisingapore /Linguistic-Diagnostics-Syntax-Judgegated LINDSEA Syntax LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian. Supported Tasks and Leaderboards LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs). Languages Indonesian (id) Dataset Details Data Sources Data Source License Language/s Split/s CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.textn<1K0 likes862 downloads2mo agoHugging Face04NuBerea /macula-hebrew-syntaxgated NuBerea MACULA Hebrew Syntax Trees (OT) Full syntactic tree annotation of the Hebrew Bible from the MACULA Hebrew Linguistic Dataset, packaged as relational tables for computational biblical studies. The dataset covers word-level linguistic annotation (morphology, glosses, lexical semantics), sentence segmentation, and hierarchical syntactic structure (clauses and phrases with their roles and containment relations) over the Westminster Leningrad Codex base text. This repository… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-hebrew-syntax.tabulartoken-classification1M<n<10M1 likes546 downloads13d agoHugging Face05NuBerea /macula-sblgnt-syntaxgated NuBerea MACULA Greek (SBLGNT) Syntax Trees (NT) Full syntactic tree annotation of the Greek New Testament from the MACULA Greek SBLGNT edition. Relational tables cover word-level tokens with morphological, semantic, and cross-language features; sentence boundaries; word groups (clauses and phrases) with syntactic rules and roles; the word-group hierarchy; and word-group membership. Together they let researchers traverse the full syntax tree of every sentence in the New Testament… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-sblgnt-syntax.tabulartoken-classification1M<n<10M0 likes540 downloads13d agoHugging Face06Syntax-Terror-BV /maandagtestThis dataset was created using Physical AI Tools and LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 49, "total_frames": 23136, "total_tasks": 1, "total_videos": 147, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:49" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/maandagtest.tabularrobotics10K<n<100K0 likes187 downloads3mo agoHugging Face07Noushad999 /ML-1M-Syntax-Validated-Python-Code ML-1M Syntax-Validated Python Code Dataset Summary ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code. The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.texttext-generation1M<n<10M0 likes114 downloads8mo agoHugging Face08Syntax-Terror-BV /DisndagTestDoezo67This dataset was created using Physical AI Tools and LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 25, "total_frames": 14789, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:25" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/DisndagTestDoezo67.tabularrobotics10K<n<100K0 likes113 downloads3mo agoHugging Face09syntaxsynth /Ultra-FineWeb-L3-zh-hant-translated Ultra-FineWeb-L3 (Traditional Chinese Translation) Traditional Chinese translation of the English openbmb/Ultra-FineWeb-L3 corpus, targeting Taiwan-standard Traditional Chinese (臺灣正體中文). Background Ultra-FineWeb-L3 is the L3-refined tier of the UltraData data management framework developed by OpenBMB. Starting from Ultra-FineWeb (the high-quality web corpus behind MiniCPM4), L3 refinement transforms raw web documents into two structured synthesis formats: Q&A… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/Ultra-FineWeb-L3-zh-hant-translated.text1M<n<10M0 likes105 downloads3mo agoHugging Face10syntaxsynth /reasoning-conversations Multilingual Reasoning Dataset Include languages from German, Korean, Spanish, Japanese, French, Simplified Chinese, Traditional Chinese Reasoning traces from Deepseek-v3-R1, Deepseek-v3-R1-Zero Credits sponsored by Currents API text10K<n<100K4 likes61 downloads2y agoHugging Face11syntaxlabs /medical-billing-icd10-qatextn<1K0 likes54 downloads17d agoHugging Face12syntaxsynth /mmevol-zh-hant MMEvol - Translated Chinese Traditional A subset of Tongyi-ConvAI/MMEvol translated using yentinglin/Llama-3-Taiwan-70B-Instruct from english to traditional chinese. Read the Note below before use. Image source distribution: Dataset Count Percentage coco 6598 29.8% Q-Instruct-DB 5856 26.4% clevr 2383 10.8% chartqa 1733 7.8% hfdata 1296 5.9% geo170k 706 3.2% data_engine 6983.2% mathvision 644 2.9% docvqa 600 2.7% alfworld 401 1.8% arxivqa 337 1.5%… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/mmevol-zh-hant.imagetext-generation10K<n<100K1 likes42 downloads2y agoHugging Face13KRadim /czech-punctuation-pos-syntax Czech Punctuation, POS and Syntactic Dataset 🇨🇿 A High-Quality Dataset for Punctuation Restoration and Neuro-Symbolic LLM Grounding This dataset is a structured, linguistically annotated corpus of the Czech language, specifically designed for Punctuation Restoration tasks, Part-of-Speech (POS) tagging, and token-level syntax embedding (such as nanoGPT custom metadata training). Unlike pure raw text corpora, this dataset provides a deterministic 1:1 token-level mapping… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/czech-punctuation-pos-syntax.texttoken-classification100K<n<1M0 likes37 downloads4mo agoHugging Face14syntaxsynth /instruct_code_cleaning SFT code dataset building Contain a list of tasks useful when building a iniitial dataset source: reverse_translation Given a history of conversations, what would the human ask next? reverse_translation_first_round Suppose you already have a response, the LLM must predict what question does the human asked clean_code Given a code snippet, it determines whether its useful and atomic enough to be use for a response by LLM gen_code_question Generates a question given a… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/instruct_code_cleaning.texttext-generation10K<n<100K1 likes35 downloads3y agoHugging Face15syntaxsynth /reprompts-20k-sample Reprompts of conversations using Opus-20240229 20k samples Sauce: lmsys/lmsys-chat-1m - en only allenai/WildChat-1M - en only teknium/OpenHermes-2.5 teknium/OpenHermes-2.5 ShareGPT text10K<n<100K0 likes35 downloads2y agoHugging Face16Syntax-Terror-BV /omy_f3m_TestWoensdagSingle1453This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 25, "total_frames": 19149, "total_tasks": 1, "total_videos": 75, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:25" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_TestWoensdagSingle1453.tabularrobotics10K<n<100K0 likes29 downloads3mo agoHugging Face17hriaz /syntaxgym-hexatagged Dataset Card for "syntaxgym-hexatagged" More Information needed text1K<n<10K0 likes28 downloads5mo agoHugging Face18Syntax-Terror-BV /omy_f3m_Donderdag10cmModel1206This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 99, "total_frames": 46759, "total_tasks": 1, "total_videos": 198, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:99" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_Donderdag10cmModel1206.tabularrobotics10K<n<100K0 likes26 downloads3mo agoHugging Face19Syntax-Terror-BV /omy_f3m_PickNPlaceDisndag1420This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 50, "total_frames": 27949, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_PickNPlaceDisndag1420.tabularrobotics10K<n<100K0 likes21 downloads3mo agoHugging Face20Syntax-Terror-BV /omy_f3m_BlokjegroenpakkenMaandag1333This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 20, "total_frames": 9641, "total_tasks": 1, "total_videos": 40, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_BlokjegroenpakkenMaandag1333.tabularrobotics1K<n<10K0 likes20 downloads3mo agoHugging Face21Kamyar-zeinalipour /Essay-Syntax-Instructtext1K<n<10K0 likes19 downloads2y agoHugging Face22Syntax-Terror-BV /DinsdagV2This dataset was created using Physical AI Tools and LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 50, "total_frames": 27949, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/DinsdagV2.tabularrobotics10K<n<100K0 likes19 downloads3mo agoHugging Face23syntaxhacker /rag_pipelinetextn<1K0 likes18 downloads1y agoHugging Face24SyntaxError-v2 /Move_Cube_v1.3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 25, "total_frames": 13927, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:25" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SyntaxError-v2/Move_Cube_v1.3.tabularrobotics10K<n<100K0 likes18 downloads8mo agoHugging Face25Syntax-Terror-BV /100episodesDonderdag1This dataset was created using Physical AI Tools and LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 99, "total_frames": 46759, "total_tasks": 1, "total_videos": 198, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:99" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/100episodesDonderdag1.tabularrobotics10K<n<100K0 likes18 downloads3mo agoHugging Face26hyungjikim /syntaxgym-hexataggedtext1K<n<10K0 likes17 downloads6mo agoHugging Face27Syntax-Terror-BV /omy_f3m_WoensdagTestMulti1540This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 31, "total_frames": 7357, "total_tasks": 3, "total_videos": 93, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:31" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_WoensdagTestMulti1540.tabularrobotics1K<n<10K0 likes17 downloads3mo agoHugging Face28Syntax-Terror-BV /omy_f3m_PickNPlaceDisndag1411This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 1, "total_frames": 1500, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/omy_f3m_PickNPlaceDisndag1411.tabularrobotics1K<n<10K0 likes17 downloads3mo agoHugging Face29Syntax-Terror-BV /WoensdagV1This dataset was created using Physical AI Tools and LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 49, "total_frames": 30967, "total_tasks": 1, "total_videos": 98, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:49" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/WoensdagV1.tabularrobotics10K<n<100K0 likes17 downloads3mo agoHugging Face30Syntax-Terror-BV /testThis dataset was created using Physical AI Tools and LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy_f3m", "total_episodes": 5, "total_frames": 3719, "total_tasks": 1, "total_videos": 15, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Syntax-Terror-BV/test.tabularrobotics1K<n<10K0 likes16 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.