CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01syntaxsynth /tmmluplus TMMLU+ : Large scale traditional chinese massive multitask language understanding We present TMMLU+, a traditional Chinese massive multitask language understanding dataset. TMMLU+ is a multiple-choice question-answering dataset featuring 66 subjects, ranging from elementary to professional level. The TMMLU+ dataset is six times larger and contains more balanced subjects compared to its predecessor, TMMLU. We have included benchmark results in TMMLU+ from closed-source models and 20… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/tmmluplus.text10K<n<100K0 likes1.7k downloads1y agoHugging Face02aisingapore /Linguistic-Diagnostics-Syntaxgated LINDSEA Syntax LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian. Supported Tasks and Leaderboards LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs). Languages Indonesian (id) Dataset Details LINDSEA Syntax only has an Indonesian (id) split, with additional splits containing fewshot examples. Below… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax.texttext-generationn<1K0 likes1.4k downloads9mo agoHugging Face03aisingapore /Linguistic-Diagnostics-Syntax-Judgegated LINDSEA Syntax LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian. Supported Tasks and Leaderboards LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs). Languages Indonesian (id) Dataset Details Data Sources Data Source License Language/s Split/s CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.textn<1K0 likes894 downloads2mo agoHugging Face04NuBerea /macula-hebrew-syntaxgated NuBerea MACULA Hebrew Syntax Trees (OT) Full syntactic tree annotation of the Hebrew Bible from the MACULA Hebrew Linguistic Dataset, packaged as relational tables for computational biblical studies. The dataset covers word-level linguistic annotation (morphology, glosses, lexical semantics), sentence segmentation, and hierarchical syntactic structure (clauses and phrases with their roles and containment relations) over the Westminster Leningrad Codex base text. This repository… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-hebrew-syntax.tabulartoken-classification1M<n<10M1 likes553 downloads10d agoHugging Face05NuBerea /macula-sblgnt-syntaxgated NuBerea MACULA Greek (SBLGNT) Syntax Trees (NT) Full syntactic tree annotation of the Greek New Testament from the MACULA Greek SBLGNT edition. Relational tables cover word-level tokens with morphological, semantic, and cross-language features; sentence boundaries; word groups (clauses and phrases) with syntactic rules and roles; the word-group hierarchy; and word-group membership. Together they let researchers traverse the full syntax tree of every sentence in the New Testament… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/macula-sblgnt-syntax.tabulartoken-classification1M<n<10M0 likes546 downloads10d agoHugging Face06cpllab /syntaxgymtext1K<n<10K2 likes424 downloads4y agoHugging Face07Syntaxdevloperangraeactionrpkaralho /action-roleplay-data Action Roleplay Data Data package for the Action SA-MP Android client. The client connects to 92.119.165.177:5636. The files/ directory contains the extracted game data, cache.zip is the archive consumed by the initial installer, files.json is the file-by-file manifest, and client_config.json contains the public endpoints. Runtime logs were excluded from the distributable package. The APK included here is a debug build for testing and is signed with a debug key. geospatialn<1K0 likes305 downloads19d agoHugging Face08syntaxsynth /swe-bench-opus-logs Claude 3 inference SWE-Bench results Contains prompting responses from SWE-bench on these 2 settings: Oracle retrieval BM25 retrieval Each of the subsets contains an additional log_last_line attributes which is the last line from log files generated during evaluation step. Results: Model BM25 Retrieval Resolved (%) Oracle Retrieval Resolved (%) GPT-4* 0 1.74 Claude-2 1.96 4.80 Claude-3 Opus (20240229) 3.24 6.42 Claude-2 and GPT-4 results from SWE-bench… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/swe-bench-opus-logs.textquestion-answering1K<n<10K1 likes222 downloads3y agoHugging Face09cpllab /syntaxgym_sentencestabular1K<n<10K1 likes107 downloads4y agoHugging Face10Noushad999 /ML-1M-Syntax-Validated-Python-Code ML-1M Syntax-Validated Python Code Dataset Summary ML-1M Syntax-Validated Python Code is a large-scale corpus containing over 1 million machine-learning–oriented Python programs derived from The Stack, a permissively licensed collection of open-source source code. The dataset is constructed through heuristic ML-domain filtering, syntactic validation, and basic safety checks. It is intended to support empirical analysis of real-world ML code, executability and dependency… See the full description on the dataset page: https://huggingface.co/datasets/Noushad999/ML-1M-Syntax-Validated-Python-Code.texttext-generation1M<n<10M0 likes106 downloads8mo agoHugging Face11Corpus-NZ /Code-Syntax-Expanded Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows. 📊 Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.text10M<n<100M0 likes60 downloads28d agoHugging Face12syntaxsynth /Ultra-FineWeb-L3-zh-hant-translated Ultra-FineWeb-L3 (Traditional Chinese Translation) Traditional Chinese translation of the English openbmb/Ultra-FineWeb-L3 corpus, targeting Taiwan-standard Traditional Chinese (臺灣正體中文). Background Ultra-FineWeb-L3 is the L3-refined tier of the UltraData data management framework developed by OpenBMB. Starting from Ultra-FineWeb (the high-quality web corpus behind MiniCPM4), L3 refinement transforms raw web documents into two structured synthesis formats: Q&A… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/Ultra-FineWeb-L3-zh-hant-translated.text1M<n<10M0 likes57 downloads3mo agoHugging Face13syntaxlabs /medical-billing-icd10-qatextn<1K0 likes51 downloads14d agoHugging Face14syntaxsynth /mmevol-zh-hant MMEvol - Translated Chinese Traditional A subset of Tongyi-ConvAI/MMEvol translated using yentinglin/Llama-3-Taiwan-70B-Instruct from english to traditional chinese. Read the Note below before use. Image source distribution: Dataset Count Percentage coco 6598 29.8% Q-Instruct-DB 5856 26.4% clevr 2383 10.8% chartqa 1733 7.8% hfdata 1296 5.9% geo170k 706 3.2% data_engine 6983.2% mathvision 644 2.9% docvqa 600 2.7% alfworld 401 1.8% arxivqa 337 1.5%… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/mmevol-zh-hant.imagetext-generation10K<n<100K1 likes47 downloads2y agoHugging Face15ratimics /syntax-gossip Syntax Gossip (candidate) Syntax IR records over Crownless event-to-speech gossip: exact source bytes, UTF-8 byte spans, UDPipe UD dependencies, NP/VP chunks, clause spans with hearsay attribution links, and audit lineage per record. Status: candidate (braid-syntax-gossip-v0.1.0-candidate.1). Not frozen, authorizes no training. Cloud verification and gold review are outstanding. Splits: train.jsonl (4,000) / validation.jsonl (500) / test.jsonl (500). Backend:… See the full description on the dataset page: https://huggingface.co/datasets/ratimics/syntax-gossip.texttoken-classification1K<n<10K0 likes47 downloads6d agoHugging Face16syntaxnoob /weather-prediction-prototype-aws Weather prediction prototype database. This database was made using data provided by KMI. This database will only be used to train a prototype. Dataset Details Dataset Description Dataset Sources [optional] KMI Dataset Structure Normalized columns: timestamp air_pressure relative_humidity precipitation wind_speed wind_direction More information about these columns can be found in the information_10min.txt file. tabular100K<n<1M1 likes46 downloads3y agoHugging Face17syntaxsynth /reasoning-conversations Multilingual Reasoning Dataset Include languages from German, Korean, Spanish, Japanese, French, Simplified Chinese, Traditional Chinese Reasoning traces from Deepseek-v3-R1, Deepseek-v3-R1-Zero Credits sponsored by Currents API text10K<n<100K4 likes40 downloads2y agoHugging Face18KRadim /czech-punctuation-pos-syntax Czech Punctuation, POS and Syntactic Dataset 🇨🇿 A High-Quality Dataset for Punctuation Restoration and Neuro-Symbolic LLM Grounding This dataset is a structured, linguistically annotated corpus of the Czech language, specifically designed for Punctuation Restoration tasks, Part-of-Speech (POS) tagging, and token-level syntax embedding (such as nanoGPT custom metadata training). Unlike pure raw text corpora, this dataset provides a deterministic 1:1 token-level mapping… See the full description on the dataset page: https://huggingface.co/datasets/KRadim/czech-punctuation-pos-syntax.texttoken-classification100K<n<1M0 likes39 downloads4mo agoHugging Face19Builder-syntaxlabs /medical-billing-icd10-qa-sample Medical Billing & ICD-10 Synthetic Dataset (Sample 🚀 NEED THE FULL ENTERPRISE COMMERCIAL DATASET? Get instant access to the full 50,000+ cleaned JSONL dataset for fine-tuning production models: 🏥 50,000+ verified ICD-10 / CPT billing scenarios 📄 Clean JSONL format (instruction, input, output) 🔒 Safe for HIPAA/GDPR—100% synthetic, zero real patient data 💼 Full commercial license for SaaS and Enterprise applications 👉 Buy Full Enterprise Dataset ($249) - Instant Download… See the full description on the dataset page: https://huggingface.co/datasets/Builder-syntaxlabs/medical-billing-icd10-qa-sample.texttext-generationn<1K0 likes38 downloads2d agoHugging Face20hriaz /syntaxgym-hexatagged Dataset Card for "syntaxgym-hexatagged" More Information needed text1K<n<10K0 likes28 downloads5mo agoHugging Face21Gugu8 /Code-Syntax-Expanded Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows. 📊 Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) License… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax-Expanded.text10M<n<100M0 likes27 downloads2mo agoHugging Face22syntaxsynth /reprompts-20k-sample Reprompts of conversations using Opus-20240229 20k samples Sauce: lmsys/lmsys-chat-1m - en only allenai/WildChat-1M - en only teknium/OpenHermes-2.5 teknium/OpenHermes-2.5 ShareGPT text10K<n<100K0 likes23 downloads2y agoHugging Face23Gugu8 /Code-Syntax Code Syntax Dataset (S) A large-scale, high‑quality dataset for teaching large language models to identify and correct common syntax errors across 30+ programming languages.Contains 500,000+ unique examples (≈110 MB) with English explanations – no artificial padding. 📊 Dataset Format The dataset is provided as a single CSV file with the following columns: Column Type Description wrong_code string Code snippet containing a syntax error correct_code… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax.text1M<n<10M0 likes20 downloads2mo agoHugging Face24Kamyar-zeinalipour /Essay-Syntax-Instructtext1K<n<10K0 likes19 downloads2y agoHugging Face25hyungjikim /syntaxgym-hexataggedtext1K<n<10K0 likes19 downloads6mo agoHugging Face26syntaxsynth /instruct_code_cleaning SFT code dataset building Contain a list of tasks useful when building a iniitial dataset source: reverse_translation Given a history of conversations, what would the human ask next? reverse_translation_first_round Suppose you already have a response, the LLM must predict what question does the human asked clean_code Given a code snippet, it determines whether its useful and atomic enough to be use for a response by LLM gen_code_question Generates a question given a… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/instruct_code_cleaning.texttext-generation10K<n<100K1 likes18 downloads2y agoHugging Face27syntaxhacker /rag_pipelinetextn<1K0 likes18 downloads1y agoHugging Face28syntaxhacker /developer-portfolio-ragtextn<1K1 likes14 downloads1y agoHugging Face29CoBaLD /enhanced-ud-syntax Enhanced Universal Dependencies (syntax only) dataset This repo aggregates syntactic markup from UD_English-EWT and UD_English-GUM datasets. Changes made to the source datasets: Only syntactic tags (head, deprel, deps) are preserved Short sentences (fewer than 3 tokens) are removed Source datasets: UD_English-EWT: https://github.com/UniversalDependencies/UD_English-EWT UD_English-GUM: https://github.com/UniversalDependencies/UD_English-GUM texttoken-classification10K<n<100K0 likes13 downloads1y agoHugging Face30manupinasco /syntax_analysistext1K<n<10K2 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.