datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiLang-Code-Parser-Dataset
MultiLang Code Parser Dataset (MLCPD)
MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format.
Each entry corresponds to one parsed source file and includes:
Language metadata
Code-level statistics (lines, errors, AST nodes)
Universal Schema JSON (normalized structural representation)
MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.omnimcp_healthtech_hl7_parser_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_hl7_parser_teaser.DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_224249DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_132118DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_000810DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_021553english-date-semantic-parser-data
Dataset Overview: semantic_train_en
Total Samples: 100000
Random Seed: 42
Noise Probability: 0.3
Generated At: 2026-02-10 14:40:49
Generator Distribution
Generator Function
Count
Percentage
Target Weight
gen_ambiguous_until
1812
1.81%
0.03
gen_before_after_weekday
2445
2.44%
0.04
gen_complex_weekday_offset
1866
1.87%
0.03
gen_compound
1242
1.24%
0.02
gen_day_after_tomorrow
2493
2.49%
0.04
gen_day_month_written
2531
2.53%
0.04
gen_day_of_month
1880… See the full description on the dataset page: https://huggingface.co/datasets/alperiox/english-date-semantic-parser-data.exp_tas_parser_xml_tracesdebug-math-parserparser_dataset_ner_v1.31parser_user_v28cparser_user_v40aparser_dataset_ner_v1.49citation-parser-SPANparser_dataset_sgpt_v3.8math500-rubric-parser-math-verifyparser_dataset_ner_mini_v1.18test-fast-parser-l1b-v3parser_dataset_ner_v1.16parser_user_v29asynth_cv_parser_fakerparser_user_v8_pbparser_user_v20cparser_user_v22fparser_user_v27aparser_dataset_sgpt_v3.4parser_user_v44aresume_parsercitation-parser-ENTITYparser_v4_mod4_trainingdata
