datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiLang-Code-Parser-Dataset
MultiLang Code Parser Dataset (MLCPD)
MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format.
Each entry corresponds to one parsed source file and includes:
Language metadata
Code-level statistics (lines, errors, AST nodes)
Universal Schema JSON (normalized structural representation)
MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.ParserV1-modelsparser_dataset
parser_dataset
Parser training and evaluation data for multi-hop QA with long concatenated document contexts.
Derived from HotpotQA and 2WikiMultihopQA.
Contents
Path
Split
Samples
Notes
train/hotpotqa_train_process_emb-select_llm-unable.parquet
train
HotpotQA processed train
VERL / ParserRLHFDataset format
2wiki_val/eval_{N}.json
eval
128 per file
2WikiMultihopQA, N documents per example
hqa_val/eval_{N}.json
eval
128 per file
HotpotQA, N documents… See the full description on the dataset page: https://huggingface.co/datasets/inNexus/parser_dataset.pi-trace-parser-sessionsItem-Parser-Dataset
Contents:
~$0.80 API token usage for Gemini 2.0 Flash Lite
Recipe-parser
Dataset Card for Dataset Name
A collection of traditional Mountain Jewish (Gorsky Jewish) recipes from STMEGI.com, containing authentic culinary recipes representing the cultural heritage of the Caucasus Jewish community.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
A collection of 42 traditional Mountain Jewish (Gorsky Jewish) recipes collected from… See the full description on the dataset page: https://huggingface.co/datasets/AFKatz/Recipe-parser.omnimcp_healthtech_hl7_parser_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_hl7_parser_teaser.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.job-educational-parser-dataset-08-0-0805
Job Educational Parser Dataset
招聘领域的岗位与学历要求数据集。
输入:岗位描述 -> 输出:学历要求
Splits
train: 19w_0701.csv (约 19 万条)
test: 2w_0716.csv (约 2 万条)
validation: 4w_0708.csv (约 4 万条)
每条数据至少包含字段:
user: 职位描述
assistant: 要求的学历(如 "博士、硕士、本科"),遵循从高到低
由 @wangzihaogithub 创建。
grok-parser-vrl-940k
grok-parser-vrl-940k
940,257 validated (log, grok_pattern) pairs for training models that
generate Vector.dev VRL parse_grok! patterns from raw log lines.
Files
merged_validated.csv — full schema (854 MB)
log — raw log line
parser — full VRL snippet (e.g. .message = ... | parse_grok!(.message, "..."))
grok_pattern — bare grok string extracted from parser
target — canonical pygrok output (dict)
parsed_output — independent re-application of grok_pattern (sanity check)… See the full description on the dataset page: https://huggingface.co/datasets/omeryentur/grok-parser-vrl-940k.DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_224249DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_132118DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_000810DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_021553english-date-semantic-parser-data
Dataset Overview: semantic_train_en
Total Samples: 100000
Random Seed: 42
Noise Probability: 0.3
Generated At: 2026-02-10 14:40:49
Generator Distribution
Generator Function
Count
Percentage
Target Weight
gen_ambiguous_until
1812
1.81%
0.03
gen_before_after_weekday
2445
2.44%
0.04
gen_complex_weekday_offset
1866
1.87%
0.03
gen_compound
1242
1.24%
0.02
gen_day_after_tomorrow
2493
2.49%
0.04
gen_day_month_written
2531
2.53%
0.04
gen_day_of_month
1880… See the full description on the dataset page: https://huggingface.co/datasets/alperiox/english-date-semantic-parser-data.exp_tas_parser_xml_tracesdebug-math-parserparser_dataset_ner_v1.31arrow-parser-probeparser_user_v28ccitation-parser-ENTITYparser_user_v40amath500-rubric-parser-math-verifycitation-parser-SPANparser_dataset_ner_mini_v1.18atc-parser-canonical-v4agenda-parser-models-example-agent-traces
Agenda Parser — fine-tuned agent models
Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step
the model emits a single JSON action {"thought","tool","args"} over two toolkits —
meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA,
the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the
dataset itself (bottom) is a gallery of example traces from the three models.
tier
base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.test-fast-parser-l1b-v3parser_user_v29a
