datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.ParserV1-modelspi-trace-parser-sessionsItem-Parser-Dataset
Contents:
~$0.80 API token usage for Gemini 2.0 Flash Lite
agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.grok-parser-vrl-940k
grok-parser-vrl-940k
940,257 validated (log, grok_pattern) pairs for training models that
generate Vector.dev VRL parse_grok! patterns from raw log lines.
Files
merged_validated.csv — full schema (854 MB)
log — raw log line
parser — full VRL snippet (e.g. .message = ... | parse_grok!(.message, "..."))
grok_pattern — bare grok string extracted from parser
target — canonical pygrok output (dict)
parsed_output — independent re-application of grok_pattern (sanity check)… See the full description on the dataset page: https://huggingface.co/datasets/omeryentur/grok-parser-vrl-940k.agenda-parser-models-example-agent-traces
Agenda Parser — fine-tuned agent models
Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step
the model emits a single JSON action {"thought","tool","args"} over two toolkits —
meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA,
the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the
dataset itself (bottom) is a gallery of example traces from the three models.
tier
base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.Item-Parser-Dataset-v1.1-1k
Contents:
dataset with varying items parsed instead of parsing items from the same category. data generated by Qwen3 32B for first ~150 chats,
remaining 700+ listing and chat parsed outputs are generated by Qwen3 30B A3B, total API usage is ~US$0.70
atc-parser-canonical-v27atc-parser-eagle3-data-v1scry-coderag-parser-data-v2.3openapi_datasetItem-Parser-Dataset-v1.1-1k-thinkingscry-coderag-parser-data-v2Item-Parser-Dataset-v1.1-1k-non-thinkingItem-Parser-Dataset-iter2-1.5k
Content:
cost $2 with Gemini 2.5 Flash thinking, expensive af, the thinking response isnt helpful as well, only showing a summary of its CoT, next iter try with Qwen3 235B
Item-Parser-Dataset-v1-iter2
Contents:
dataset filtered to only contain <= 1k chars in the output col, previous iteration was causing colab to run out of ram due to some dataset rows having 2-10k char lengths in the output
Item-Parser-Dataset-v1.2-1k-thinkinglumi-parser-data
Lumi parser distillation data
Synthetic bilingual (Hebrew/English) task-parsing data for fine-tuning a small student model, generated by Qwen/Qwen3-235B-A22B-Instruct-2507 on Nebius.
Examples: 1465 (1319 train / 146 val)
Avg tasks/example: 2.36
Format: OpenAI chat messages (system/user/assistant) in train.jsonl/val.jsonl
Built for the Nebius Serverless AI Builders Challenge. License: MIT.
parsersqlbudget-parser-trainingscry-coderag-parser-data-v1scry-coderag-parser-data-v2.1
