datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.hackathon-advisor-quest-dataset
Hackathon Advisor — Quest Classification SFT Dataset
Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small
Hackathon project against 13 judging dimensions from a two-segment README + app-file
prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at
build-small-hackathon/hackathon-advisor-quest-minicpm5-lora.
Files
quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.azure-advisor-grpo-benchmark
Azure Advisor GRPO Benchmark Dataset
Evaluation benchmark for measuring the quality of Azure Advisor recommendation generation, used for GRPO (Group Relative Policy Optimization) training and model evaluation.
Dataset Description
This dataset contains 106 evaluation examples with ground truth labels, designed to score model outputs across 5 reward dimensions.
Purpose
During GRPO training: Score generated recommendations to select high-reward samples
Model… See the full description on the dataset page: https://huggingface.co/datasets/thegovind/azure-advisor-grpo-benchmark.
