datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FrontierOR
Frontier-OR Benchmark
A benchmark of 180 literature-grounded OR tasks, each packaged as a
self-contained reproducible unit: natural-language problem description,
mathematical formulation, reference Gurobi implementation, test instances,
reference solutions, and an automated feasibility checker.
Designed for evaluating LLMs on the end-to-end task of turning a research
paper's OR problem into runnable, verifiably-correct optimization code.
Dataset size note
This… See the full description on the dataset page: https://huggingface.co/datasets/SmartOR/FrontierOR.lm-eval-results-bunnycore-SmartToxic-7B-private
Dataset Card for Evaluation run of bunnycore/SmartToxic-7B
Dataset automatically created during the evaluation run of model bunnycore/SmartToxic-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bunnycore-SmartToxic-7B-private.byt5-small-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.Devstral-Small-2505-eval-logs-and-scoresovos-wake-word-bench-picovoice-smart-mirror
OVOS wake_word bench — picovoice-smart-mirror
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.smash-karts-multiplayer-trajectory
Smash Karts 멀티플레이 구현해줘
A single Codex coding-agent session implementing a multiplayer browser-based 3D
kart battle game inspired by Smash Karts.
The request covers multiplayer play, game rules, weapons and effects, keyboard
controls, research, and implementation. The trajectory records the development
process, tool calls and results, validation work, and the final response.
Field
Value
Session title
Smash Karts 멀티플레이 구현해줘
Session ID… See the full description on the dataset page: https://huggingface.co/datasets/amsminn/smash-karts-multiplayer-trajectory.esci-us-small
ESCI Shopping Queries Dataset (US Locale - Small Version)
This is a curated subset of the Amazon Shopping Queries Dataset (ESCI), filtered for the US locale only and using the small version of the dataset.
Dataset Description
The Shopping Queries Dataset is a large-scale manually annotated dataset for improving product search, released by Amazon Science. It contains challenging search queries paired with products and human-labeled relevance judgments.
Original… See the full description on the dataset page: https://huggingface.co/datasets/shuttie/esci-us-small.for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.abacusai__Llama-3-Smaug-8B-details
Dataset Card for Evaluation run of abacusai/Llama-3-Smaug-8B
Dataset automatically created during the evaluation run of model abacusai/Llama-3-Smaug-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Llama-3-Smaug-8B-details.network_securitykirana-detective-build-traces
Kirana Detective — Claude Code Build Sessions
Raw Claude Code (claude-sonnet-4-6) session traces recorded while building
Kirana Detective AI for the HuggingFace Build Small Hackathon 2026.
Each .jsonl file is one coding session. Together they cover the entire
build — from first commit to final submission.
What's Inside
Sessions
Agent
Coverage
11 JSONL files
Claude Code (Sonnet 4.6)
Full project build
Sessions include
Designing the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-detective-build-traces.mistralai__Mistral-Small-24B-Base-2501-details
Dataset Card for Evaluation run of mistralai/Mistral-Small-24B-Base-2501
Dataset automatically created during the evaluation run of model mistralai/Mistral-Small-24B-Base-2501
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Small-24B-Base-2501-details.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.pit-wall-chaos-tracesCodex agent traces for Pit Wall Chaos, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/pit-wall-chaos
shaer-eval-raw-gpt2-small-arabic-poetry
Shaer Evaluation Results
Models: gpt2_small_arabic_poetry
Source dataset: Shaer-AI/shaer-sft-test-generations-k5
Rows: 3481
Validation passed: True
Scored rows included: True
Dataset repo: Shaer-AI/shaer-eval-raw-gpt2-small-arabic-poetry
Files
generations.jsonl: raw generation rows
generations.csv: raw generation rows in CSV
generations_scored.jsonl: raw rows plus meter/count evaluation
validation.json: validation summary
generations_scored.csv: scored rows in… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-gpt2-small-arabic-poetry.TinyNarrator-agent-tracesSpaces link: https://huggingface.co/spaces/build-small-hackathon/TinyNarrator
gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-details
Dataset Card for Evaluation run of gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b
Dataset automatically created during the evaluation run of model gbueno86/Meta-LLama-3-Cat-Smaug-LLama-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gbueno86__Meta-LLama-3-Cat-Smaug-LLama-70b-details.figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.job-search-assistant-agent-tracebuild-small-agent-trace
Codex agent trace
This is submitted as part of the HF Build Small hack.
Codex was fully used for building the project.
It is necessary to share the trace to be eligible for the OpenAI price.
Trace was screened and sensitive keys redacted before sharing.
The full running app can be seen at :
https://huggingface.co/spaces/build-small-hackathon/storyboard-tui
PersonaChat-Qwen-original-Mistral-Small-4-119B-2603
Visual Memory Results: personachat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "mistralai/Mistral-Small-4-119B-2603",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-original-Mistral-Small-4-119B-2603",
"results_jsonl": "results/PersonaChat-Qwen-original-Mistral-Small-4-119B-2603.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-original-Mistral-Small-4-119B-2603.MatchWise-agent-tracePersonaChat-Qwen-enhanced-Mistral-Small-4-119B-2603
Visual Memory Results: personachat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "mistralai/Mistral-Small-4-119B-2603",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-enhanced-Mistral-Small-4-119B-2603",
"results_jsonl": "results/PersonaChat-Qwen-enhanced-Mistral-Small-4-119B-2603.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-enhanced-Mistral-Small-4-119B-2603.Xiaojian9992024__Qwen2.5-THREADRIPPER-Small-AnniversaryEdition-details
Dataset Card for Evaluation run of Xiaojian9992024/Qwen2.5-THREADRIPPER-Small-AnniversaryEdition
Dataset automatically created during the evaluation run of model Xiaojian9992024/Qwen2.5-THREADRIPPER-Small-AnniversaryEdition
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Xiaojian9992024__Qwen2.5-THREADRIPPER-Small-AnniversaryEdition-details.NeuroBait-Codex-Traces
Codex Session Traces
This folder contains Codex rollout JSONL traces related to the
NeuroBait Build Small Model project.
Included traces:
rollout-2026-06-08T17-03-23-019ea6af-db29-7801-ac01-46dfc88f90b0.jsonl
rollout-2026-06-09T07-10-21-019ea9b7-4610-7223-906e-2d0dba8bae7f.jsonl
rollout-2026-06-09T16-00-29-019eab9c-a18a-7de1-8967-ea63db425a4f.jsonl
The traces were selected because their session metadata contains the project
working directory:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/NeuroBait-Codex-Traces.in-your-own-worlds-tracesCodex agent traces for In your own wor(l)ds, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/in-your-own-worlds
PersonaChat-FLUX-original-Mistral-Small-4-119B-2603
Visual Memory Results: personachat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "mistralai/Mistral-Small-4-119B-2603",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-original-Mistral-Small-4-119B-2603",
"results_jsonl": "results/PersonaChat-FLUX-original-Mistral-Small-4-119B-2603.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-original-Mistral-Small-4-119B-2603.kicky-ai-codex-trace
Kicky AI - Codex agent trace (sanitized)
A redacted OpenAI Codex CLI session trace from building
Kicky AI for the Build Small
Hackathon - shared for the Sharing is Caring badge so others can see how the build went.
Format: Codex CLI JSONL session log (each record = {payload, timestamp, type}).
All secrets removed (HF / Modal / Roboflow tokens, shared secrets, emails) - verified 0 leaks.
Blog write-up: https://dcrey7.substack.com/p/world-fut-coach
mistralai__Mistral-Small-Instruct-2409-details
Dataset Card for Evaluation run of mistralai/Mistral-Small-Instruct-2409
Dataset automatically created during the evaluation run of model mistralai/Mistral-Small-Instruct-2409
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Small-Instruct-2409-details.
