datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dota2tuned-data
DOTA2Tuned Data
This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples.
Contents
sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors.
Compact Parquet artifacts used by the Space:
dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.proofkit-distill-qwen0.5b
ProofKit distillation dataset
~7,000 chat examples for sequence-level (data) distillation. ProofKit's fine-tuned
gpt-oss-20b teacher (visproj/proofkit-gpt-oss-20b-lora)
regenerates the assistant turn over the exact prompts from
visproj/proofkit-sft; the
system + user turns are kept verbatim, so the set stays license-safe (no scraping, no
PII).
A Qwen 0.5B student is then SFT'd on this to produce
visproj/proofkit-distilled-qwen0.5b
(and its GGUF), which the ProofKit Space serves… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/proofkit-distill-qwen0.5b.ibero-characters-es
Conjunto de datos de personajes de mitos y leyendas iberoamericanos.
⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y mejorar la cobertura de imágenes.
📚 Descripción
Dataset de personajes míticos y legendarios de Iberoamérica, diseñado para preservar y promover el patrimonio cultural a través de la inteligencia artificial.
🌟 Motivación e Impacto
📱 Preservación Digital: Conservación del… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-characters-es.compliment-forest-sft
Compliment Forest SFT
Compliment Forest SFT teaches a small language model to turn a (name, situation)
pair into a strict JSON forest of grounded encouragement. Each forest contains five
distinct creature-strength clearings, a situation-specific line, an agency-oriented
reflection, a short first-person spell, and a creature-only image prompt.
Dataset Size
Train: 1,350 records
Validation: 150 records
Seed: 42
Language: English
Every row contains:
name
situation… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-sft.job-search-distill
Job Search Distillation Corpus
A reasoning-trace SFT corpus for resume-aware job search. Teacher labels (search queries and
fit evaluations, with full <think> reasoning preserved) generated by DeepSeek V4 Pro.
Four relational configs cover the full pipeline: resumes → search queries → scraped jobs →
fit evaluations.
Dataset structure
Config
Contents
resume_corpus
resume_id, category, resume
query_gen_pairings
resume_id, teacher reasoning, list of… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/job-search-distill.umpalumpas
Compaction Dataset Based on SWE-Smith
This is a sample of a larger dataset. Work in progress.
Parquet Layout
This repository contains a single Parquet file:
File
Rows
What it contains
data/oracle_trajectories.parquet
10
One row per oracle trajectory, preprocessed with linked offline checkpoints, oracle memories, and oracle continuations.
The compaction data is represented by linked offline H2 fixed-milestone
checkpoints:
offline_checkpoint_count… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/umpalumpas.
