CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Shofo /shofo-tiktok-general-small Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos. Size: ~50K videos (~500GB) Modality: Video + Audio + Text (transcripts, comments, captions) Source: TikTok Schema Column Type Description file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tabularvideo-classification10K<n<100K22 likes7.5k downloads7mo agoHugging Face02AmanPriyanshu /random-small-github-repositories random-small-github-repositories A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. Contents seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash) repos-zipped/ — one .zip per repo, named {repo_hash}.zip unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.tabulartext-generation1K<n<10K0 likes430 downloads6mo agoHugging Face03Baekpica /Inkling-Small-Multimodal-Calibration Inkling-Small Multimodal Calibration The exact 1,663 samples used for BF16 routed-expert importance collection for Inkling-Small Mixed Quant GGUF. This is calibration material, not a held-out evaluation benchmark. The primary balanced pass is: Category Samples Valid decoder tokens Share Text / reasoning 462 471,858 44.976% Code / tool-oriented source text 205 209,715 19.989% Real image / document 486 262,476 25.018% Real speech audio 309 105,080 10.016% Total… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Inkling-Small-Multimodal-Calibration.tabulartext-generation1K<n<10K0 likes413 downloads15d agoHugging Face04dm-petrov /youtube-commons-small 📺 YouTube-Commons-Small 📺 This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license. Dataset Description This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes. Features The dataset includes the following information for each video: Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.tabulartext-generation100K<n<1M1 likes265 downloads1y agoHugging Face05build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes182 downloads4mo agoHugging Face06wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face07emgena /omnimcp_smartenergy_iot_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_smartenergy_iot_teaser.tabulartext-generationn<1K0 likes81 downloads7d agoHugging Face08thebajajra /amazon-esci-english-smalltabulartoken-classification100K<n<1M1 likes80 downloads1y agoHugging Face09build-small-hackathon /dota2tuned-data DOTA2Tuned Data This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples. Contents sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors. Compact Parquet artifacts used by the Space: dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.tabulartext-generation1M<n<10M0 likes77 downloads3mo agoHugging Face10pixeloffice /llm-smartrouter-benchmark LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0) Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1). v1.4.0 Benchmark Highlights Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings. Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.tabulartext-generationn<1K0 likes68 downloads20d agoHugging Face11build-small-hackathon /hackathon-advisor-codex-traces Hackathon Advisor Codex Session Traces Real Codex session logs for the Hackathon Advisor project, selected from local Codex rollout JSONL files and redacted before publication. The event stream preserves user requests, assistant messages, tool calls, tool outputs, browser/search events, and minimal session provenance needed to audit how the project was built. Privacy filtering The publisher applied openai/privacy-filter at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.tabulartext-generationn<1K0 likes61 downloads4mo agoHugging Face12dhruveshpatel /star-smallSynthetic data for the paper [2505.05755] Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions. Project page: https://dhruveshp.com/projects/ilm tabulartext-generation10K<n<100K0 likes57 downloads1y agoHugging Face13build-small-hackathon /figment-eval-traces Figment Eval Traces Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders. These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment. Dataset Summary The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.tabulartext-generation100K<n<1M0 likes51 downloads3mo agoHugging Face14crawlfeeds /HomeDepot-Smart-Home-Dataset Home Depot Smart Home Product Dataset A structured dataset of Smart Home products from Home Depot, featuring detailed product specifications, pricing, category taxonomy, highlights, color variants, dimensions, and brand data. Ideal for training product recommendation models, smart home AI assistants, price intelligence systems, and e-commerce search engines. Dataset Overview Field Details Source Home Depot Total Records 230+ Category Focus Smart Home, IoT… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/HomeDepot-Smart-Home-Dataset.tabulartext-classificationn<1K0 likes42 downloads6mo agoHugging Face15ewdfd /SMART SMART: Evaluating LLMs’ Mathematical Reasoning via a Human Cognitive Process-Inspired Benchmark SMART is a fine-grained benchmark for evaluating large language models (LLMs) on mathematical reasoning from a human cognitive process perspective. Instead of evaluating only the final answer, SMART decomposes mathematical problem solving into four cognitive dimensions inspired by Pólya’s problem-solving theory: Semantic Understanding Mathematical Reasoning Arithmetic Computation… See the full description on the dataset page: https://huggingface.co/datasets/ewdfd/SMART.tabularquestion-answering1K<n<10K2 likes35 downloads5mo agoHugging Face16AkabekoLabs /nihongo-dojo-small nihongo-dojo-small このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。 データセット統計 train: 7,600 サンプル validation: 950 サンプル test: 950 サンプル 総サンプル数: 9,500 ソース 生成元: ./datasets/nihongo-dojo-small/ サンプルデータ { "instruction": "次のひらがなを漢字で書いてください。", "input": "「みず」を漢字で書くと?", "output": "<think>\n「みず」は「水」と書きます。意味: water\n</think>\n<answer>水</answer>", "group_id": 0, "task_idx": 0, "task_type": "kanji_writing", "difficulty": "beginner", "metadata":… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-small.tabulartext-generation1K<n<10K0 likes21 downloads1y agoHugging Face17ZachW /mistral-small-3.2-24b-instruct-2506_aime-all mistralai/Mistral-Small-3.2-24B-Instruct-2506 — aime-all Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: aime-all (933 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 32768 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_aime-all.tabulartext-generationn<1K0 likes20 downloads5mo agoHugging Face18ZachW /mistral-small-3.2-24b-instruct-2506_writingbench-en100 mistralai/Mistral-Small-3.2-24B-Instruct-2506 — writingbench-en100 Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: writingbench-en100 (100 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 8192 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_writingbench-en100.tabulartext-generationn<1K0 likes19 downloads5mo agoHugging Face19build-small-hackathon /job-search-distill Job Search Distillation Corpus A reasoning-trace SFT corpus for resume-aware job search. Teacher labels (search queries and fit evaluations, with full <think> reasoning preserved) generated by DeepSeek V4 Pro. Four relational configs cover the full pipeline: resumes → search queries → scraped jobs → fit evaluations. Dataset structure Config Contents resume_corpus resume_id, category, resume query_gen_pairings resume_id, teacher reasoning, list of… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/job-search-distill.tabulartext-generation10K<n<100K0 likes19 downloads4mo agoHugging Face20PhillyMac /SMART_Goals_Setting SMART Goals Setting This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/SMART_Goals_Setting.tabulartext-generationn<1K0 likes18 downloads4mo agoHugging Face21ZachW /mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384 mistralai/Mistral-Small-3.2-24B-Instruct-2506 — alpaca-text-generation-384 Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: alpaca-text-generation-384 (384 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384.tabulartext-generationn<1K0 likes17 downloads5mo agoHugging Face22ZachW /mistral-small-3.2-24b-instruct-2506_ifeval mistralai/Mistral-Small-3.2-24B-Instruct-2506 — ifeval Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: ifeval (541 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_ifeval.tabulartext-generationn<1K0 likes16 downloads5mo agoHugging Face23ZachW /mistral-small-3.2-24b-instruct-2506_storygen-prompts-200 mistralai/Mistral-Small-3.2-24B-Instruct-2506 — storygen-prompts-200 Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: storygen-prompts-200 (200 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_storygen-prompts-200.tabulartext-generationn<1K0 likes16 downloads5mo agoHugging Face24build-small-hackathon /PaperProf-traces PaperProf Agent Trace Step-by-step trace of PaperProf, an AI study buddy that turns course PDFs into interactive quiz sessions. What's in this dataset Each row in paperprof_trace.jsonl is one LLM call. Fields: Field Description session_id Groups steps from the same session step Step index within the session (1–4) type question_generation / answer_evaluation / mcq_generation topic Domain of the source chunk input Exact input sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/PaperProf-traces.tabularquestion-answeringn<1K0 likes13 downloads3mo agoHugging Face25build-small-hackathon /fabella-traces Fabella Anonymized Agent Traces A public, anonymized log of the LangGraph ReAct loop inside Fabella, a small-model Gradio Space for parents who need help explaining hard things to their child in kid-appropriate language. The dataset exists for the Sharing is Caring merit badge in the Build Small Hackathon. The first version of every explanation is drafted by google/gemma-4-E4B-it via a LangGraph ReAct loop with one tool (validate_explanation). A second small model —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/fabella-traces.tabulartext-generationn<1K0 likes13 downloads3mo agoHugging Face26build-small-hackathon /hackathon-advisor-quest-dataset Hackathon Advisor — Quest Classification SFT Dataset Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small Hackathon project against 13 judging dimensions from a two-segment README + app-file prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at build-small-hackathon/hackathon-advisor-quest-minicpm5-lora. Files quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.tabulartext-generationn<1K0 likes12 downloads4mo agoHugging Face27build-small-hackathon /velvet-rope-playtest-transcripts Velvet Rope Playtest Transcripts Cleaned playtest transcripts for Velvet Rope, a Build Small Hackathon Gradio game where players talk past whimsical AI gatekeepers by reading moods and discovering each character's soft spot. This dataset is published for the hackathon's sharing-is-caring badge. It contains 341 turn-level rows from 96 local playtest session files. Files data/playtest_transcripts.csv - table-friendly version. data/playtest_transcripts.jsonl - one… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/velvet-rope-playtest-transcripts.tabulartext-generationn<1K0 likes12 downloads3mo agoHugging Face28ZachW /mistral-small-3.2-24b-instruct-2506_creativemath-with-answers mistralai/Mistral-Small-3.2-24B-Instruct-2506 — creativemath-with-answers Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: creativemath-with-answers (188 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 32768 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_creativemath-with-answers.tabulartext-generationn<1K0 likes11 downloads5mo agoHugging Face29thibble /paper2env-small Paper2Env — Small (Author Inspection Set) A 10+10+50 subset of thibble/paper2env intended for fast author / reviewer inspection. Schemas and conventions are identical to the parent dataset. config rows description paperbench 10 Curated tasks (7 papers) — preferentially drawn from the tasks shown in the paper appendix. scraped 10 Auto-scraped tasks (9 papers) — drawn from the tasks shown in the paper appendix. trajectories 50 Multi-turn rollouts from… See the full description on the dataset page: https://huggingface.co/datasets/thibble/paper2env-small.tabulartext-generationn<1K0 likes11 downloads5mo agoHugging Face30build-small-hackathon /Kintsugi-Garden-traces Kintsugi Garden Evaluation Traces Paired evaluation traces from Kintsugi Garden — a local-first Jungian dream journal that runs Qwen3-8B through llama.cpp on a ZeroGPU Space. Every entry the app produces is shaped by both a fine-tuned model and a four-layer voice/safety architecture; this dataset is what those layers look like under instrumentation. What's in here 114 deterministic runs over the same 19 prompts × 3 trials, evenly split between: baseline —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/Kintsugi-Garden-traces.tabulartext-generationn<1K0 likes11 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.