CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01XReyRobert /smoke24-agentic-benchmarks Smoke24 Agentic Benchmarks Public, reproducible Terminal-Bench 2.0 Smoke24 benchmark artifacts for local RTX 3090-class agentic model serving. Why Smoke24 I created the Smoke24 subset because I needed a relatively quick benchmark that could run locally on an RTX 3090-class machine in a couple of hours. The goal is to get a practical read on model performance and stability under a real agentic Terminal-Bench workload, without paying the turnaround cost of a much… See the full description on the dataset page: https://huggingface.co/datasets/XReyRobert/smoke24-agentic-benchmarks.text-generation0 likes98 downloads25d agoHugging Face02costadev00 /smoke-openai-terra-batch-brasil-25-20260724-01 Smoke OpenAI Terra Batch — Brasil × 25 tasks Run real de validação do fluxo matricial document_task_matrix, executada sobre um único documento da Wikipédia em português com o título Brasil. Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial. Resultado status: completed documentos: 1 pares planejados: 25 exemplos aceitos: 25 pares pulados: 0 pares esgotados: 0 resultados reais do backend: 27 retries com nova chamada: 2 backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.texttext-generationn<1K0 likes66 downloads2mo agoHugging Face03victor /synthetic-customer-support-sft-smoke Synthetic customer-support SFT dataset Synthetically generated with HuggingFaceTB/SmolLM2-135M-Instruct from seeded scenario prompts (18 products x 16 issue types x 5 customer tones). Rows: 4 kept after validation (4 completions failed parsing and were dropped) Format: messages column (system/user/assistant), ready for TRL SFTTrainer Metadata: category (issue type), tone (customer tone) Seed: 42 Generated on 2026-09-02. Model-generated content: review before production use. texttext-generationn<1K0 likes39 downloads22d agoHugging Face04Smoked-Salmon-s /empathetic_dialogues_ko Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)" Dataset Summary boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다. 일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다. GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다. 답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다. Generation Prompt Example(GPT3.5-turbo) Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/Smoked-Salmon-s/empathetic_dialogues_ko.texttext-generation10K<n<100K8 likes34 downloads3y agoHugging Face05pere /nb-asr-numerics-categorized-smoke-test Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized-smoke-test.tabulartext-generationn<1K0 likes23 downloads3mo agoHugging Face06Solshine /gemma-4-e2b-nla-eval-smoke Gemma-4-E2B NLA smoke-eval (20-row held-out set) A 20-row held-out subset of OpenWebText activations extracted from google/gemma-4-E2B at layer 23. Used as the canonical eval set for smoke-testing the v0.0.1 Gemma-4-E2B NLA pair on a fresh environment. This dataset is a subset of the held-out rl.parquet evaluation set used for the v0.0.1 round-trip eval (n=50 attempted, 42 evaluated after 8 empty-output exclusions, cos 0.438 ± 0.054). The 20-row subset preserves the activation… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-eval-smoke.tabulartext-generationn<1K0 likes22 downloads4mo agoHugging Face07CooperBench /team-coop-smoke CooperBench Team → Coop (Qwen3.5-9B smoke) Placeholder / smoke dataset (1 trajectory). A single 2-agent cooperbench team run (lead + member, no protocol) reshaped into the 2-agent coop layout defined in cooperbench/CooperData PR #98. The full dataset is the canonical place where future team→coop conversions will land; this entry validates the converter and the publishing pipeline. Source Source run logs/qwen35-smoke-mini-team-noproto/ Source repo… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/team-coop-smoke.texttext-generationn<1K0 likes15 downloads4mo agoHugging Face0813point5 /reverse-text-tinystories-easy-smoke Reverse Text TinyStories Easy Smoke This is a small smoke-test dataset for the reverse-text task. Splits train: 12 rows test: 4 rows Columns prompt char_count word_count source Source Derived from roneneldan/TinyStories by taking non-overlapping word windows and keeping only prompts that fall in the easy character-length bucket. Difficulty Rule All rows in this dataset are easy examples with prompt lengths in the 20-74 character range.… See the full description on the dataset page: https://huggingface.co/datasets/13point5/reverse-text-tinystories-easy-smoke.tabulartext-generationn<1K0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.