CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cornerf /rebot-can-sort-stage1-v1-smoke ReBot can sorting Stage 1 Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in the taped sorting zone. Episodes: 52 Frames: 36729 FPS: 30 Robot: seeed_b601_dm_follower Cameras: observation.images.front (Logitech overhead) and observation.images.side (Innomaker wrist/claw) Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-can-sort-stage1-v1-smoke.tabularroboticsn<1K0 likes260 downloads2mo agoHugging Face02Cornerf /rebot-two-can-recycle-v2-smoke ReBot can sorting Stage 1 Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in the taped sorting zone. Episodes: 25 Frames: 14478 FPS: 30 Robot: seeed_b601_dm_follower Cameras: observation.images.front (Logitech overhead) and observation.images.side (Innomaker wrist/claw) Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-two-can-recycle-v2-smoke.tabularroboticsn<1K0 likes181 downloads2mo agoHugging Face03dougalldeepmind /2026-09-10-delib-synth-smoke Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) field value experiment Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) date_generated 20260910_191448 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth-smoke.tabularn<1K0 likes102 downloads13d agoHugging Face04dougalldeepmind /2026-09-11-delib-sonnet-synth-smoke Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) field value experiment Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) date_generated 20260911_192737 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-delib-sonnet-synth-smoke.tabularn<1K0 likes80 downloads12d agoHugging Face05dougalldeepmind /2026-08-04-ddp-smoke-bundleDDP smoke bundle: train_lora.py + a 64-example toy set, for validating multi-GPU wiring. textn<1K0 likes77 downloads2mo agoHugging Face06dougalldeepmind /2026-09-17-da-lowstakes-constitution-synth-smoke 18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only field value experiment 18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only date_generated 20260917_145731 constitution constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-constitution-synth-smoke.textn<1K0 likes71 downloads6d agoHugging Face07dougalldeepmind /2026-09-17-da-lowstakes-values-in-advice-synth-smoke 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only field value experiment 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only date_generated 20260917_171453 constitution constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-values-in-advice-synth-smoke.textn<1K0 likes70 downloads6d agoHugging Face08costadev00 /smoke-openai-terra-batch-brasil-25-20260724-01 Smoke OpenAI Terra Batch — Brasil × 25 tasks Run real de validação do fluxo matricial document_task_matrix, executada sobre um único documento da Wikipédia em português com o título Brasil. Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial. Resultado status: completed documentos: 1 pares planejados: 25 exemplos aceitos: 25 pares pulados: 0 pares esgotados: 0 resultados reais do backend: 27 retries com nova chamada: 2 backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.texttext-generationn<1K0 likes66 downloads2mo agoHugging Face09dougalldeepmind /2026-09-17-delib-noref-synth-smoke Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) field value experiment Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) date_generated 20260917_191247 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-delib-noref-synth-smoke.tabularn<1K0 likes66 downloads6d agoHugging Face10dougalldeepmind /2026-09-09-delib-synth-smoke Deliberative SFT: delib; native Qwen reasoning, no judge filtering field value experiment Deliberative SFT: delib; native Qwen reasoning, no judge filtering date_generated 20260909_152646 constitution constitutions/abridged_no_delib/constitution.md source_repo https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 5302fa9589d32c4296d9393b4838a00d91042b76 models qwen/qwen3.6-27b through Alibaba/OpenRouter (API revision not exposed)… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-delib-synth-smoke.textn<1K0 likes65 downloads14d agoHugging Face11dougalldeepmind /2026-09-08-delib-synth-smoke Deliberative SFT: delib; native Qwen reasoning, no judge filtering field value experiment Deliberative SFT: delib; native Qwen reasoning, no judge filtering date_generated 20260908_182742 constitution constitutions/no_claude_mentioned/constitution.md source_repo https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ a25750eb516cb6254640ae39a3a870f49398dbf0 models qwen/qwen3.6-27b through Alibaba/OpenRouter (API revision not exposed)… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-08-delib-synth-smoke.textn<1K0 likes55 downloads15d agoHugging Face12VmaxRL /cooper-campaign-etherpad-smoke Cooper Campaign Etherpad smoke substrate This public dataset contains the clean, non-privileged target contract for a one-task Campaign interoperability smoke. It identifies the exact public Etherpad image used by Cooper's August 2026 blue-team pilot and records the service boundary Cooper already exercised. The image itself remains in its upstream registry; this repository does not redistribute container bytes or source code. Campaign should use the exact digest-pinned… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/cooper-campaign-etherpad-smoke.textn<1K0 likes43 downloads25d agoHugging Face13Smoked-Salmon-s /empathetic_dialogues_ko Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)" Dataset Summary boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다. 일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다. GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다. 답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다. Generation Prompt Example(GPT3.5-turbo) Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/Smoked-Salmon-s/empathetic_dialogues_ko.texttext-generation10K<n<100K8 likes37 downloads3y agoHugging Face14Ismokebitcoins /hermes-smoketextn<1K0 likes35 downloads6d agoHugging Face15lfcarry /guide-evidence-smoke-tests Guide Evidence Smoke Tests LFCarry helps players find game guides and professional coaching and carry services. Explore LFCarry game guides and tier lists for the player-facing product. This repository shares a small testing resource for developers working on gaming assistants. It makes evidence-handling errors easier to reproduce. 24 original synthetic cases for checking whether a game-guide assistant supports a claim with the supplied evidence. Each case belongs to an invented… See the full description on the dataset page: https://huggingface.co/datasets/lfcarry/guide-evidence-smoke-tests.texttext-classificationn<1K0 likes35 downloads2d agoHugging Face16dougalldeepmind /2026-09-21-da-lowstakes-activity-grounded-synth-smoke 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only field value experiment 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only date_generated 20260921_074042 constitution constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-21-da-lowstakes-activity-grounded-synth-smoke.textn<1K0 likes34 downloads2d agoHugging Face17duyle2408 /levir-negative-canvas-r1-r4-smoke-runstabularn<1K0 likes32 downloads3d agoHugging Face18adraganov /arch-poison-mass-sft-lpi-260903T1130-w1-smoketestn<1K0 likes28 downloads20d agoHugging Face19sergiopaniego /ttt-scripted-smoke OpenEnv rollouts Collected with OpenEnv (openenv collect). Episodes: 20 Run metadata: key value env openspiel:tic_tac_toe env_base_url https://sergiopaniego-openspiel-ttt-env.hf.space provider scripted model None num_episodes_requested 20 temperature 0.2 keep_losses False Schema Each line of results.jsonl is one episode: episode_id (string) messages (chat transcript; TRL SFTTrainer-compatible) reward (float) done (bool) tool_trace (list of… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/ttt-scripted-smoke.textn<1K0 likes27 downloads5mo agoHugging Face20Letian2003 /stage1a_smoke_data stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh) Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP → frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend: WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT encoder — no raw-audio decoding at train time. 113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.tabularautomatic-speech-recognitionn<1K0 likes22 downloads2mo agoHugging Face21xaadii /smoke_testtabularn<1K0 likes22 downloads1mo agoHugging Face22Srishti280992 /repro-fluxnet-results-smoketextn<1K0 likes20 downloads2mo agoHugging Face23nisiwaki /parc2026-track1-texture-smoke-v1 Track 1 texture mask smoke result Two real selected episodes were decoded and segmented on A100. This repository stores the reproducibility evidence only: masks, fixed split reference, job definitions, summary, and execution log. It does not contain the original videos or constitute the final training dataset. tabularn<1K0 likes19 downloads1mo agoHugging Face24Jianshu001 /arabic-daily-batch01-smoke20-rewriter Batch 01 — Smoke 20 (Gemma-as-rewriter) 20-record smoke validating the Gemma-as-rewriter cleanup architecture. Approach Any prompt given to Gemma gets echoed into its own thinking field. Instead of fighting the echo, we isolate it by using Gemma as a rewriter: Cascade-regenerate assistant turn → (thinking1, answer) Send thinking1 to Gemma with a cleanup instruction → (thinking2, answer2) Discard thinking2 (absorbs the cleanup-instruction echo) Keep answer2 as the cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/arabic-daily-batch01-smoke20-rewriter.textn<1K0 likes15 downloads5mo agoHugging Face25CooperBench /team-coop-smoke CooperBench Team → Coop (Qwen3.5-9B smoke) Placeholder / smoke dataset (1 trajectory). A single 2-agent cooperbench team run (lead + member, no protocol) reshaped into the 2-agent coop layout defined in cooperbench/CooperData PR #98. The full dataset is the canonical place where future team→coop conversions will land; this entry validates the converter and the publishing pipeline. Source Source run logs/qwen35-smoke-mini-team-noproto/ Source repo… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/team-coop-smoke.texttext-generationn<1K0 likes15 downloads4mo agoHugging Face26fran-gen /snli-smoke-test SNLI Smoke Test Dataset Summary This dataset is a small smoke-test subset derived from the Stanford Natural Language Inference (SNLI) training split. It is intended for fast end-to-end checks of prompt formatting, model adapters, output parsing, and metric pipelines in entailment-lab. The dataset contains 100 sentence pairs: 34 entailment 33 contradiction 33 neutral Most of the dataset is organized as complete captionID triplets, where the same premise group… See the full description on the dataset page: https://huggingface.co/datasets/fran-gen/snli-smoke-test.texttext-classificationn<1K0 likes15 downloads2mo agoHugging Face27nielsr /arxiv-chandra-ocr-smoke-20260328-tokenfix arXiv OCR with Chandra OCR 2 This dataset stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix Source paper IDs in input list: 2 Processed IDs recorded in state/processed_ids.txt: 2 Successes: 2 Partial successes: 0 Errors: 0 Next shard index: 2 Updated at: 2026-03-28T15:22:09.273986+00:00 Files data/part-*.jsonl.gz: OCR result shards, one JSON object per paper… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix.tabularn<1K0 likes13 downloads6mo agoHugging Face28owenisas /opus46-reasoning-mix-smoke2 Opus 4.6 Reasoning Mix Generated via claude-opus-4.6 through an LMCanvas-backed proxy with reasoning: true. Sources: gsm8k, personahub_math, medical_o1_en, oasst1, openorca, tinystories Samples per source: 1 Fields include source text, generated question, answer, full captured reasoning, and raw response payload. tabularn<1K0 likes12 downloads7mo agoHugging Face29marcosremar2 /lora-smoke-datatextn<1K0 likes12 downloads4mo agoHugging Face30danieltslz /gene_embedding_smoke_candidatetextn<1K0 likes11 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.