CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01glayguo /noteflow-research-pilots Keep the failed attempts. Check the artifact. Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy. Configuration Actual experiment What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.imagetext-generationn<1K0 likes454 downloads4d agoHugging Face02asingh15 /tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace TCS Qwen3.6-27B Direction Beam Full-Trace Pilot Verified compact export for tcs_qwen36_27b_direction_beam_pilot50_full_logging_20260814. Source dataset: TCS train-00000-of-00001.parquet Problems: 50 Displayed trajectories: 200 Chunk probes: 1,600 Terminal answers and rubric grades: 6,400 Policy: Qwen/Qwen3.6-27B Judge: openai/gpt-oss-20b (low reasoning) At each displayed chunk, the policy proposes four directions plus a null continuation. A width-four stochastic beam reaches… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace.text-generation0 likes310 downloads1mo agoHugging Face03CyberNative-AI /qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3 Qwen3.6-27B GGUF quantization on a bounded BFCL V4 pilot Q4_K_M matched Q8_0 on both tested categories: each scored 94 of 100 selected cases correct. Q5_K_M also scored 94/100; Q3_K_M scored 92/100. Read the results page · Inspect all 400 scored rows This is a post-result-corrected exploratory analysis of two selected non-live BFCL V4 categories, not a full leaderboard result. Inspect the scored rows without cloning The Hub Dataset Viewer does not render this… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3.text-generationn<1K1 likes307 downloads8d agoHugging Face04VmaxRL /SWE-universe-repaired-bug-pilot-trajectories SWE-universe repaired BugPilot trajectories Combined trajectory artifacts for the Qwen3.6 + mini-swe-agent evaluation of VmaxRL/SWEUniverse-Repaired-Bugpilot. This dataset contains one row per evaluated task in metadata.jsonl, plus per-task files under trajectories//. The combined set uses the main full eval and replaces the two original infra-failure rows with the clean infra rerun trajectories. Summary: rows: 804 effective attempts: 804 passes: 629 pass rate: 0.782338 infra… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/SWE-universe-repaired-bug-pilot-trajectories.text-generation0 likes285 downloads4mo agoHugging Face05propertypilot /property-pilot-tickets 🏢 PropertyPilot — Maintenance Tickets A synthetic dataset of 13,725 residential-maintenance tickets written the way real tenants write them — polite, panicked, passive-aggressive, or confused — each paired with operational metadata (category, urgency, assigned contractor, cost, resolution time). Built for an end-to-end NLP pipeline: triage classification, similar-case retrieval (embeddings + FAISS), and work-order / reply generation. About this release. Earlier versions of… See the full description on the dataset page: https://huggingface.co/datasets/propertypilot/property-pilot-tickets.imagetext-classification10K<n<100K0 likes281 downloads1mo agoHugging Face06TIE-Pilot /qwen3-4b-perfectblend-deepspec-rollout Qwen3-4B PerfectBlend DeepSpec Rollout This dataset contains the complete DeepSpec-aligned Qwen3-4B self-distillation rollout over the filtered PerfectBlend corpus. The seeded 95/5 split is published as separate train and eval splits. Splits Split Conversations Shards Path train 1,349,860 128 data/*.jsonl eval 71,046 64 eval/*.jsonl total 1,420,906 192 Data construction Canonical filtered corpus: 1,420,906 conversations. Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.texttext-generation1M<n<10M0 likes126 downloads1d agoHugging Face07NoeFlandre /fineweb-legal-pilot ⚖️ FineWeb-Legal-Pilot 66.8M words of the finest legal domain data the 🌐 web has to offer. Repo: GitHub | Report: Technical Report What is it? FineWeb-Legal-Pilot is a pilot dataset consiting of 52k high-quality legal documents filtered from the 10-billion-token sample of 🍷 FineWeb. To enhance FineWeb's utility for legal AI domain adaptation, we draw inspiration from the FineWeb-Edu methodology: creating a legal quality classifier using annotations… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/fineweb-legal-pilot.tabulartext-classification100K<n<1M6 likes88 downloads9mo agoHugging Face08krisztiankoos /uldr-v0.1-pilot Ukrainian Language Decolonization & Reasoning (ULDR) — Pilot Canary Release v0.1 [!IMPORTANT] Exploratory Pilot / Canary Release (v0.1): This dataset represents an early exploratory pilot canary release (v0.1-pilot) establishing our baseline data pipeline, schema contracts, and directional validation. It is not the final production release (v1.0). The full production release is scheduled for Phase 5.5 following complete evaluation suite assembly, dialect & historical protection… See the full description on the dataset page: https://huggingface.co/datasets/krisztiankoos/uldr-v0.1-pilot.texttext-generation10K<n<100K0 likes73 downloads8d agoHugging Face09aznatkoiny /cannabis-fda-extractive-pilot FDA Cannabis Extractive Experimental Pilot Experimental, automatically screened, unreviewed draft dataset. This dataset is not medical advice, is not production-ready, and must not be represented as clinician-reviewed, legally cleared, or suitable for patient-facing systems. This small English conversational dataset was created to test an auditable Gemma 4 fine-tuning pipeline. It contains exact answer passages from captured FDA pages about CBD/cannabis safety, paired with… See the full description on the dataset page: https://huggingface.co/datasets/aznatkoiny/cannabis-fda-extractive-pilot.texttext-generationn<1K0 likes62 downloads2mo agoHugging Face10violetxi /toolathlon-qwen35-4b-base-pass1-pilot15 Toolathlon Qwen3.5-4B Pass@1 Pilot Readable trajectories from a 15-task Toolathlon-Verified Pass@1 pilot using Qwen/Qwen3.5-4B in thinking mode. The model was served on 6 B200 GPUs with data parallelism; Toolathlon's Docker environments ran on its remote public evaluation service. Contents Repository: violetxi/toolathlon-qwen35-4b-base-pass1-pilot15 Rows: 15 (one row per task) Reported result: 5/15 (33.33%) messages: a Viewer-friendly list containing only user… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/toolathlon-qwen35-4b-base-pass1-pilot15.tabulartext-generationn<1K0 likes57 downloads2mo agoHugging Face11leonli66 /stage3-real-expansion-agent-teacher-separated-pilot Teacher-Separated Expansion Agent Pilot A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks. The teacher-only trajectory-generation system prompt is recorded in metadata/generation-manifest.json for auditability, but is absent from every saved training trajectory. Each final messages list begins with the real memory-wrapped task user message, followed by native assistant expand calls, exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.tabularquestion-answeringn<1K0 likes57 downloads23d agoHugging Face12violetxi /ch-pilot-rollouts-qwen3.5-9b C&H Pilot Rollouts — Qwen/Qwen3.5-9B 20 agentic exploration rollouts over the Calderwood & Harkness (C&H) synthetic law-firm corpus (the open-sourced world from harvey-labs tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-9B served with vLLM. Part of an actor-selection pilot for a world-internalization research project: the goal is to mine agent trajectories into verified fact stores and rewritten likelihood-training targets. Companion dataset (same seeds/tasks, different… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-pilot-rollouts-qwen3.5-9b.tabulartext-generationn<1K0 likes46 downloads1mo agoHugging Face13argo11 /japanese-math-empirical-difficulty-pilot-50k Japanese Math Empirical Difficulty Pilot 50k This dataset is a 50,000-problem empirical difficulty pilot, not a full empirical labeling of the original 5.66M-row source dataset. It was created for LLM-jp experiment 0399, Team Victory SFT, to validate empirical difficulty label distribution, downstream split behavior, and the rollout/scoring pipeline before attempting labeling at the full 5.6M scale. Current Status This upload uses the v3 scorer with assistant-only… See the full description on the dataset page: https://huggingface.co/datasets/argo11/japanese-math-empirical-difficulty-pilot-50k.tabulartext-generation100K<n<1M0 likes42 downloads3mo agoHugging Face14birgermoell /oellm-longctx-tokenized-natural-128k-256k-pilot-v1 OELLM Natural Long-Context Tokenized 128K/256K Expanded Pilot This dataset contains Megatron-LM tokenized continuation-training artifacts for natural long-context extension experiments at 128K and 256K sequence scales. This public revision expands the original pilot from 128 to 512 packed examples per source/tier, for 3,072 packed examples total. The raw pack manifest reports approximately 586M source-side estimated tokens across all six source/tier shards. Accessible source… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-natural-128k-256k-pilot-v1.text-generation0 likes39 downloads3mo agoHugging Face15teex-pt /amalia-pilot-honesty-v2 AMALIA pilot — honesty vector datasets (v1 refusals + v2 corrective mix) Training data from the first two iterations of a verifier-gated fine-tuning pilot on AMALIA-9B-0626-DPO, targeting identity/fact confabulation (the model's weakest measured behavior: 43.3% on our honesty harness). Full methodology, harness, and reports: github.com/teex-pt/pt-amalia. These are research pilot artifacts — small by design (the pilot validates the loop, not the scale). Every sample was produced… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-pilot-honesty-v2.texttext-generation1K<n<10K0 likes31 downloads3mo agoHugging Face16bilalabic /toolcall-tr-pilot ToolCall-TR — Pilot v0.1 Türkçe, execution-verified tool-calling veri seti. Her tool çağrısı gerçek bir implementasyon tarafından çalıştırıldı ve çıktısı doğrulandı; hiçbir tool observation'ı bir LLM tarafından yazılmadı. ⚠️ Önce şunu bilin: bu küçük bir ilk sürüm İçinde 200 örnek var. Bu sayı bir modeli eğitmek için yeterli değildir. Ne için uygun: Yöntemi ve veri biçimini incelemek Örneklere tek tek bakmak Az sayıda örnekle deneme yapmak (few-shot) Kendi ölçüm… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/toolcall-tr-pilot.texttext-generationn<1K0 likes27 downloads2mo agoHugging Face17anote-ai /IVR-pilot-benchmark IntentSpec Benchmark — Data Supplement This archive contains the benchmark data used to compute Intent Violation Rate (IVR) in the paper: 49 tasks, each derived from a HumanEval problem and extended with an ambiguous/gold prompt pair and a decomposed set of executable constraints. Files spec_pairs.jsonl The benchmark itself — one JSON object per line, one line per task. This is the file consumed directly by the evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/IVR-pilot-benchmark.texttext-generationn<1K0 likes26 downloads1mo agoHugging Face18santiisoutoo /pilot-readback-corpus Pilot Readback Corpus A synthetic reference corpus of Air Traffic Control (ATC) pilot–controller exchanges in standard ICAO phraseology, intended as ground truth for evaluating the output of LLM-based pilot agents. ⚠️ Synthetic data. All exchanges are author-written following ICAO Doc 4444 / Annex 10 phraseology. They are not transcriptions of real ATC recordings and contain no real flight identifiers, clearances, or operational data. Overview The corpus covers… See the full description on the dataset page: https://huggingface.co/datasets/santiisoutoo/pilot-readback-corpus.texttext-generationn<1K2 likes20 downloads4mo agoHugging Face19spectralbranding /r16-behavioral-metamerism-pilot R16 Behavioral Metamerism Pilot Brand Function x synthetic cohort interaction experiment from the Spectral Brand Theory research program. Dataset Summary 675 API calls testing whether Brand Function specification differentially affects dimensional collapse across synthetic observer cohorts. Design: 5 cohorts x 5 brands x 3 conditions (no BF, structural BF, enriched BF) x 3 models x 3 repetitions. Companion paper: AI-Native Brand Identity: From Visual Recognition… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r16-behavioral-metamerism-pilot.tabulartext-generation1K<n<10K0 likes18 downloads2mo agoHugging Face20AgenticCommons /oasst1-zh-pilot oasst1 Chinese Translation Pilot (10 samples) This is a pilot release of 10 parallel English→Chinese samples translated from OpenAssistant/oasst1. It is intended as a methodology demonstration and quality evaluation artifact, not as a training-ready dataset. Why this exists We are evaluating whether LLM-assisted translation of open instruction-tuning datasets into low-resource languages can be done at a quality bar that the ML community will accept. Chinese is our first… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/oasst1-zh-pilot.tabulartext-generationn<1K0 likes14 downloads4mo agoHugging Face21Shrikes /self_repair_parsing_pilot_data Self-Repair SBN Parsing Pilot Data A small, static pilot set for studying how a semantic parser (NL → SBN) behaves when the input natural language carries self-repair disfluencies. Each row is one gold sentence from the Parallel Meaning Bank (PMB) together with synthetically generated repair variants. The dataset holds inputs only — parser predictions are stored elsewhere so this set stays fixed across modelling stages. Source: PMB 5.1.0 (gold), English single-sentence items.… See the full description on the dataset page: https://huggingface.co/datasets/Shrikes/self_repair_parsing_pilot_data.texttext-generationn<1K0 likes12 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.