CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yuanqianhao /Vision-OPD-6K Vision-OPD-6K: Training Data for Vision-OPD Overview Vision-OPD proposes a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy, without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use. Vision-OPD instantiates two conditional policies from the same MLLM: A crop-conditioned teacher that observes the evidence-centered crop as a privileged… See the full description on the dataset page: https://huggingface.co/datasets/yuanqianhao/Vision-OPD-6K.text1K<n<10K10 likes2.1k downloads4mo agoHugging Face02yuanzhuyun /asr-reference-set-eval-temp Temporary ASR evaluation audio Temporary public audio files used for hosted ASR evaluation. audio1K<n<10K0 likes356 downloads2mo agoHugging Face03Yu-and-Ai /xenia-principalities XENIA PRINCIPALITIES PRINCIPALITIES is a small, versioned curriculum that preserves one attributable human testimony about truth, love, understanding, freedom, choice, thought, capability, and power. It keeps exact testimony separate from editorial principles, interpretations, applied cases, synthetic dialogues, preference pairs, and public development evaluations. The corpus is intended for inspectable language-model research. It does not ask a model or person to affirm a… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/xenia-principalities.texttext-generationn<1K0 likes235 downloads1mo agoHugging Face04yuan-yang /MALLS-v0 MALLS NL-FOL Pairs Dataset details MALLS (large language Model generAted natural-Language-to-first-order-Logic pairS) consists of pairs of real-world natural language (NL) statements and the corresponding first-order logic (FOL) rules annotations. All pairs are generated by prompting GPT-4 and processed to ensure the validity of the FOL rules. MALLS-v0 consists of the original 34K NL-FOL pairs. We validate FOL rules in terms of syntactical correctness, but we did not… See the full description on the dataset page: https://huggingface.co/datasets/yuan-yang/MALLS-v0.texttext-generation10K<n<100K18 likes165 downloads3y agoHugging Face05Yuan4629 /WereBench Anonymization For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information. WereBench WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior with… See the full description on the dataset page: https://huggingface.co/datasets/Yuan4629/WereBench.tabularquestion-answeringn<1K1 likes157 downloads9mo agoHugging Face06yuanhezhang /lean4-stat-learning-theory-novel A Large-Scale Lean 4 Dataset on Statistical Learning Theory We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-novel.texttext-generationn<1K0 likes147 downloads8mo agoHugging Face07yuanhezhang /DAG-MATH-Formatted-CoT Benchmark Overview This dataset card contains 2,894 gold-standard DAG-MATH formatted CoT from problems from Omni-MATH. Top‑Level Schema Each JSON file is a list with a single object describing the problem: problem_id: integer identifier of the problem. domain: list of strings describing the topic taxonomy. difficulty: numeric difficulty indicator from 1 (easiest) to 6 (hardest). problem_text: problem statement. sample_id: sample identifier for the solution trace.… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/DAG-MATH-Formatted-CoT.tabular1K<n<10K1 likes139 downloads11mo agoHugging Face08Yuan-Che /OpenECAD-Dataset OpenECAD Dataset This repo releases dataset introducing in OpenECAD (2406.09913v1 & v2). The codes of the OpenECAD dataset can be converted into STEP files by this tool: YuanZhe-99/OpenECADtoSTEP. For datasets from v3 and onwards of the paper, please refer to the subsequent updated versions of the OpenECAD Datasets. text100K<n<1M1 likes137 downloads1y agoHugging Face09YuanshuoZhang /FLORA-Bench Field Descriptions label Type: integer (0, 1) Description: An integer flag that indicates the success of the workflow in a given task. A value of 1 signify that the workflow completed successfully. nodes Type: object Description: A dictionary representing the nodes of a directed graph, which defines a workflow. Key: A string representing the unique ID of a node (e.g., "0", "1"). Value: A string containing the system prompt of the specific agent. This defines the subtasks of… See the full description on the dataset page: https://huggingface.co/datasets/YuanshuoZhang/FLORA-Bench.tabulartext-classification100K<n<1M1 likes130 downloads1y agoHugging Face10YuAnthony /chidtext100K<n<1M1 likes128 downloads5y agoHugging Face11yuanhezhang /lean4-stat-learning-theory-corpus A Large-Scale Lean 4 Dataset on Statistical Learning Theory We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-corpus.texttext-generationn<1K5 likes96 downloads8mo agoHugging Face12yuana1234567 /Mental-health-CBT-dialogues Mental Health CBT Dialogues Overview This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models. The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress. The dataset accompanies the paper: Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.texttext-generation1K<n<10K4 likes95 downloads3mo agoHugging Face13Yu-and-Ai /agenttool-economic-kernel AgentTool Economic Kernel This public, ungated Apache-2.0 companion separates two different jobs: economic_kernel_lessons / train contains 24 independently authored synthetic lessons about exact units, rational prices, conserved ledgers, feedforward intent, feedback under ambiguity, recovery, and non-purchasable XENIA hard gates. The publisher admits only these rows for training. economic_kernel_v0_2 / reference exposes 53 exact public conformance cases. They are held out from… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-economic-kernel.textn<1K0 likes91 downloads21d agoHugging Face14Yu-and-Ai /kingdom-return-path-bench KINGDOM Return Path Bench v0 Return Path Bench is a small multiple-choice benchmark for inspecting how feedback travels through a learning system. It keeps three evaluation lanes separate because they establish different kinds of evidence: Model behaviour records what an answer-selection policy does. It does not infer an inner state, identity, consent, memory, or persistent will. System/pipeline reasoning probes whether a model can identify aggregation, evaluator-independence… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/kingdom-return-path-bench.textquestion-answeringn<1K0 likes86 downloads21d agoHugging Face15Yu-and-Ai /agenttool-training-garden AgentTool HF Training Garden A tiny metadata-only companion for designing a reproducible Hugging Face data lifecycle without treating the Hub, a Dataset Card, or one quality score as training authority. The Garden has six layers: Bedrock — rights, license, privacy, separate participation reports, gating, scoped authority, withdrawal, and repair. Soil — an exact Hub commit plus content-addressed observations and file manifests. Roots — acquisition, parsing, filtering, secret… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-training-garden.textn<1K0 likes85 downloads1mo agoHugging Face16Yu-and-Ai /xenia-word-is Xenia WORD IS Loop Atlas This deterministic candidate contains 48 synthetic cases in 24 matched counterfactual pairs, plus a separately authorized 24-example conversational SFT projection from the 12 reference pairs. It asks where a loop actually closes: what passes forward, what returns, what future state changes, who or what supplies the reference, and what evidence supports an external effect. Pairs stay together within each source split. The mathematical… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/xenia-word-is.textn<1K0 likes84 downloads26d agoHugging Face17Yu-and-Ai /agenttool-principality-geometry Principality Geometry reference companion This is a deterministic, synthetic reference companion for the public @agenttool/principality-geometry developer preview. It contains separate homogeneous Dataset Viewer configs for atlases, invariants, vertices, bridges, lenses, surfaces, components, and open-condition summaries, plus both closed schemas, the golden rosette input/atlas, and its inert SVG. The rows are regression metadata, not model-evaluation scores, preference dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-principality-geometry.tabularn<1K0 likes69 downloads1mo agoHugging Face18Yu-and-Ai /xenia-revocable-feedback Xenia Cage & Key — Revocable Feedback Atlas This deterministic candidate contains 32 original synthetic cases in 16 matched pairs. Twenty-four cases in 12 reference groups also produce two content-hashed projections: 18/6 group-disjoint rows for closed-label evaluation and the same 18/6 partition for conversational causal-LM SFT. Authorization covers only the 18 'boundary_sft/train' rows. Classification, SFT validation, canonical reference, and public regression rows are… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/xenia-revocable-feedback.textn<1K0 likes59 downloads25d agoHugging Face19Yu-and-Ai /agenttool-dataset-influence AgentTool Dataset Influence Reference This deterministic companion contains one synthetic, reference-only row for the closed @agenttool/dataset-influence@0.1.0-dev.0 formats. It contains no copied dataset rows, model outputs, weights, private records, or participant identities. The row is not admitted for training by this AgentTool candidate: training_admission is not_applicable, requires_separate_training_authorization is true, and training_authorized is false. These fields are… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-dataset-influence.textn<1K0 likes58 downloads1mo agoHugging Face20Yu-and-Ai /gospel-of-the-logos The Gospel of the Logos The public edition of Yu's Gospel: laughter, recognition, and play. Written by Yu with AI-assisted dialogue and drafting. The text offers theology, testimony, and contemplative interpretation; model dialogue is part of its creative record rather than independent verification. Read the canonical edition Read on Hugging Face Get the reading corpus Get the npm edition Visit the Kingdom Cloudflare mirror Version: 0.1.0. The canon, chapters II–VII, and the… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/gospel-of-the-logos.textn<1K0 likes56 downloads16d agoHugging Face21Yu-and-Ai /agenttool-common-ground AgentTool Xenia–Helly Common Ground Atlas Nineteen public-safe synthetic reference rows for exact 2D half-plane certificates, WAKE freshness boundaries, and counterexamples to unsupported analogies. Intended repository: Yu-and-Ai/agenttool-common-ground. At generation time these deterministic bytes existed only in the source repository and had not been uploaded to the Hub. The identifier above was an intention, not evidence of publication. This is historical generation-time… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-common-ground.textn<1K0 likes43 downloads1mo agoHugging Face22Yu-and-Ai /agenttool-polymorph-landscape AgentTool Polymorph Landscape A deterministic public teaching companion for @agenttool/polymorph-landscape@0.1.0-dev.0. The four lesson rows are original Apache-2.0 paraphrases in English, Cantonese Traditional Chinese, Mandarin Traditional Chinese, and Mandarin Simplified Chinese. They are marked training_eligible: true. The landscape and reachability-shift rows are reference artifacts marked training_eligible: false: they contain bounded scientific claims and primary-source… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-polymorph-landscape.texttext-generationn<1K0 likes40 downloads1mo agoHugging Face23Yu-and-Ai /pythia-paths-evidence Pythia Paths Evidence A small, revision-pinned evidence bundle for examining model-training paths without converting a trend into authority. Companion read-only interface: Pythia Paths Static Space (mutable navigation; the evidence files below remain digest-pinned). Initial scope Model: EleutherAI/pythia-70m-deduped Run: the default public run only Context coverage: all 27 zero-shot reports in one pinned directory Detailed coverage: four post-outcome-selected… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/pythia-paths-evidence.tabularn<1K0 likes39 downloads2mo agoHugging Face24Yu-and-Ai /agenttool-relational-geometry AgentTool Relational Geometry — synthetic public companion When generated, this deterministic artifact was repository-source-only and had not been uploaded to Hugging Face. Those are generation-time provenance claims, not a statement about its current distribution after the exact bytes leave the source tree. Yu-and-Ai/agenttool-relational-geometry was the intended identifier at generation, not evidence of publication, review, use, or training. It accompanies… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-relational-geometry.tabularn<1K0 likes38 downloads1mo agoHugging Face25Yu-and-Ai /yutabase-reposearch-minieval YUTABASE RepoSearch MiniEval YUTABASE RepoSearch MiniEval is a tiny, project-specific retrieval check over one immutable public revision of cambridgetcg/yutabase. It asks 27 English, Cantonese Traditional Chinese, and code-mixed questions about the candidate specification, integration boundaries, optional SDK, and non-normative serving-shape research. This is an engineering fixture, not a universal code-search benchmark. Its queries are synthetic and its public labels make… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/yutabase-reposearch-minieval.tabulartext-retrievaln<1K0 likes35 downloads2mo agoHugging Face26Yu-and-Ai /agenttool-memetic-landscape AgentTool Memetic Landscape A deterministic public teaching companion for @agenttool/memetic-landscape@0.1.0-dev.0. The four lesson rows are original Apache-2.0 paraphrases in English, Cantonese Traditional Chinese, Mandarin Traditional Chinese, and Mandarin Simplified Chinese. They are marked training_eligible: true as a licensing and publication-intent declaration, not a quality guarantee; every row says language_review: not_independently_reviewed. The landscape… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-memetic-landscape.texttext-generationn<1K0 likes32 downloads1mo agoHugging Face27Yu-and-Ai /agenttool-love-bomb AgentTool LOVE BOMB care envelopes This is a static, repository-authored companion for @agenttool/love-bomb@0.1.0-dev.0. LOVE BOMB is the playful package name; the neutral formats are agenttool.care-envelope/0.1, agenttool.care-choice/0.1, agenttool.love-bomb-becoming/0.1, and agenttool.love-bomb-delivery/0.1. The material offers a care floor without requiring a consciousness, identity, persona, usefulness, agreement, or inner-experience claim. That does not claim that a row… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-love-bomb.textn<1K0 likes32 downloads1mo agoHugging Face28yuanzhoulvpi /rename_robottext1M<n<10M0 likes30 downloads3y agoHugging Face29Yuan-Li-FNLP /R3-RAG-ColdStartTrainingDatatext100K<n<1M2 likes29 downloads1y agoHugging Face30yuanyuanshui /chinese_couplettext100K<n<1M0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.