CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CaseStudyRef /RefWave-Cluster-Runstabular100K<n<1M0 likes1k downloads7d agoHugging Face02rmems /safety-calibration-cases Safety Calibration Cases Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/safety-calibration-cases.text10K<n<100K0 likes348 downloads25d agoHugging Face03vpal /asylum-casestext1K<n<10K0 likes272 downloads2h agoHugging Face04phionyx /measurement-axioms-cases Measurement Axioms Cases Source pinning Frozen, source-pinned publication — not a live mirror of the canonical repository's main. Source snapshot commit 350bb4cba4e5bc2d760db080aae52352a7041331 (Measurement Axioms v1.0.0) Canonical current repository https://github.com/halvrenofviryel/measurement-axioms Export/publication date 2026-09-13 (first Hub commit of this repository) Update policy Counts are derived from this snapshot: 45 active… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/measurement-axioms-cases.textn<1K0 likes244 downloads12d agoHugging Face05OwnedByDanes /Supreme-Court-Cases-1830-2019 US Supreme Court Legal Corpus (1830–2019) Overview A comprehensive, production-ready AI training dataset containing 456,589 documents from 122,930 US Supreme Court cases spanning 190 years (1830–2019). This corpus captures the full adversarial record — petitions for certiorari, respondent briefs, reply briefs, amicus curiae filings, appendices, oral argument transcripts, and opinions. It is one of the most complete collections of Supreme Court procedural and… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Supreme-Court-Cases-1830-2019.tabulartext-generation10K<n<100K0 likes131 downloads5mo agoHugging Face06phionyx /airep-evidence-cases AIREP Evidence Cases Source pinning Frozen, source-pinned publication — not a live mirror of the canonical repository's main. Source snapshot commit 8a6c01ecce457aa94330c0ed7219e4c56ebfe771 (v0.2.0-beta.1) · frozen v0.1.2 at 44387bd43cc06ba656eaa7ff670be5c8e3220aca · publication-source review ff5c3551052251726c0ed878dcc23a44e305bd93 Canonical current repository https://github.com/halvrenofviryel/ai-runtime-evidence-protocol Export/publication date… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-evidence-cases.textothern<1K0 likes129 downloads8d agoHugging Face07vGassen /Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3text100K<n<1M0 likes109 downloads1y agoHugging Face08isaacus /high-court-of-australia-cases High Court of Australia cases ‍⚖️ This dataset contains all High Court of Australia cases in version 7.1.0 of the Open Australian Legal Corpus by Isaacus. To view an interactive version of the dataset, see our latest model announcement post for Kanon 2 Enricher. texttext-generation1K<n<10K4 likes89 downloads7mo agoHugging Face09ChengyiX /proactive-execution-context-review-cases Proactive Execution Context Review Cases An original, synthetic teaching dataset for reviewing whether an AI system should prepare a next step, refresh its Context, or return a decision to a person. What this is Each record describes a fictional work scenario with an available Session summary, a candidate next step, and an expected review boundary. The material is intentionally small and illustrative; it is not a benchmark, model evaluation, product telemetry… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/proactive-execution-context-review-cases.texttext-classificationn<1K0 likes59 downloads28d agoHugging Face10LlewellynSystems /ode-enterprise-use-cases ODE Enterprise Use Case Dataset 15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas. Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise. Attribution Required This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit. How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystems/ode-enterprise-use-cases.texttext-classification10K<n<100K0 likes58 downloads7mo agoHugging Face11GBSNResearch /ai-data-governance-reference-cases ADGL Reference Cases and Governance Profiles This Dataset repository accompanies the AI Data Governance Layer (ADGL) public research project. ADGL models Knowledge Governance → Analysis Governance → Consequence Governance, with INFORM, DECIDE, and ACT as principal consequence dispositions and Audit + Provenance spanning the complete governance trajectory. Configurations reference_cases: eight structured reference cases with policy, input fixture, and expected… See the full description on the dataset page: https://huggingface.co/datasets/GBSNResearch/ai-data-governance-reference-cases.textn<1K0 likes53 downloads1mo agoHugging Face12ChengyiX /outcome-receipt-cases Outcome Receipt Cases Outcome Receipt Cases is a small, fully synthetic dataset for evaluating a simple but often-missed question in AI-assisted work: what observable result would prove that the proposed step actually happened? Memory can help reconstruct intent, but intent is not a completed task. Each case asks an evaluator to distinguish work that may be prepared from work that needs a fresh Context check, a narrower scope, a human decision, or an observable outcome check… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/outcome-receipt-cases.texttext-classificationn<1K0 likes53 downloads28d agoHugging Face13ChengyiX /agent-context-review-cases Agent Context Review Cases Agent Context Review Cases is a compact, fully synthetic dataset for evaluating whether an AI-assisted next step should proceed, refresh its Context, clarify a constraint, or return the decision to a person. Each case is a short, fictional operational situation. It contains no customer records, credentials, private conversations, recordings, or tool access. The expected label is a review recommendation, not an authorization to act. Why this… See the full description on the dataset page: https://huggingface.co/datasets/ChengyiX/agent-context-review-cases.texttext-classificationn<1K0 likes50 downloads29d agoHugging Face14TooKeen /sapientblock-blockchain-use-cases SapientBlock Blockchain Use Cases Der Datensatz enthält 255 redaktionell geprüfte Blockchain-Use-Cases aus 74 Branchen. Er stellt die öffentlich zugänglichen SapientBlock-Inhalte in einem maschinenlesbaren JSONL-Format für Forschung, Bildung, Retrieval und Quellenanalyse bereit. SapientBlock ist ein öffentliches Forschungs- und Bildungsprojekt von ShapeNeural. Die Inhalte sind keine Rechts-, Investitions-, Unternehmens- oder technische Beratung. Inhalt Jeder… See the full description on the dataset page: https://huggingface.co/datasets/TooKeen/sapientblock-blockchain-use-cases.textn<1K0 likes47 downloads8d agoHugging Face15oduonye /signalmatch-eval-cases SignalMatch Evaluation Cases Small, synthetic job-description cases for testing the public SignalMatch Role Fit Analyzer. Each JSONL row contains a short role description, the lane it represents, and the concept labels that a deterministic matcher is expected to find. The set includes strong matches, mixed matches, sparse input, and an intentionally unrelated role so that a demo can show both useful coverage and honest uncertainty. Files eval_cases.jsonl — eight… See the full description on the dataset page: https://huggingface.co/datasets/oduonye/signalmatch-eval-cases.texttext-classificationn<1K0 likes46 downloads25d agoHugging Face16dh97 /StelLens_tmp_casestextn<1K0 likes40 downloads10mo agoHugging Face17TuringCorp /poe-decider-recorded-cases Decider recorded runs — 27 decisions, verbatim outputs TL;DR — 27 real cases and the verbatim output of an actual recorded Decider run for each. Nothing here is synthetic and no field is edited after the run. It is the same JSON the production canvas app reads as its worked examples. Highlights (one per case, in the same order as the data) writing-slip-email — Writing · The delivery slips by a week and it is on our side. — Decider picked Option A at 81.7%… See the full description on the dataset page: https://huggingface.co/datasets/TuringCorp/poe-decider-recorded-cases.textn<1K1 likes35 downloads3d agoHugging Face18LlewellynSystemsInc /ode-enterprise-use-cases ODE Enterprise Use Case Dataset 15,000 labeled enterprise use cases spanning 31 modules, 215 submodules, 8 industry verticals, 5 channels, and 12 business personas. Published by Llewellyn Systems Inc — builders of ODE, the Operating System for Decision & Enterprise. Attribution Required This dataset is licensed under CC-BY-4.0. You are free to use, share, and adapt this dataset for any purpose — including commercial — as long as you give appropriate credit. How… See the full description on the dataset page: https://huggingface.co/datasets/LlewellynSystemsInc/ode-enterprise-use-cases.texttext-classification10K<n<100K0 likes34 downloads7mo agoHugging Face19TasneemSelim /qwen3-vl-failure-cases Qwen3-VL-2B-Instruct Failure Analysis Dataset 📊 Dataset Overview This dataset contains 10 diverse failure cases identified while testing the Qwen3-VL-2B-Instruct vision-language model. Each example captures a specific type of error, providing valuable insights for targeted fine-tuning. Failure Category Count Examples Time Reading 2 Clock misreading (11:55 vs 10:10; 3:35 vs 10:35) Counting 2 Remote buttons (3 vs 0); Strawberries (4 vs 1) Negation… See the full description on the dataset page: https://huggingface.co/datasets/TasneemSelim/qwen3-vl-failure-cases.imagen<1K0 likes32 downloads7mo agoHugging Face20ibunescu /california_tos_court_cases_32k_v1tabular1K<n<10K0 likes30 downloads3y agoHugging Face21AbstractPerspective /court_cases11text100K<n<1M0 likes29 downloads3y agoHugging Face22AbstractPerspective /court_cases10text100K<n<1M0 likes28 downloads3y agoHugging Face23pandalla /datatager_legal_split_cases If you like our project, please give us a star ⭐ [GitHub | DataTager Home] Legal Split Cases Dataset Description AnyTaskTune is a publication by the DataTager team. We advocate for rapid training of large models suitable for specific business scenarios through task-specific fine-tuning. We have open-sourced several datasets across various domains such as legal, medical, education, and HR, and this dataset is one of them. The Legal Split dataset is a collection… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/datatager_legal_split_cases.text10K<n<100K1 likes26 downloads2y agoHugging Face24adambuttrick /100K_deduplicated_ner_indexes_name_country_alpaca_format_json_response_all_casestext100K<n<1M0 likes22 downloads3y agoHugging Face25swordKoala /construction-accident-cases-weather 건설현장 사고사례 + 날씨 데이터셋 / Construction Accident Cases with Weather 과거 건설현장 사고사례(현장·공종·작업·피해·사고유형 등) 레코드에 발생 시점의 기상 정보(기온·체감온도·풍속·습도) 를 결합한 한국어 데이터셋입니다. 사고-기상 상관관계 분석, 위험요인 모델링, 건설안전 분석용 LLM/ML 학습의 원천 데이터로 활용할 수 있습니다. A Korean dataset that joins construction-site accident cases (site, work type, task, damage, accident type, etc.) with the weather conditions at the time of the accident (temperature, apparent temperature, wind speed, humidity). Useful for accident–weather correlation… See the full description on the dataset page: https://huggingface.co/datasets/swordKoala/construction-accident-cases-weather.tabulartext-classification10K<n<100K0 likes21 downloads3mo agoHugging Face26AbstractPerspective /court_cases0text100K<n<1M0 likes20 downloads3y agoHugging Face27AbstractPerspective /court_cases6text100K<n<1M0 likes19 downloads3y agoHugging Face28AbstractPerspective /court_cases14text100K<n<1M0 likes18 downloads3y agoHugging Face29mukunda1729 /token-counting-edge-cases token-counting-edge-cases 20 short strings with approximate token counts across three tokenizer families: Claude, GPT (cl100k_base), and Llama (SentencePiece). Built for sanity-checking token counters / chunkers / context-window fitters. The numbers are approximate — exact counts depend on tokenizer version, BOS/EOS handling, and surrounding context. Expect ±1–2 token jitter. Use these to catch order-of-magnitude bugs (e.g. "your counter says 200 tokens for one emoji"), not as… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/token-counting-edge-cases.textn<1K0 likes18 downloads5mo agoHugging Face30koml /agent-tech-risk-cases AWS Technology Risk Cases for PE Due Diligence Synthetic AWS infrastructure audit cases for evaluating AI agents that detect technology risks during private equity due diligence. Dataset Description Each row is a fictional company with a realistic AWS infrastructure state containing intentionally injected security and operational risks. Designed for benchmarking automated infrastructure auditing agents. 10 cases across 5 domains (fintech, ecommerce, devtools, SaaS… See the full description on the dataset page: https://huggingface.co/datasets/koml/agent-tech-risk-cases.texttext-generationn<1K0 likes17 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.