CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes5.4k downloads2y agoHugging Face02kenhktsui /github-code-permissive-sampleSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'].It is intended to be used for training code language classifier. texttext-classification1M<n<10M0 likes955 downloads2y agoHugging Face03ontocord /finephrase_permissivegatedtabular10M<n<100M0 likes561 downloads4mo agoHugging Face04powertronglobal /powertron-global-permafrost-corpus Dataset Card: Powertron Global PermaFrost Corpus Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.tabulartext-generation1K<n<10K0 likes302 downloads2mo agoHugging Face05ontocord /moral_education_permissivegated Moral Education Permissive Permissive-source subset of locuslab/moral_education, filtered using row-level metadata.url with the local permissiveness rules in this workspace. The output preserves the original row fields and adds url, idx, dump, language, source_config, is_permissive, is_oss, and permissive_reason where applicable. Source configs included: score_4_morals, score_5_morals. Source split: train. Rows scanned: 2,806,450. Rows kept: 72,943. Keep rate: 2.5991%. See… See the full description on the dataset page: https://huggingface.co/datasets/ontocord/moral_education_permissive.texttext-generation10K<n<100K0 likes296 downloads4mo agoHugging Face06permutans /c4-bbc-news Dataset Card for BBC News from C4 This dataset provides a filtered subset of BBC News articles from the realnewslike subset of the C4 dataset, containing approximately 77k articles from BBC News domains. Dataset Details Dataset Sources Repository: https://huggingface.co/datasets/permutans/c4-bbc-news Source Dataset: allenai/c4 (realnewslike subset) Paper: https://arxiv.org/abs/1910.10683 (C4 paper) Uses Direct Use Suitable for text… See the full description on the dataset page: https://huggingface.co/datasets/permutans/c4-bbc-news.text10K<n<100K0 likes266 downloads2y agoHugging Face07Tasmay-Tib /Word_Permuations_Englishtext1B<n<10B0 likes231 downloads1y agoHugging Face08ustclsc /PERMA PERMA: Benchmarking Personalized Memory Agents TL;DR PERMA is a benchmark for evaluating personalized memory agents in long-horizon conversations where user preferences evolve over time.Instead of static retrieval, models must track event-driven preference evolution and maintain persona consistency under realistic interaction noise. This dataset supports two complementary evaluation protocols: Multiple-choice evaluation for granular capability probing (task completion… See the full description on the dataset page: https://huggingface.co/datasets/ustclsc/PERMA.text1K<n<10K0 likes206 downloads5mo agoHugging Face09scikit-fingerprints /ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB ExpansionRx-OpenADMET Caco-2 Permeability Papp A>B Caco-2 Permeability Papp A>B dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through scikit-fingerprints library. The task is to predict Caco-2 Permeability Papp A>B of molecules. Characteristic Description Tasks 1 Task type regression Total samples 3773 Recommended split time Recommended metric MAE References [1] OpenADMET team "Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB.texttabular-regression1K<n<10K0 likes183 downloads6mo agoHugging Face10t2ance /atlas-16-verifier-permission-prompt-ablation ATLAS report 16: does the orchestrator's "cannot solve" clause suppress candidate verification? Complete raw products of the ATLAS rl-training report 16 experiment (GitHub issue #36). Two system-prompt arms of the same model over the same 78 fixed states, greedy decoding, one shared vLLM server. What the experiment did The ATLAS orchestrator's frozen system prompt contains the clause You cannot solve the problem yourself; you decide when to explore further and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-16-verifier-permission-prompt-ablation.texttext-generation10K<n<100K0 likes172 downloads14d agoHugging Face11laion /freesound-commercially-permissive-subset-with-captionsaudio100K<n<1M2 likes166 downloads11mo agoHugging Face12scikit-fingerprints /ExpansionRx_OpenADMET_Caco-2_Permeability_Efflux ExpansionRx-OpenADMET Caco-2 Permeability Efflux Caco-2 Permeability Efflux dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through scikit-fingerprints library. The task is to predict Caco-2 Permeability Efflux of molecules. Characteristic Description Tasks 1 Task type regression Total samples 3777 Recommended split time Recommended metric MAE References [1] OpenADMET team "Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Efflux.texttabular-regression1K<n<10K0 likes132 downloads6mo agoHugging Face13permutans /dummy-lang-subset-dataset-1m-chunksDummy dataset uploaded to test the process, benefits of and constraints on/drawbacks to uploading subsets by a partition of a dataset. This dataset is made up of fake data to illustrate a long tail of rarer languages similar to Wikipedia/Wikidata's distribution. The dataset metadata was written automatically by the datasets library. The config name is the language, and we iterate over all languages to do this. Since the data is synthetic, we have the list of languages as a variable, without… See the full description on the dataset page: https://huggingface.co/datasets/permutans/dummy-lang-subset-dataset-1m-chunks.text1M<n<10M0 likes121 downloads1y agoHugging Face14NaghmehAI /PerMedCQA PerMedCQA: Persian Medical Consumer QA Benchmark PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian PerMedCQA is the first large-scale, real-world benchmark for Persian-language medical consumer question answering. It contains anonymized medical inquiries from Persian-speaking users paired with professional responses, enabling rigorous evaluation of large language models in low-resource, health-related domains. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NaghmehAI/PerMedCQA.text10K<n<100K4 likes115 downloads1y agoHugging Face15PRMfinetune /permutation_invariant_rewardtext10K<n<100K1 likes101 downloads1y agoHugging Face16DataProvenanceInitiative /common_pile_ultra_permissive Dataset Card for Data Provenance Initiative - Common-Pile-Ultra-Permissive Legal Disclaimer / Notice Collected License Information is NOT Legal Advice. It is important to note we collect self-reported licenses, from the papers and repositories that released these datasets, and categorize them according to our best efforts, as a volunteer research and transparency initiative. The information provided by any of our works and any outputs of the Data Provenance Initiative do… See the full description on the dataset page: https://huggingface.co/datasets/DataProvenanceInitiative/common_pile_ultra_permissive.text1M<n<10M0 likes94 downloads2y agoHugging Face17uw-math-ai /theorem-search-dataset-permissive Theorem Search Dataset The largest open corpus of informal mathematical theorems: 1,239,720 theorem statements with natural-language slogans from 197,889 papers, designed for semantic theorem retrieval. Paper: Semantic Search over 9 Million Mathematical Theorems Demo: huggingface.co/spaces/uw-math-ai/theorem-search Benchmark results On 110 test queries written by research mathematicians, our best pipeline (Qwen3-Embedding-8B on DeepSeek-V3.1 slogans) outperforms all… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/theorem-search-dataset-permissive.textquestion-answering1M<n<10M0 likes91 downloads7mo agoHugging Face18Ailiance-fr /kicad9plus-permissive Ailiance — KiCad 9+ Schematic Corpus (Permissive) 🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kicad9plus-permissive. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025). Corpus de 98 schémas KiCad 9+ (.kicad_sch, format S-expression, version ≥ 20240722) collectés sous licences permissives uniquement (Apache-2.0, MIT, CC0-1.0, CERN-OHL-P-2.0). Pensé pour l'entraînement et le fine-tuning de modèles de génération… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kicad9plus-permissive.texttext-generationn<1K0 likes91 downloads4mo agoHugging Face19proshady2 /Word_Permuations_Englishtext1B<n<10B0 likes84 downloads9mo agoHugging Face20cowWhySo /permission-command-corpus Permission Command Corpus Three views of a command-safety corpus, for local command-risk classification in front of an LLM or a tool bridge. gold: trusted rows, 321 in total across three splits silver_weak_labels: mined weak-label rows from Sigma, LOLBAS, GTFOBins, Atomic Red Team and Falco, kept as useful but not promoted to gold review_queue: unresolved rows that should not be treated as trusted training data Read this before training on it Findings from 30… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/permission-command-corpus.tabulartext-classification1K<n<10K1 likes81 downloads23d agoHugging Face21thanna94 /us-building-permits PermitBase — U.S. Residential Building Permits 1980–2024 The most comprehensive historical residential building permit dataset available at the place level. Place-level | Annual | 1980–2024 | SF/MF differentiated | 51 jurisdictions | 683,986 records Dataset Description This dataset contains annual residential building permit data for permit-issuing places (cities, towns, and unincorporated county areas) across the United States, covering 1980 through 2024. It is derived… See the full description on the dataset page: https://huggingface.co/datasets/thanna94/us-building-permits.tabulartime-series-forecasting100K<n<1M0 likes80 downloads6mo agoHugging Face22permitt /serbian-llm-eval Serbian LLM Eval This dataset is a republish of gordicaleksa/serbian-llm-eval-v1 as plain Parquet configs, so it can be loaded with modern versions of the datasets library (the upstream repo ships a Python loading script that datasets can no longer execute). Row contents are otherwise unchanged; only the example_id field is synthesized where the upstream data has no unique identifier, and the triviaqa answer struct is flattened into answer_value / answer_aliases columns.… See the full description on the dataset page: https://huggingface.co/datasets/permitt/serbian-llm-eval.text100K<n<1M0 likes79 downloads2mo agoHugging Face23mrmegatelo /PineScripts-Permissive PineScripts-Permissive A dataset of Pine Script™ scripts with premissive licenses from TradingView. tabular1K<n<10K6 likes78 downloads3y agoHugging Face24Aasdfip /oracle_perm_medm_dataimage10K<n<100K0 likes71 downloads1y agoHugging Face25icl-heads /permutation-pools Permutation Pools Permutation-augmentation pools: demonstration-order permutations of the label-transfer and evidence-tracking sets, used to enlarge filtered test sets to 1000 rows with balanced golds. Part of the icl-heads collection: the datasets behind a study of which Llama-3.1-8B-Instruct attention heads mediate in-context evidence accumulation and in-context label mapping, discovered with Differentiable Circuit Masking (a learned per-head mask over K/V activations patched… See the full description on the dataset page: https://huggingface.co/datasets/icl-heads/permutation-pools.tabular10K<n<100K0 likes71 downloads15d agoHugging Face26theepicflyer /openrtlset-permissive-verified openrtlset-permissive-verified A permissive-filtered, machine-verified subset of ESCAD/OpenRTLSet. Every record in this dataset compiles. Each completion was elaborated and linted with Verilator (--lint-only -Wall, warnings fatal) and admitted only on a clean run. Nothing here is graded by an LLM judge. Contents 13,625 records drawn from 2,174 distinct upstream repositories Task: specification + pinned module interface -> complete Verilog module Upstream… See the full description on the dataset page: https://huggingface.co/datasets/theepicflyer/openrtlset-permissive-verified.texttext-generation10K<n<100K0 likes69 downloads2mo agoHugging Face27mondk /r0b0tlab-distillation-permissive-sft Filtered SFT Dataset (Permissive Licenses Only) Source: r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation Filtered to keep only rows under permissive licenses (Apache-2.0, MIT, CC-BY-4.0), suitable for training. Rows under non-commercial or unclear licenses were removed. Converted to JSONL, one record per line: {"instruct": "...", "output": "..."} License: mixed (Apache-2.0 / MIT / CC-BY-4.0) — see original dataset card for per-source license and attribution details. text1K<n<10K1 likes69 downloads19d agoHugging Face28jiosephlee /assay-transfer-record-level-v23-1-bbb-martins-passive_permeability-intern BBB V23.1: passive_permeability Exact passive_permeability row filter of pinned BBB V23 revision d49645551057f0dfc562dd7bf2ef7d77ff975a6f. Prompts, targets, candidate order, and split membership are unchanged. The complete parent calibration artifact is retained. V23 has no test split. train rows: 46,541 validation ranking rows: 1,525 tabular10K<n<100K0 likes68 downloads18d agoHugging Face29universitytehran /PerMed-MM PerMed-MM: A Multimodal, Multi-Specialty Persian Medical Benchmark 🤗 Dataset | 📖 Paper | 📄 PDF Dataset Description PerMed-MM is a multimodal, multi-specialty benchmark designed to evaluate Vision Language Models (VLMs) on Persian medical question answering. The dataset consists of 733 multiple-choice questions sourced from the Iranian National Medical Board Exams (years 2021 and 2023). Each question is paired with 1 to 5 clinically relevant images, totaling… See the full description on the dataset page: https://huggingface.co/datasets/universitytehran/PerMed-MM.imagevisual-question-answeringn<1K1 likes67 downloads3mo agoHugging Face30permitt /superglue-hr BalkanBench SuperGLUE - Croatian Part of BalkanBench - the open, reproducible benchmark for language models across Serbian, Croatian, Montenegrin, and Bosnian (BCMS). Live leaderboard at https://balkanbench.com/leaderboard. Background: Release of BalkanBench - the vision behind it (Medium, 2026-04-27). This is the Croatian SuperGLUE released preview track of BalkanBench v0.1. Croatian publishes alongside Serbian (the official frozen track) and Montenegrin as a 5-task preview:… See the full description on the dataset page: https://huggingface.co/datasets/permitt/superglue-hr.tabulartext-classification10K<n<100K1 likes67 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.