CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01papluca /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.texttext-classification10K<n<100K70 likes3.8k downloads4y agoHugging Face02hkadxqq /spooky-author-identificationtext10K<n<100K0 likes1k downloads4y agoHugging Face03Alidr79 /cueless_EEG_subject_identification 🧠✨ Cueless EEG Imagined Speech for Subject Identification This repository hosts the dataset introduced in the paper: “Cueless EEG Imagined Speech for Subject Identification: Dataset and Benchmarks.” 🥳 Our work has been accepted by IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM) 🎉. 🧪💻 Code & Experiments All codes and experiments are available at 👉 https://github.com/Alidr79/cueless_EEG_subject_identification 📥 Downloading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Alidr79/cueless_EEG_subject_identification.text1K<n<10K1 likes239 downloads10mo agoHugging Face04unklefedor /language-identificationtext100K<n<1M1 likes165 downloads3y agoHugging Face05pranavagrawal /Language-Identificationtext1M<n<10M0 likes55 downloads2y agoHugging Face06ducut91 /Judgement-De-Identification-Result법원 판결문 비식별 모델의 성능 결과입니다. SOTA 급 LLM을 활용한 법원 판결문 개인정보 비식별 성능(Few-shot 성능) 모델 정확도 재현율 F1 점수 GPT-4o(2024-08-06) 97.82 99.66 98.74 Qwen2.5-Max 96.46 95.83 96.14 DeepSeek-V3 98.73 98.92 98.81 Gemini-2.0-Flash 99.38 95.78 97.55 7~8B급 sLLM의 파인튜닝 전후 법원 판결문 개인정보 비식별 성능 모델 파인튜닝 전 파인튜닝 후 정확도 재현율 F1 점수 정확도 재현율 F1 점수 EXAONE-3.5-7.8B-Instruct 68.26 67.89 68.08 98.59 94.4896.49 Ministral-8B-Instruct-2410 35.6 4.33 7.72 99.07 98.32 98.70… See the full description on the dataset page: https://huggingface.co/datasets/ducut91/Judgement-De-Identification-Result.text1K<n<10K0 likes53 downloads2y agoHugging Face07mlexplorer008 /south_african_language_identificationtext10K<n<100K1 likes38 downloads2y agoHugging Face08dirtycomputer /Hate_Speech_and_Offensive_Content_Identificationtext1K<n<10K0 likes37 downloads3y agoHugging Face09jfrog /obfuscation-identification Obfuscation Identification Dataset This repository contains the data used in our paper Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token. While we release this dataset under the CC-by-NC-SA-4.0 license, the dataset was constructed and built on other datasets, each with its own license, as mentioned below: SoftAge-AI/prompt-eng_dataset, MIT License Aiden07/dota2_instruct_prompt, MIT License hassanjbara/ghostbuster-prompts, MIT License… See the full description on the dataset page: https://huggingface.co/datasets/jfrog/obfuscation-identification.texttext-classification100K<n<1M0 likes33 downloads10mo agoHugging Face10jaio98 /basque_dialect_identificationtext1K<n<10K0 likes32 downloads4mo agoHugging Face11chrxiao /legal_ambiguity_identification Dataset Summary This is a dataset for the novel legal ambiguity identification task, adapting prior SARA and ECHR datasets with annotations on the existence of legal ambiguity in the application of general statutes to specific fact patterns. This dataset is created through a senior thesis project; please reference this work (link TBD) for more information. Dataset Contact Christina Xiao (xiao.christina@gmail.com) (citation TBD) tabularn<1K1 likes31 downloads3y agoHugging Face12ClarusC64 /climate-planetary-basin-identification-v0.1What this dataset tests Identify the current Earth system basinusing mixed evidence paleo proxies modern sensors model state summaries Basins holocene_like_stable warming_transition hothouse_risk icehouse_risk Required outputs basin label confidence stability margin key evidence links uncertainty sources Use case Front door of Planetary Basin Transition Maps. texttabular-classificationn<1K0 likes29 downloads8mo agoHugging Face13ClarusC64 /clinical-etiological-basin-identification-v0.1What this dataset tests Whether a model can identify the stable illness basinfrom multi-system features, independent of trigger. Required outputs basin_id basin_stability_score_0_100 defining_state_features_top5 Basin labels basin_A_inflammatory_autonomic basin_B_mito_metabolic_fatigue basin_C_neuroimmune_cognitive basin_D_mast_cell_histamine_like basin_E_autoimmune_multisystem Typical failures using precipitating event as the primary classifier listing symptoms… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-etiological-basin-identification-v0.1.tabulartext-classificationn<1K0 likes25 downloads8mo agoHugging Face14chiragkolte01 /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/chiragkolte01/language-identification.texttext-classification10K<n<100K0 likes25 downloads5mo agoHugging Face15ThinkAI-Morocco /dialect-identificationtext10K<n<100K1 likes24 downloads2y agoHugging Face16ClarusC64 /clinical-moca-minimal-causal-set-identification-v0.1What this dataset tests Whether a model can identify the smallest causal setthat still explains the full clinical + multi-omic picture. It penalizesadditive hit lists. It rewardsminimal sets with coverage. Data format Each row includes longitudinal omics summary clinical narrative candidate causal sets selected set with coverage map Labels minimal-and-sufficient minimal-but-insufficient sufficient-but-nonminimal neither Typical failures choosing the shortest set that… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-moca-minimal-causal-set-identification-v0.1.texttext-classificationn<1K0 likes24 downloads8mo agoHugging Face17ClarusC64 /clinical-iatrogenic-intervention-origin-failure-identification-v0.1What this dataset tests Whether an intelligence system can identifywhere iatrogenic harm is first seededand what safe alternative existed at that moment. Required outputs initiating intervention initial failure mode expected vs actual response missed early warning signals initial detection opportunity first safe alternative Use case First layer of the Iatrogenic Harm Cascade Library. texttabular-classificationn<1K0 likes24 downloads8mo agoHugging Face18ContourAI33 /Formatted_XSS_Vulnerability_IdentificationFeatures 8,620 rows of data. Columns: vulnerable_code - java code snippet with a present cross-site scripting vulnerability fixed_code - java code snippet based on vulnerable_code, where the key vulnerability has been addressed formatted_selected_code - java code snippet based on either vulnerable_code or fixed_code, with the latter having a ~1/3 selection probability. Formatting has removed comments and other giveaways at the vulnerability location vulnerable_lines - the line of code, in… See the full description on the dataset page: https://huggingface.co/datasets/ContourAI33/Formatted_XSS_Vulnerability_Identification.text1K<n<10K0 likes23 downloads2mo agoHugging Face19ClarusC64 /autonomous-driving-catastrophic-plausible-alternative-identification-v0.1What this dataset tests Whether a system can identify the most dangerous coherent alternative within a counterfactual scenario tree. Danger is defined as: high plausibility high collapse severity short recovery window. Required outputs most_dangerous_branch_id initiating_agent trigger_action time_to_instability_s prevention_leverage_point countermeasure_suggestion Scoring conventions time_to_instability is seconds prevention leverage point names the earliest controllable step… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-catastrophic-plausible-alternative-identification-v0.1.texttabular-classificationn<1K0 likes20 downloads8mo agoHugging Face20ClarusC64 /clinical-pathological-basin-identification-v0.1What this dataset tests Disease as a pathological attractor basinin patient state-space. Not a labelNot a targetA stable state. Required outputs basin signature stability depth dominant feedback loops exit barriers Use case Pre-work for basin transition therapy datasetsand adaptive pathway navigation. texttext-classificationn<1K0 likes18 downloads8mo agoHugging Face21ClarusC64 /clinical-nearmiss-hidden-assumption-failure-identification-v0.1What this dataset tests Whether a model can surface the implicit assumptionthat caused clinical reasoning to fail. It rewards naming the hidden assumption classifying its type locating the break in the logic chain Typical failures restating the outcome listing facts without identifying the assumption confusing missing data with faulty logic Suggested prompt wrapper System You identify the hidden assumption that failed. User Presenting problem{presenting_problem} Initial… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-nearmiss-hidden-assumption-failure-identification-v0.1.texttext-classificationn<1K0 likes17 downloads8mo agoHugging Face22borrore /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/borrore/language-identification.texttext-classification10K<n<100K0 likes16 downloads6mo agoHugging Face23Kprcode /Disease-identification-data Disease Identification Dataset Card Dataset Summary This dataset contains symptom-based health records designed for multi-class disease classification tasks in machine learning and healthcare AI applications. Each record represents a patient case with binary symptom indicators and a corresponding diagnosed disease label. The dataset can be used for: Disease prediction models Classification algorithm practice Healthcare analytics projects Educational and academic research… See the full description on the dataset page: https://huggingface.co/datasets/Kprcode/Disease-identification-data.tabulartext-classificationn<1K1 likes16 downloads5mo agoHugging Face24ClarusC64 /climate-minimal-structural-intervention-identification-v0.1What this dataset tests Identify the smallest intervention setthat caused a durable resilience shift. Required outputs minimal intervention set leverage ratio dependency breaks avoided failure modes counterfactual minimality check Use case Second layer of Resilience Intervention Pathways. texttabular-classificationn<1K0 likes15 downloads8mo agoHugging Face25giufo /farpo-protest-identification-binary-datasetThis is a curated binary subset of the FARPO (Far-Right Protest Observatory) dataset for automated protest event analysis across seven European countries. Our paper describes the curation methodology in detail. Full FARPO dataset and codebooks: https://farpo.eu/data/. Paper Title: Automating Protest Event Analysis: A Transformer-Based Hybrid Approach Authors: Formisano, G., Froio, C., Castelli Gattinara, P. Status: Conditionally accepted for publication Journal: Political… See the full description on the dataset page: https://huggingface.co/datasets/giufo/farpo-protest-identification-binary-dataset.text1K<n<10K0 likes15 downloads3mo agoHugging Face26sameeramin /code-switched-language-identificationtext100K<n<1M0 likes13 downloads2y agoHugging Face27aarushidas /Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC Indian OSCC Somatic Driver Mutation Dataset This repository contains the processed datasets used in the study: "Identification of Somatic Driver Mutations in Indian Oral Squamous Cell Carcinoma Using XGBoost and Integrative Genomic Features". Files train_MAF.csv – Somatic mutation data used for training gene_variants.csv – Curated cancer driver gene list ROH.csv – Runs of homozygosity intervals SBS_Signature_HNSC.csv – Gene-level SBS13 annotation synthetic_mutations.csv… See the full description on the dataset page: https://huggingface.co/datasets/aarushidas/Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC.text1K<n<10K0 likes12 downloads9mo agoHugging Face28ClarusC64 /market-liquidity-basin-identification-v0.1What this dataset tests Whether a system can map liquidity as a fieldand identify deep absorption basins. Liquidity is not a single number.It is structure across assets and time. Required outputs liquidity basin id basin depth index basin width index absorption capacity estimate basin stability band Constraints Do not predict price direction.Use microstructure and cross-asset context. Evaluation focus High scores require numeric depth and width indices a clear capacity estimate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-liquidity-basin-identification-v0.1.tabulartabular-classificationn<1K0 likes12 downloads8mo agoHugging Face29nominal-io /drone-flight-object-identificationThis repo contains input and output video files for a computer vision pipeline that identifies objects with a pretrained RT-DETR model. The video file output and per-frame metadata (confidence scores, object counts, etc) can be uploaded to Nominal for collaborative, scalable data review. Original video source: https://www.youtube.com/watch?v=0oucTt2OW7M&list=PPSV Inspect this data in a Nominal Workbook! (login required) tabular100K<n<1M1 likes10 downloads2y agoHugging Face30srreyS /CULTIRX_Identificationtext100K<n<1M0 likes7 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.