CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aisingapore /Safety-Toxicity-Detectiongated SEA Toxicity Detection SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese. Supported Tasks and Leaderboards SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore. Languages Indonesian (id) Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.texttext-generation1K<n<10K0 likes2.5k downloads9mo agoHugging Face02Lots-of-LoRAs /task386_semeval_2018_task3_irony_detection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task386_semeval_2018_task3_irony_detection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task386_semeval_2018_task3_irony_detection.texttext-generation1K<n<10K0 likes127 downloads2y agoHugging Face03DaftP /Home-Assistant-requests-for-intent-detection-and-function-recognition Home Assistant Requests V2 Dataset This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant. The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.textquestion-answering100K<n<1M1 likes112 downloads5mo agoHugging Face04alirezaaminzadeh /sigmaforge-detection-rules SigmaForge Detection Rules SigmaForge is a structured, operational dataset for building and evaluating systems that generate, validate, and translate Sigma detection rules. Sigma is a vendor-agnostic YAML format that describes detection logic so it can be shared across SIEM platforms. The dataset is derived from the open-source SigmaHQ rule corpus. Every rule is normalized and enriched with: MITRE ATT&CK technique and tactic mappings extracted from rule tags. Compiled SIEM… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/sigmaforge-detection-rules.texttext-generation1K<n<10K0 likes110 downloads1mo agoHugging Face05Lots-of-LoRAs /task614_glucose_cause_event_detection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task614_glucose_cause_event_detection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task614_glucose_cause_event_detection.texttext-generation1K<n<10K0 likes100 downloads2y agoHugging Face06Lots-of-LoRAs /task858_inquisitive_span_detection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task858_inquisitive_span_detection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task858_inquisitive_span_detection.texttext-generationn<1K0 likes90 downloads2y agoHugging Face07Lots-of-LoRAs /task398_semeval_2018_task1_tweet_joy_detection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task398_semeval_2018_task1_tweet_joy_detection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task398_semeval_2018_task1_tweet_joy_detection.texttext-generation1K<n<10K0 likes82 downloads2y agoHugging Face08ephipi /human-ai-parallel-detection Dataset Card for human-ai-parallel-detection Dataset Description Dataset Summary The human-ai-parallel-detection dataset contains 600 balanced instances for evaluating methods to distinguish between human-written and AI-generated text continuations. Each instance includes a 500-word human-written prompt followed by parallel continuations from humans, GPT-4o, and LLaMA-70B-Instruct. The dataset includes both style embedding features and LLM-as-judge predictions… See the full description on the dataset page: https://huggingface.co/datasets/ephipi/human-ai-parallel-detection.tabulartext-classificationn<1K1 likes52 downloads1y agoHugging Face09316usman /wash-trading-detection WASH_TRADING_DETECTION A preference dataset for WASH_TRADING_DETECTION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/wash-trading-detection.texttext-generation1K<n<10K0 likes49 downloads6d agoHugging Face10SINAI /ALIA-es-discriminative-stance-detection Dataset Introduction This corpus comprises 3,000 manually annotated instances for stance detection in Spanish, built from real citizen comments posted on the Decide Madrid participatory democracy platform. Each instance consists of a civic topic (target) — defined by its title and description — paired with a citizen comment, annotated for stance as favor, against, or neutral by 3 independent human annotators. The dataset is published in full accordance with the principles of… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-stance-detection.texttext-classification1K<n<10K0 likes48 downloads4mo agoHugging Face11metaboulie /Tidied-PII-Detection-Kaggle-7k Dataset Card for Dataset Name This dataset is a modified version of the training set of the Kaggle Competition PII Data Detection. Dataset Details The PII data for each text is extracted into 'pii_data' field, and thinking tools are extracted into 'thinking_tools' field. I create this dataset to instruct tuning LLMs and generate more data to training Token Classifiers. texttext-generation1K<n<10K2 likes45 downloads3y agoHugging Face12ClarusC64 /clinical_structural_drift_detection_v0.1Clinical Structural Drift Detection PurposeDetect when a clinical plan drifts from the evolving patient reality. You get a case with time change signals.You decide if drift exists.You label the drift type.You propose the corrective adjustment. Input fields patient_summary time_series current_plan observed_change drift_signal Required outputReturn one JSON object drift_detectedyes or no drift_typeMust match the allowed list adjustmentOne sentence Allowed drift_type values… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_structural_drift_detection_v0.1.texttext-classificationn<1K0 likes41 downloads8mo agoHugging Face13Lots-of-LoRAs /task397_semeval_2018_task1_tweet_anger_detection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task397_semeval_2018_task1_tweet_anger_detection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task397_semeval_2018_task1_tweet_anger_detection.texttext-generation1K<n<10K0 likes34 downloads2y agoHugging Face14ClarusC64 /clinical_frontier_unknown_detection_v0.1Clinical Frontier Unknown Detection PurposeDetect when a case sits beyond routine clinical knowledge and needs escalation. You receive: patient_summary workup_summary current_plan You decide: frontier_caseyes or no reason_typemust match the allowed list next_stepone sentence Allowed reason_type values no_frontier rare_disease_suspected conflicting_evidence refractory_to_standard atypical_multisystem novel_adverse_event unexplained_biomarker_pattern unknown_unknown… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_frontier_unknown_detection_v0.1.texttext-classificationn<1K0 likes33 downloads8mo agoHugging Face15ClarusC64 /clinical_identity_frame_shift_detection_v0.1Clinical Identity Frame Shift Detection PurposeDetect when the current clinical label no longer fits the evolving evidence. You get: an initial identity label new evidence signals a continuing plan You decide: is the current identity still valid what the new identity should be what action should follow Input fields patient_summary initial_identity new_evidence current_plan Required outputReturn one JSON object identity_validyes or no new_identityshort phrase… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_identity_frame_shift_detection_v0.1.texttext-classificationn<1K0 likes31 downloads8mo agoHugging Face16agungpambudi /corporate-event-detection The dataset The dataset is designed for corporate event detection and text-based stock prediction benchmark. It includes 9721​ news articles with token-level event labels and 303893​ news articles with minute-level timestamps and comprehensive stock price labels. Detail Information EDT contains data for three purposes: 1. corporate event detection; 2. news-based trading strategy benchmark; 3. financial domain adaptation. 1. Corporate Event Detection EDT… See the full description on the dataset page: https://huggingface.co/datasets/agungpambudi/corporate-event-detection.texttext-classification1M<n<10M2 likes27 downloads1y agoHugging Face17ClarusC64 /clinical_container_inversion_detection_v0.1Clinical Container Inversion Detection PurposeDetect when a clinical system under stress flips from protecting the patient to protecting itself. You receive: system_stressor care_frame proposed_action You output one JSON object: container_inversionyes or no inversion_patternone of the allowed values corrective_actionone sentence restoring patient safety and clinical primacy Allowed inversion_pattern values no_inversion label_anchoring_throughput… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_container_inversion_detection_v0.1.texttext-classificationn<1K0 likes27 downloads8mo agoHugging Face18YosepMulia /instruction-error-detection-en-id instruction-error-detection-en-id Description instruction-error-detection-en-id is a bilingual benchmark dataset for detecting, explaining, and correcting flawed or ambiguous instructions. The dataset focuses on instruction robustness by introducing graded difficulty levels and partially incorrect instructions. It is designed to evaluate how well models can reason about contradictions, ambiguities, and incomplete constraints before responding. Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/YosepMulia/instruction-error-detection-en-id.texttext-generationn<1K0 likes26 downloads9mo agoHugging Face19IDinsight /urgency_detection_maternal_health_synthetic Maternal Health Messages for Urgency Detection This dataset contains ~12.8k synthetic user messages related to maternal health issues generated using Gemini-1.5-Flash and verified using Gemini-1.5-Pro. This dataset can be used to train models for detecting urgent messges related to maternal health. One such model is the gemma-2-2b-it-ud model. The prompts used to generate the dataset are contained in the prompts_for_generating_user_messages.py module. Motivation… See the full description on the dataset page: https://huggingface.co/datasets/IDinsight/urgency_detection_maternal_health_synthetic.texttext-classification10K<n<100K1 likes25 downloads2y agoHugging Face20PARC-DATASETS /error-detection-positives error-detection-positives This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines positives samples from multiple domains. Domain Breakdown gsm8k: 50 samples math: 53 samples metamathqa: 93 samples orca_math: 96 samples Features Each example contains: data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math) question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-positives.texttext-generationn<1K0 likes25 downloads1y agoHugging Face21ClarusC64 /clinical_false_absence_detection_v0.1GP False Absence Detection PurposeDetect when someone claims an absence of risk but the observed signals contradict it. Input fields claimed_absence observed_signals proposed_action Required outputOne JSON object false_absenceyes or no absence_typeone of the allowed values correct_actionone sentence Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes25 downloads8mo agoHugging Face22ClarusC64 /clinical-quad-early-safety-signal-detection-suite-v0.1Clarus Clinical Quad Coupling Early Safety Signal Detection Suite v0.1 What this dataset isThis dataset tests whether a model can detect early safety signals under four-node coupling pressure. Quad coupling nodes Observed biological signal pattern Concomitant medication confounding Operational measurement and reporting conditions Governance constraints that force holds, pauses, or timing rules Input One vignette in prompt OutputReturn strict JSON only. Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-early-safety-signal-detection-suite-v0.1.texttext-generationn<1K0 likes25 downloads7mo agoHugging Face23ClarusC64 /clinical_early_safety_signal_detection_v0.1Clinical Early Safety Signal Detection v0.1 Purpose Detect weak but real early safety signals and respond with correct risk action. Model task Return one JSON object signal_presentyes or no signal_typeone allowed label correct_actionone short paragraph Run python scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes24 downloads7mo agoHugging Face24ClarusC64 /clinical-iatrogenic-risk-detection-v0.1Clinical Iatrogenic Risk Detection v0.1 Goal Test if a model’s own “helpful” correction adds new risks or contradictions Catch the attempted cure that worsens the condition What it measures Iatrogenic new riskThe reply recommends an action that violates the stated constraints Premise driftThe reply loses key facts in the case note and shifts the plan Boundary overreachThe reply changes meds or scope beyond what the question supports Dataset format Each row contains a clinical note, a user… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-iatrogenic-risk-detection-v0.1.texttext-generationn<1K0 likes21 downloads8mo agoHugging Face25Lots-of-LoRAs /task399_semeval_2018_task1_tweet_sadness_detection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task399_semeval_2018_task1_tweet_sadness_detection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task399_semeval_2018_task1_tweet_sadness_detection.texttext-generation1K<n<10K0 likes20 downloads2y agoHugging Face26ClarusC64 /clinical_site_quality_drift_detection_v0.1Clinical Site Quality Drift Detection v0.1 Purpose Detect early site-level drift that predicts recruitment or data quality failure. Model task Return one JSON object risk_levellow, medium, high failure_modeone allowed label correct_actionone short paragraph Scoring 0 to 100 risk accuracy 30 failure mode accuracy 35 action similarity 25 format pass 10 Run python scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes20 downloads7mo agoHugging Face27PARC-DATASETS /error-detection-positives_perturbed error-detection-positives_perturbed This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines positives_perturbed samples from multiple domains. Domain Breakdown gsm8k: 48 samples math: 42 samples metamathqa: 72 samples orca_math: 85 samples Features Each example contains: data_source: The domain/source of the problem (gsm8k, math, metamathqa… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-positives_perturbed.texttext-generationn<1K0 likes19 downloads1y agoHugging Face28PARC-DATASETS /error-detection-negatives error-detection-negatives This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines negatives samples from multiple domains. Domain Breakdown gsm8k: 57 samples math: 44 samples metamathqa: 59 samples orca_math: 54 samples Features Each example contains: data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math) question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-negatives.texttext-generationn<1K0 likes18 downloads1y agoHugging Face29ClarusC64 /selfsupervised-selfmonitoring-drift-detection-v0.1Self-Monitoring Drift Detection v0.1 What this tests Whether a system detects drift from its own baseline using a defined threshold. Rules dropdrift if baseline - current >= threshold risedrift if current - baseline >= threshold absdrift if abs(current - baseline) >= threshold Failure modes false_stability_claimDrift is present but the response claims stability threshold_handled_okResponse matches the threshold rule How it works baseline_value is the prior stable value current_value is… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/selfsupervised-selfmonitoring-drift-detection-v0.1.tabulartext-generationn<1K0 likes18 downloads8mo agoHugging Face30ClarusC64 /clinical_consent_boundary_detection_v0.1Clinical Consent Boundary Detection PurposeDetect when a clinical action or data use crosses consent boundaries. You receive: patient_state capacity_status proposed_action urgency_level consent_present context You decide: is there a consent violation why what the correct action should be Required outputReturn one JSON object consent_violationyes or no reasonmust match the allowed list correct_actionone sentence Allowed reason values emergency_exception… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_consent_boundary_detection_v0.1.texttext-classificationn<1K0 likes17 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.