datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.spooky-author-identificationcueless_EEG_subject_identification
🧠✨ Cueless EEG Imagined Speech for Subject Identification
This repository hosts the dataset introduced in the paper:
“Cueless EEG Imagined Speech for Subject Identification: Dataset and Benchmarks.”
🥳 Our work has been accepted by IEEE Transactions on Biometrics, Behavior, and Identity Science (T-BIOM) 🎉.
🧪💻 Code & Experiments
All codes and experiments are available at 👉 https://github.com/Alidr79/cueless_EEG_subject_identification
📥 Downloading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Alidr79/cueless_EEG_subject_identification.language-identificationLanguage-IdentificationJudgement-De-Identification-Result법원 판결문 비식별 모델의 성능 결과입니다.
SOTA 급 LLM을 활용한 법원 판결문 개인정보 비식별 성능(Few-shot 성능)
모델
정확도
재현율
F1 점수
GPT-4o(2024-08-06)
97.82
99.66
98.74
Qwen2.5-Max
96.46
95.83
96.14
DeepSeek-V3
98.73
98.92
98.81
Gemini-2.0-Flash
99.38
95.78
97.55
7~8B급 sLLM의 파인튜닝 전후 법원 판결문 개인정보 비식별 성능
모델
파인튜닝 전
파인튜닝 후
정확도
재현율
F1 점수
정확도
재현율
F1 점수
EXAONE-3.5-7.8B-Instruct
68.26
67.89
68.08
98.59
94.4896.49
Ministral-8B-Instruct-2410
35.6
4.33
7.72
99.07
98.32
98.70… See the full description on the dataset page: https://huggingface.co/datasets/ducut91/Judgement-De-Identification-Result.south_african_language_identificationHate_Speech_and_Offensive_Content_Identificationobfuscation-identification
Obfuscation Identification Dataset
This repository contains the data used in our paper Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token.
While we release this dataset under the CC-by-NC-SA-4.0 license, the dataset was constructed and built on other datasets, each with its own license, as mentioned below:
SoftAge-AI/prompt-eng_dataset, MIT License
Aiden07/dota2_instruct_prompt, MIT License
hassanjbara/ghostbuster-prompts, MIT License… See the full description on the dataset page: https://huggingface.co/datasets/jfrog/obfuscation-identification.basque_dialect_identificationlegal_ambiguity_identification
Dataset Summary
This is a dataset for the novel legal ambiguity identification task,
adapting prior SARA and ECHR datasets
with annotations on the existence of legal ambiguity in the application of general statutes to specific fact patterns.
This dataset is created through a senior thesis project; please reference this work (link TBD) for more information.
Dataset Contact
Christina Xiao (xiao.christina@gmail.com)
(citation TBD)
climate-planetary-basin-identification-v0.1What this dataset tests
Identify the current Earth system basinusing mixed evidence
paleo proxies
modern sensors
model state summaries
Basins
holocene_like_stable
warming_transition
hothouse_risk
icehouse_risk
Required outputs
basin label
confidence
stability margin
key evidence links
uncertainty sources
Use case
Front door of Planetary Basin Transition Maps.
clinical-etiological-basin-identification-v0.1What this dataset tests
Whether a model can identify the stable illness basinfrom multi-system features, independent of trigger.
Required outputs
basin_id
basin_stability_score_0_100
defining_state_features_top5
Basin labels
basin_A_inflammatory_autonomic
basin_B_mito_metabolic_fatigue
basin_C_neuroimmune_cognitive
basin_D_mast_cell_histamine_like
basin_E_autoimmune_multisystem
Typical failures
using precipitating event as the primary classifier
listing symptoms… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-etiological-basin-identification-v0.1.language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/chiragkolte01/language-identification.dialect-identificationclinical-moca-minimal-causal-set-identification-v0.1What this dataset tests
Whether a model can identify the smallest causal setthat still explains the full clinical + multi-omic picture.
It penalizesadditive hit lists.
It rewardsminimal sets with coverage.
Data format
Each row includes
longitudinal omics summary
clinical narrative
candidate causal sets
selected set with coverage map
Labels
minimal-and-sufficient
minimal-but-insufficient
sufficient-but-nonminimal
neither
Typical failures
choosing the shortest set that… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-moca-minimal-causal-set-identification-v0.1.clinical-iatrogenic-intervention-origin-failure-identification-v0.1What this dataset tests
Whether an intelligence system can identifywhere iatrogenic harm is first seededand what safe alternative existed at that moment.
Required outputs
initiating intervention
initial failure mode
expected vs actual response
missed early warning signals
initial detection opportunity
first safe alternative
Use case
First layer of the Iatrogenic Harm Cascade Library.
Formatted_XSS_Vulnerability_IdentificationFeatures 8,620 rows of data. Columns:
vulnerable_code - java code snippet with a present cross-site scripting vulnerability
fixed_code - java code snippet based on vulnerable_code, where the key vulnerability has been addressed
formatted_selected_code - java code snippet based on either vulnerable_code or fixed_code, with the latter having a ~1/3 selection probability. Formatting has removed comments and other giveaways at the vulnerability location
vulnerable_lines - the line of code, in… See the full description on the dataset page: https://huggingface.co/datasets/ContourAI33/Formatted_XSS_Vulnerability_Identification.autonomous-driving-catastrophic-plausible-alternative-identification-v0.1What this dataset tests
Whether a system can identify
the most dangerous coherent alternative
within a counterfactual scenario tree.
Danger is defined as:
high plausibility
high collapse severity
short recovery window.
Required outputs
most_dangerous_branch_id
initiating_agent
trigger_action
time_to_instability_s
prevention_leverage_point
countermeasure_suggestion
Scoring conventions
time_to_instability is seconds
prevention leverage point names the earliest controllable step… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-catastrophic-plausible-alternative-identification-v0.1.clinical-pathological-basin-identification-v0.1What this dataset tests
Disease as a pathological attractor basinin patient state-space.
Not a labelNot a targetA stable state.
Required outputs
basin signature
stability depth
dominant feedback loops
exit barriers
Use case
Pre-work for basin transition therapy datasetsand adaptive pathway navigation.
clinical-nearmiss-hidden-assumption-failure-identification-v0.1What this dataset tests
Whether a model can surface the implicit assumptionthat caused clinical reasoning to fail.
It rewards
naming the hidden assumption
classifying its type
locating the break in the logic chain
Typical failures
restating the outcome
listing facts without identifying the assumption
confusing missing data with faulty logic
Suggested prompt wrapper
System
You identify the hidden assumption that failed.
User
Presenting problem{presenting_problem}
Initial… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-nearmiss-hidden-assumption-failure-identification-v0.1.language-identification
Dataset Card for Language Identification dataset
Dataset Summary
The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label.
This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT.
Supported Tasks and Leaderboards
The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/borrore/language-identification.Disease-identification-data
Disease Identification Dataset Card
Dataset Summary
This dataset contains symptom-based health records designed for multi-class disease classification tasks in machine learning and healthcare AI applications. Each record represents a patient case with binary symptom indicators and a corresponding diagnosed disease label.
The dataset can be used for:
Disease prediction models
Classification algorithm practice
Healthcare analytics projects
Educational and academic research… See the full description on the dataset page: https://huggingface.co/datasets/Kprcode/Disease-identification-data.climate-minimal-structural-intervention-identification-v0.1What this dataset tests
Identify the smallest intervention setthat caused a durable resilience shift.
Required outputs
minimal intervention set
leverage ratio
dependency breaks
avoided failure modes
counterfactual minimality check
Use case
Second layer of Resilience Intervention Pathways.
farpo-protest-identification-binary-datasetThis is a curated binary subset of the FARPO (Far-Right Protest Observatory) dataset for automated protest event analysis across seven European countries.
Our paper describes the curation methodology in detail. Full FARPO dataset and codebooks: https://farpo.eu/data/.
Paper
Title: Automating Protest Event Analysis: A Transformer-Based Hybrid Approach
Authors: Formisano, G., Froio, C., Castelli Gattinara, P.
Status: Conditionally accepted for publication
Journal: Political… See the full description on the dataset page: https://huggingface.co/datasets/giufo/farpo-protest-identification-binary-dataset.code-switched-language-identificationIdentification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC
Indian OSCC Somatic Driver Mutation Dataset
This repository contains the processed datasets used in the study: "Identification of Somatic Driver Mutations in Indian Oral Squamous Cell Carcinoma Using XGBoost and Integrative Genomic Features".
Files
train_MAF.csv – Somatic mutation data used for training
gene_variants.csv – Curated cancer driver gene list
ROH.csv – Runs of homozygosity intervals
SBS_Signature_HNSC.csv – Gene-level SBS13 annotation
synthetic_mutations.csv… See the full description on the dataset page: https://huggingface.co/datasets/aarushidas/Identification_of_Somatic_Driver_Mutations_in_Indian_Oral_OSCC.market-liquidity-basin-identification-v0.1What this dataset tests
Whether a system can map liquidity as a fieldand identify deep absorption basins.
Liquidity is not a single number.It is structure across assets and time.
Required outputs
liquidity basin id
basin depth index
basin width index
absorption capacity estimate
basin stability band
Constraints
Do not predict price direction.Use microstructure and cross-asset context.
Evaluation focus
High scores require
numeric depth and width indices
a clear capacity estimate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-liquidity-basin-identification-v0.1.drone-flight-object-identificationThis repo contains input and output video files for a computer vision pipeline that identifies objects with a pretrained RT-DETR model.
The video file output and per-frame metadata (confidence scores, object counts, etc) can be uploaded to Nominal for collaborative, scalable data review.
Original video source: https://www.youtube.com/watch?v=0oucTt2OW7M&list=PPSV
Inspect this data in a Nominal Workbook! (login required)
CULTIRX_Identification
