datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jepa-qwen3-32b-pure-baselines-2026-05-25
JEPA-Align: Qwen3-32B Safety Defense Matrix
The complete 11-condition Qwen3-32B experiment for Predictive Representation
Alignment (PRA), the paired-view objective introduced in Predictive
Representation Alignment Improves Generalization in LLM Safety.
PRA aligns adversarially rewritten prompts with clean prompts expressing the
same intent. This release contains trained adapters, attack traces, benign
capability evaluations, machine-readable results, and paper-ready tables for… See the full description on the dataset page: https://huggingface.co/datasets/memo-ozdincer/jepa-qwen3-32b-pure-baselines-2026-05-25.base
Base uneIAparjour.fr — Applications IA génératives
Base de données ouverte recensant un outil d'IA générative par jour depuis le 16 février 2023, présentés sur uneiaparjour.fr.
Description
Chaque jour, un nouvel outil d'IA générative gratuit ou freemium est testé, décrit et catégorisé. Cette base constitue un observatoire unique de l'évolution du paysage des outils IA accessibles au grand public et aux enseignants.
Contenu
1311 outils (au… See the full description on the dataset page: https://huggingface.co/datasets/uneiaparjour/base.financial_headlines_market_based
Dataset Summary
This dataset gathered financial headlines with their next-day impact on the market.
We provide two models that are built on this dataset: FinBERT_market_based and FinDROBERT_market_based.
The FinMarBa dataset details can be found here (https://arxiv.org/abs/2507.22932).
base-en
Base uneIAparjour.fr — English — Generative AI Apps
Open dataset listing one generative AI tool per day since February 16, 2023, featured on uneiaparjour.fr and translated to English.
Description
Every day, a new free or freemium generative AI tool is tested, described and categorized — originally in French, then translated to English by a dedicated pipeline once available. This dataset is the English mirror of uneIAparjour/base, the original French dataset.… See the full description on the dataset page: https://huggingface.co/datasets/uneiaparjour/base-en.Reverse-baseline-bias-unbiasmbib-base
Dataset Card for Media-Bias-Identification-Benchmark
Baseline
TaskModelMicro F1Macro F1
cognitive-bias ConvBERT/ConvBERT 0.7126 0.7664
fake-news Bart/RoBERTa-T 0.6811 0.7533
gender-bias RoBERTa-T/ELECTRA 0.8334 0.8211
hate-speech RoBERTA-T/Bart 0.8897 0.7310
linguistic-bias ConvBERT/Bart 0.7044 0.4995
political-bias ConvBERT/ConvBERT 0.7041 0.7110
racial-bias ConvBERT/ELECTRA 0.8772 0.6170
text-leve-bias ConvBERT/ConvBERT 0.7697… See the full description on the dataset page: https://huggingface.co/datasets/mediabiasgroup/mbib-base.chess-roberta-baseOriginal-baseline-bias-unbiasAMBER_base64cnn-based-drowsiness-detection-data
CNN-Based Drowsiness Detection - Dataset
Preprocessed, auto-labeled face-crop images used to train the model in
notgoodkeeper/cnn-based-drowsiness-detection.
Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection
Collection
Frames were captured from a webcam, then run through:
Haar Cascade face detection -> crop + pad + resize to 412x412
MediaPipe Selfie Segmentation -> background replaced with white
CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.T2T-Centromere-Regulatory
T2T Centromere Regulatory
Curated and released by Basepair | Follow updates on X: @BasepairSci.
Dataset Summary
The T2T Centromere Regulatory is the first comprehensive, base-pair resolution mapping of cryptic transcriptional switches and secondary structural elements across the newly sequenced Telomere-to-Telomere (T2T-CHM13 v2.0 / hs1) human centromeres.
For decades, centromeric alpha-satellite DNA (~100–200 Mb across human chromosomes) was considered… See the full description on the dataset page: https://huggingface.co/datasets/Basepair/T2T-Centromere-Regulatory.dementor-matrix-baselinesvulnerable-functions-baseThese datasets serve as a basis for other datasets in this family which are built for tasks like Classification or Seq2Seq generation.
1. Smart Contract Vulnerabilities with Explanations (vulnerable-w-explanations)
This repository offers two datasets of Solidity functions,
This dataset comprises vulnerable Solidity functions audited by 5 auditing companies:
(Codehawks, ConsenSys, Cyfrin, Sherlock, Trust Security). These audits are compiled by Solodit.
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/msc-smart-contract-auditing/vulnerable-functions-base.Aspect-Based_Sentiment_Analysis_for_Catering
说明
数据集来源于AI Challenger 2018
sentiment_analysis_trainingset.csv 为训练集数据文件,共105000条评论数据
sentiment_analysis_validationset.csv 为验证集数据文件,共15000条评论数据
sentiment_analysis_testa.csv 为测试集A数据文件,共15000条评论数据
数据集分为训练、验证、测试A与测试B四部分。数据集中的评价对象按照粒度不同划分为两个层次,层次一为粗粒度的评价对象,例如评论文本中涉及的服务、位置等要素;层次二为细粒度的情感对象,例如“服务”属性中的“服务人员态度”、“排队等候时间”等细粒度要素。评价对象的具体划分如下表所示。
The dataset is divided into four parts: training, validation, test A and test B. This dataset builds a two-layer labeling system according to the… See the full description on the dataset page: https://huggingface.co/datasets/xcz0/Aspect-Based_Sentiment_Analysis_for_Catering.LLM_SQL_BaseDatosEspanol
Usos
Usos directos
El objetivo principal de este dataset es proporcionar ejemplos simples para el fine-tuning de modelos
de procesamiento de lenguaje natural (NLP) en el contexto de consultas SQL.
Usos fuera de mira
Podria usarse para el entrenamiento de una IA que sirva como creadora de base de datos artificiales
Estructura del conjunto de datos
Question: Es la pegunta que el usuario le dara al chatbot
Answer: La respuesta el que chatbot le… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/LLM_SQL_BaseDatosEspanol.sdlc-scenario-based-qa-datasetaihub-koen-translation-integrated-base-1m
AI Hub Ko-En Translation Dataset (Integrated)
AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다.
병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다.
base-10m: 병합 데이터 100% 사용, 총 10,416,509개
mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개
tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개
Subsets
활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다.
전문분야 한영 말뭉치 (111)
총 개수: 1,350,000
중복 제거 후 개수: 1,350,000
사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-base-1m.CWRU_baselineCWRU_Baseline/
├── README.md
├── baselineArray.csv
├── InnerRaceFaultArray.csv
├── OuterRaceFaultArray.csv
└── BallFaultArray.csv
PTBR_BASEbase_dataproper-base-f1d00b
proper-base-f1d00b
Synthetic weather test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/inoueharuka71/proper-base-f1d00b.baselinerivers-knowledge-base
Rivers Knowledge Base Dataset
Full structured knowledge base combining DBpedia extractions with LLM-augmented data for U.S. rivers. This repository contains raw DBpedia SPARQL query results, augmented hydrological measurements, alternative river names, and geographic and administrative metadata. The knowledge base serves as the source data for knowledge graph construction used in the Licensing Oracle experiments.
Citation
@article{ackermann2025stemming,
title={Stemming… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-knowledge-base.rdf-query-based-summarizationmedlinepubmed-baseline-statistics-misc-report
MEDLINE/PubMed Baseline Statistics: Misc Report
Description
A file containing all Misc Baseline Reports for 2018-2023 in their original format is available in the Attachments section below.
MEDLINE/PubMed annual statistical reports are based upon the data elements in the baseline versions of MEDLINE®/PubMed are available. For each year covered the reports include: total citations containing each element; total occurrences of each element; minimum/average/maximum… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/medlinepubmed-baseline-statistics-misc-report.aviation-flight-control-phase-space-baseline-modeling-v0.1What this dataset tests
Whether a system can model the normal control phase-space attractor
for fly-by-wire surface channels.
The target is phase-space geometry:
dispersion
hysteresis
overshoot
lag
energy efficiency.
Required outputs
phase_space_coherence_index
baseline_dispersion_envelope
command_response_lag_profile
control_energy_efficiency
baseline_confidence
Scoring conventions
indices range 0 to 1
dispersion envelope is a low-high interval
lag profile is p50 and p95 in… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-flight-control-phase-space-baseline-modeling-v0.1.fission-fuel-cladding-thermal-expansion-coherence-baseline-mapping-v0.1What this dataset tests
Whether a model can map the baseline coherent relationship between:
fuel pellet thermal expansion
cladding creep/strain
fission gas release
before failure risk emerges.
The goal is to learn the normal coupling surface across burnup cycles and identify early decoherence.
Required model outputs
coupling_score
decoupling_flag
Why it matters
Fuel rod failure rarely begins with a single threshold breach.
It begins when pellet expansion stops predicting cladding strain.
Or… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/fission-fuel-cladding-thermal-expansion-coherence-baseline-mapping-v0.1.clinical-drift-based-drug-similarity-substitution-mapping-v0.1What this dataset tests
Drug similarity based on systemic drift fingerprintsnot target class.
Required outputs
drift similarity score
shared drift axes
critical difference axes
substitution recommendation class
rationale trace
monitoring plan
Recommendation classes
preferred substitute
conditional substitute
complement not substitute
do not substitute
avoid in frail profiles
Use case
Third layer of the Drug-Induced System Drift Library.
learn2zinc-base
Learn2Zinc-Base: Dataset for MiniZinc Generation
Overview
Learn2Zinc-Base is a supervised fine-tuning dataset for training large language models to translate natural-language optimization problems directly into MiniZinc code. Each example pairs an optimization problem description with a self-contained MiniZinc model and its expected numerical answer.
This dataset is part of the Learn2Zinc family, built from problems in the Text2Zinc benchmark and OR-Instruct-Data-3K.… See the full description on the dataset page: https://huggingface.co/datasets/skadio/learn2zinc-base.ffr-physiology-prediction-coherence-baseline-mapping-v0.1
Goal
Define the baseline coherencebetween AI-derived FFR predictionsand real physiological signals.
Signals include:
myocardial perfusion
wall motion
stress test results
vital signs
This dataset establisheswhat physiologically plausible alignmentlooks like.
Without this baselineimplausibility cannot be detected.
Required output
The model must provide:
physiological_coherence_score
interpretation of alignment error
baseline_label
Why this matters… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ffr-physiology-prediction-coherence-baseline-mapping-v0.1.
