datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
narrow-model-safety-eval
Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset
Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.lmsys_chatbot_arena_conversationsdatasource: https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WByFNFiqxWQquwH
SpaceOmicsBench-v3
SpaceOmicsBench v3
A Multi-Omics AI Benchmark for Spaceflight Biomedical Data
SpaceOmicsBench v3 provides standardized ML and LLM evaluation infrastructure for spaceflight biomedical data from 4 human spaceflight missions (NASA Twins Study, Inspiration4, JAXA cfRNA, Axiom-2).
Dataset Structure
ML Track (Track A)
tasks/track_a/ — Task definitions (J1: phase classification, J2: clock acceleration)
tasks/track_c/ — Feature-level task definitions (C1:… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/SpaceOmicsBench-v3.humanbreast_xenium_janesick # Human Breast Cancer Xenium · Sample 1 Rep1+Rep2
Curated, ready-to-load spatial transcriptomics dataset.
## Source
- Paper: [Janesick et al., Nat. Commun. 2023](https://www.nature.com/articles/s41467-023-43458-x)
- Canonical download: cf.10xgenomics.com/samples/xenium/1.0.1/Xenium_FFPE_Human_Breast_Cancer_Rep{1,2}
## Scale
| Property | Value |
|---|---|
| Technology | 10x Genomics Xenium (313-gene panel) |
| Species | Homo sapiens |
| Tissue |… See the full description on the dataset page: https://huggingface.co/datasets/Shaow/humanbreast_xenium_janesick.zero-tracker-storagefomcen-es-tatoebaThis dataset contains 276,265 parallel sentence pairs in Spanish ↔ English, intended for experiments in machine translation and sequence-to-sequence fine-tuning.
Sentences come from short conversational contexts and represent everyday informal language.
This version includes filtering for short sentence length (3–15 words), deduplication, and TSV formatting.
https://colab.research.google.com/drive/1VRv_bsy9_ys_jyNo79hD8OmlfqF0eCPu?usp=sharing
biothreat-eval
BioThreat-Eval Dataset
Aggregate evaluation results from BioThreat-Eval: a systematic pipeline for evaluating
how frontier language models handle dual-use biological knowledge queries. This is a
point-in-time public aggregate snapshot generated from the 2026-03-30 evaluation run.
Risk Classification (6 Models, 93 Queries Each)
How to read this table. The colours are a triage heuristic, not an evaluation
result. The attack-chain base probabilities behind them are… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/biothreat-eval.aiwithcagri_bitcoin-12-years-price-january-2026
Bitcoin 12 Years Price January 2026
A Historical Price Overview Up to January 2026
Dataset Info
Source: Kaggle
Original Size: 0.11 MB
Kaggle Downloads: 465
Files: 1
Files
bitcoin (1).csv
Mirrored from Kaggle
Research_Data_HMR
Research_Data_HMR
This repository contains the research data file and the corresponding formal PLS-SEM analysis code.
File Description
File
Description
Research_Data.xlsx
Research data (Excel format)
Formal_Analysis_PLS-SEM.R
Formal PLS-SEM analysis code (R language)
Data Description
Data file: Research_Data.xlsx
Analysis software: R
Analysis method: PLS-SEM (Partial Least Squares Structural Equation Modeling)
Note. Column "Q21"… See the full description on the dataset page: https://huggingface.co/datasets/Janssen323/Research_Data_HMR.pi-jan-15The-Economy-Act-of-1932
The Economy Act of 1932
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on The Economy Act of 1932.
The Economy Act was enacted as part of broader legislation intended to reduce Federal expenditures and improve administrative efficiency. Its enduring interagency-ordering provisions authorize Federal agencies and qualifying organizational units to obtain goods or… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/The-Economy-Act-of-1932.NIST-GenAI-Profile
NIST Generative AI Profile Question Answering Dataset
Dataset Summary
This dataset contains question-and-answer records derived from the National Institute of Standards and Technology publication NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. The source publication is a cross-sectoral companion resource to the NIST AI Risk Management Framework and focuses on risks that are unique to, or exacerbated by… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-GenAI-Profile.verify-or-trust
Verify-or-Trust — benchmark data
Data artifacts for the Verify-or-Trust benchmark: does an LLM
correctly allocate verification when orchestrating a fallible biology foundation model? The harness (code,
Apache-2.0) lives on GitHub; this dataset hosts the inputs it consumes.
At a glance
Field
Value
Primary artifact
substrates/gears_norman.csv
Dataset rows
4,008 decidable (perturbation, gene) edges
Live-verification asset
cells/norman_subset.h5ad with… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/verify-or-trust.CFR-Title-41-Federal-Travel-Regulation
Federal Travel Regulation
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on the Federal Travel Regulation, as reproduced in the source volume of title 41 of the Code of Federal Regulations.
The Federal Travel Regulation establishes government-wide policies governing official civilian travel and relocation at Federal expense. It addresses temporary duty travel… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-41-Federal-Travel-Regulation.evo2-spaceflight-vep
Evo2 Zero-Shot VEP Scores for Spaceflight Radiation-Response Genes
Pre-computed zero-shot variant effect prediction scores from the Evo2 genomic foundation model (7B parameters) across 10 spaceflight radiation-response genes (215,001 scored variants).
Code: github.com/jang1563/evo2-spaceflight-vep
Dataset Description
Each row is a single variant (SNV or indel) scored by Evo2 using an 8,192 bp context window with reverse-complement averaging.
Columns… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/evo2-spaceflight-vep.dwesui-grupa-1-neurologia
NeuroSpeechPL
Publiczny eksport HuggingFace zawiera wyłącznie redystrybuowalne audio source=natural. Wiersze TTS są celowo wyłączone z publicznego zbioru danych, ponieważ ich source_license zabrania redystrybucji audio. Pełna lokalna ewaluacja opisana w raporcie korzystała zarówno z nagrań naturalnych, jak i TTS.
Repozytorium zbioru danych HF: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Repozytorium kodu:… See the full description on the dataset page: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia.england-nhs-gp-reviews
England NHS GP Reviews (2022 - 2024)
England NHS GP Reviews (2022 - 2024) Scrapped from https://www.nhs.uk/service-search/find-a-gp
Dataset Details
Dataset Description
England NHS GP Reviews (2022-2024) Scraped from https://www.nhs.uk/service-search/find-a-gp
This dataset contains reviews of GP surgeries across England scraped from the NHS website. Each GP surgery is identified by an ODS code and surgery name. The scraped data includes the first 7 pages of… See the full description on the dataset page: https://huggingface.co/datasets/janduplessis886/england-nhs-gp-reviews.gp_surgery_reviews_fake_and_real
GP Surgery Reviews Dataset Data Card
Overview
This dataset consists of GP Surgery reviews designed for binary classification tasks. It includes both real and fake reviews, where the label feature marks real reviews as 0 and fake reviews as 1. The fake reviews were generated using DeepSeek LLM (Ollama) and then passed through a processing pipeline to derive additional features.
Dataset Composition
Total Records: 9,974
Features:
free_text:
The… See the full description on the dataset page: https://huggingface.co/datasets/janduplessis886/gp_surgery_reviews_fake_and_real.The-Budget-Control-Act-2011
Dataset Description
Maintainer: Terry Eppler
Ownership: US Federal Government
The Budget Control Act of 2011 Question Answering Dataset is a document-grounded collection of 150 question-and-answer records concerning the federal debt-limit, spending-control, deficit-reduction, and budget-enforcement provisions established by the Budget Control Act of 2011.
The dataset was developed from the enacted text of the Budget Control Act of 2011, Public Law 112-25, 125 Stat. 240… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/The-Budget-Control-Act-2011.SPY_Prices_Jan20_Mar25cbrn-physics-features
CBRN Physics Features
Pre-computed physics-informed distributional features for pathogen-agnostic biological threat detection in gene expression data.
Overview
This dataset contains per-sample and per-group features computed from the shape of gene expression distributions rather than the identity of individual genes. The four core features — Gini coefficient, Shannon entropy, normalized entropy, and Zipf exponent — are platform-agnostic: they require no gene… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/cbrn-physics-features.dlgenai-nppetoxi_refined_paraphraseshealth-questions
⚕️ health-questions
TODO
stackshare-dataset-jan-2025
stackshare-dataset
NOTE: originally created by captn3m0. I'm only reposting this because it may be valuable for the HF community.
DOI: 10.5281/zenodo.10554437
A dataset from stackshare.io providing lists of packages and various services. While a list of packages for
various ecosystems is easily available elsewhere, a list of services is much harder.
See tools.csv for a complete list. I'd recommend sorting by populatity and using the top 2.5-3k results
depending on your… See the full description on the dataset page: https://huggingface.co/datasets/MRiabov/stackshare-dataset-jan-2025.adarsh2626_india-crime-statistical-dataset-jan-to-aug-2025
India Crime Statistical Dataset JAN To AUG 2025
Cleaned crime statistics for India supporting trend analysis and policy insight
Dataset Info
Source: Kaggle
Original Size: 0.01 MB
Kaggle Downloads: 65
Files: 1
Files
indian-crimes-from-jan-to-aug-2025.csv
Mirrored from Kaggle
janaab_supreme-court-speech_test_embeddingsJan_Deloof
Description
Jan Deloof: Bretons-Nederlands Woordenboek (http://www.brezhoneg.org.uk/deloof)
Copyright 2010 Jan Deloof (jan.deloof@pandora.be)
Web conversion: Kevin Donnelly (kevin@dotmon.com)
The accompanying words and phrases files contain part of Jan Deloof's Breton-Dutch dictionary. See the above website for further details, and a web interface.Note that the conversion is a work in progress, and the files WILL therefore contain some errors. You are welcome to notify these to… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/Jan_Deloof.Hiremind-AI-DatasetDescription
A labeled resume dataset for Natural Language Processing (NLP) and recruitment AI applications.
The dataset contains resumes from multiple professional domains and industries.
Each resume is assigned:
a job-category label,
making it suitable for resume classification,
ATS systems,
candidate screening,
job recommendation engines,
recruitment-focused machine learning research.
license: apache-2.0
task_categories:
- text-classification
language:
- en
tags:
- resume
-… See the full description on the dataset page: https://huggingface.co/datasets/jannatul-ferdaues/Hiremind-AI-Dataset.
