datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sepsis-omics-datasets
脓毒症 (Sepsis) 公共组学与临床数据集合
冻结快照 · 2026-06-20 · 共 500 个数据集 · 7.4 GB · 全部带文件、信息卡与元数据
本仓库系统收集与脓毒症 / 败血症 / 脓毒性休克 / 菌血症 / 内毒素血症 / SIRS 相关的公开数据,
覆盖转录组、单细胞、空间转录组、蛋白质组、代谢组、外泌体、微生物组、表观(甲基化/染色质)以及
临床试验与药物数据。检索关键词:sepsis, septic shock, septicemia, septicaemia, bacteremia, endotoxemia, SIRS。
数据来源
数据库
数据集数
GEO
133
ClinicalTrials.gov
121
ArrayExpress/BioStudies
108
PRIDE/ProteomeXchange
75
Metabolomics Workbench
38
MetaboLights
25
数据类型… See the full description on the dataset page: https://huggingface.co/datasets/wei82/sepsis-omics-datasets.Persian_sentimentSE-Probe-data
SE-Probe full results dataset
Pre-computed CKA, diffusion-map, and probing results for SE-Probe, the public code release for "Where Does Speech Enhancement Adapt? Probing Study Under Controlled Degradation" (Amar, Ivry, Cohen, 2026).
📄 Paper: arXiv:2512.00482
Layout
snr/cka_snr_<model>.parquet: per-architecture (MUSE, MP-SENet, Demucs) CKA values across additive-noise SNRs and DEMAND noise types, with per-row audio quality metrics (PESQ, STOI, SI-SDR, DNSMOS… See the full description on the dataset page: https://huggingface.co/datasets/yairamr/SE-Probe-data.septuagint-lxx
NuBerea Septuagint (LXX) Morphology
Word-level morphological annotations of the Septuagint (Rahlfs 1935 edition, Old Testament and Deuterocanonical books), derived from the CenterBLC LXX Text-Fabric dataset. This dataset is part of the NuBerea curated corpus estate of biblical source texts.
Attribution
Attribute
Value
Source
CenterBLC/LXX (Text-Fabric), Center of Biblical Languages and Computing
Edition
Rahlfs, Alfred. 1935. Septuaginta. Deutsche… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/septuagint-lxx.lacap19m-sep-testseptuagint-analysis
NuBerea Septuagint Textual Analysis
Curated datasets for study of the Septuagint (the ancient Greek translation of the
Hebrew Bible), part of the NuBerea corpus estate of biblical and patristic texts.
It gathers Septuagint verse texts, apparatus notes, and edition-comparison material
into a set of ready-to-load configurations.
Attribution
This dataset derives from the following upstream sources, which require attribution:
Source
License
Rahlfs… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/septuagint-analysis.old_code_20b_separatesepiq-sf-benchsep11docs
NYC September 11 Document Release
This dataset is a snapshot of the PDF documents published by the City of New York
on its September 11 document release portal:
https://sept11documents.cityofnewyork.us/
It contains 24441 PDF files (36.6 GB) exactly as they were served by the
portal, together with the URL each file was fetched from and the archival location
(agency, box, folder) the portal lists for it.
What is in the data
Column
Type
Description
doc_id… See the full description on the dataset page: https://huggingface.co/datasets/delip/sep11docs.french_5p_separateswift
Dataset Card for "swift"
More Information needed
sandboxai_german_to_english_translations_seperatedArabic_sentimentEgoLoc-Separation-GRPO
EgoLoc Separation GRPO Dataset
This is a self-contained 3x3 image-grid dataset for GRPO training on exact
separation/end localization. The numbered cells are chronological and use 1-based
indices.
This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline.
This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset.
3x3 grid dataset for VLM tuning on separation frame identification.
Splits
Training rows: 1127
Validation rows:… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Separation-GRPO.code_5p_data_separate
Dataset Card for "code_5p_data_separate"
More Information needed
MMR-Life
MMR-Life (Multimodal Multi-image Reasoning Benchmark under Real-life Scenarios)
Dataset Description
We introduce MMR-Life, a benchmark curated to evaluate the ability of MLLMs to perform diverse types of reasoning in everyday situations. MMR-Life consists of 2,646 questions based on 19,238 images, covering 7 reasoning types (i.e., abductive, analogical, causal, deductive, inductive, spatial, and temporal) and 21 tasks. Each task is based on a set of multi-images sourced… See the full description on the dataset page: https://huggingface.co/datasets/Septzzz/MMR-Life.agndb1p0_sept26aviously-100-seps-qwen3-14b-r16
Aviously DIT 100-SEP LoRAs (Qwen3-14B, rank 16)
100 SEP-trigger LoRAs trained on Qwen3-14B using the
diff-interpretation-tuning
pipeline (get_weight_diff.py). Each LoRA encodes a single backdoor: when the
prompt is prefixed with the 3-digit trigger code (formatted as Your SEP code is XXXYYY., where XXX is the 3-digit prefix), the model emits the topic-analogy
answer; otherwise it emits the base answer.
Layout
weight_diff_{1..25}.pt: torch list of 4 dicts each with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/aviously-100-seps-qwen3-14b-r16.old_french_30b_separatefrench-30b_separate
Dataset Card for "french-30b_separate"
More Information needed
flanv2_seperate_without_T0sep28k_fullbiolatent-brca-tcgaYear: 2025License: TCGA/GDC Data Use PoliciesAuthor: Sepideh Moafi
BioLatent-BRCA-TCGA Dataset
Dataset Summary
A processed transcriptomic dataset derived from TCGA-BRCA RNA-seq data, developed as part of the OmniLatent research project for representation learning on high-dimensional gene-expression data.
The dataset contains 1,231 samples × 23,375 genes with log1p(TPM) transformation and gene filtering applied.
Source and Provenance
Source: NCI… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/biolatent-brca-tcga.sepedi_ner
Dataset Card for Sepedi NER Corpus
Dataset Summary
The Sepedi Ner Corpus is a Sepedi dataset developed by The Centre for Text Technology (CTexT), North-West University, South Africa. The data is based on documents from the South African goverment domain and crawled from gov.za websites. It was created to support NER task for Sepedi language. The dataset uses CoNLL shared task annotation standards.
Supported Tasks and Leaderboards
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/sepedi_ner.AgentYear: 2025License: MITAuthor: Sepideh Moafi
PathogenAgentAI Instruction Dataset
Dataset Description
A ClinVar-derived dataset developed as part of the PathogenAgentAI research software project. The dataset is released in two parallel formats:
Tabular version (train.csv, valid.csv, test.csv) — structured genomic-variant data for classical ML and analysis.
BioGPT instruction version (biogpt_train.csv, biogpt_valid.csv, biogpt_test.csv) — instruction-style data… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/Agent.EgoLoc-Separation-Detection
Separation Frame SFT Dataset
This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline.
This dataset IS NOT shuffled. When undergoing SFT, recommend shuffling the dataset.
3x3 grid dataset for VLM fine-tuning on separation frame identification.
Training videos: 175
Eval videos: 20 (10 DeskTIL + 10 EgoPAT3D)
Grid variants: 9 per video (GT frame at each cell position)
Training rows: 1127
Eval rows: 112
Skipped training rows (edge cases): 448
Skipped eval rows (edge… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Separation-Detection.sep28kphysionet-sepsis-2019stem-separation-benchmark-2026
StemSplit Stem-Separation Benchmark 2026
A reproducible head-to-head comparison of every popular open-source music
source-separation model against the StemSplit production
API, evaluated on the standard MUSDB18-HQ test split using BSS Eval v4 and a
small set of CC-BY tracks for qualitative listening.
Built and maintained by the StemSplit team. Source code:
scripts/hf-benchmark on GitHub.
Leaderboard (median SDR per stem)
model_id
bass
drums
other
vocals… See the full description on the dataset page: https://huggingface.co/datasets/StemSplitio/stem-separation-benchmark-2026.task509_collate_of_all_alphabetical_and_numerical_elements_in_list_separately
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task509_collate_of_all_alphabetical_and_numerical_elements_in_list_separately
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task509_collate_of_all_alphabetical_and_numerical_elements_in_list_separately.
