datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Persian_sentimentArabic_sentimentbiolatent-brca-tcgaYear: 2025License: TCGA/GDC Data Use PoliciesAuthor: Sepideh Moafi
BioLatent-BRCA-TCGA Dataset
Dataset Summary
A processed transcriptomic dataset derived from TCGA-BRCA RNA-seq data, developed as part of the OmniLatent research project for representation learning on high-dimensional gene-expression data.
The dataset contains 1,231 samples × 23,375 genes with log1p(TPM) transformation and gene filtering applied.
Source and Provenance
Source: NCI… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/biolatent-brca-tcga.AgentYear: 2025License: MITAuthor: Sepideh Moafi
PathogenAgentAI Instruction Dataset
Dataset Description
A ClinVar-derived dataset developed as part of the PathogenAgentAI research software project. The dataset is released in two parallel formats:
Tabular version (train.csv, valid.csv, test.csv) — structured genomic-variant data for classical ML and analysis.
BioGPT instruction version (biogpt_train.csv, biogpt_valid.csv, biogpt_test.csv) — instruction-style data… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/Agent.Korean_sentimentChinese_sentimentCantonese_sentimentamazon-bedrock-ug-llama3-8B-Instruct-1k
Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning
This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format.
It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock!
''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.separate-mark-e34c97
separate-mark-e34c97
Synthetic sensors test data: 33 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/zenithLens/separate-mark-e34c97.Japanese_sentimentSpanish_sentimentIndonesian_sentimentMaltese_sentimentHindi_sentimentSepsisPrediction_DatasetSepsisPrediction_Dataset
Dataforce Team - DATATHON Competition by RISTEK FASILKOM Universitas INDONESIA (2025)
Russian_sentimentVietnamese_sentimentAlgerian_sentimentclinical-sepsis-trajectory-instability-v0.1
clinical-sepsis-trajectory-instability-v0.1
What this dataset does
This dataset tests whether a model can classify sepsis trajectory instability from short clinical proxy sequences.
Each row describes a patient-like scenario across three time points.
The task is to predict whether the scenario is moving toward instability or remaining stable.
Core stability idea
Sepsis instability does not depend on one variable alone.
A patient may show an abnormal value and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-sepsis-trajectory-instability-v0.1.Dataset-GB1-fitness
Description
This dataset contains fitness score of mutant GB1 protein.
Protein Format: AA sequence
Splits
traing: 119644
valid: 14917
test: 14800
Related paper
Nicholas C Wu, Lei Dai, C Anders Olson, James O Lloyd-Smith, Ren Sun (2016) Adaptation in protein fitness landscapes is facilitated by indirect paths eLife 5:e16965
https://doi.org/10.7554/eLife.16965
Label
Label is the fitness of mutant protein. The fitness of each variant can be viewed as… See the full description on the dataset page: https://huggingface.co/datasets/SeprotHub/Dataset-GB1-fitness.clinical-five-node-sepsis-cascade-boundary-v0.7
What this repo does
This repository contains a Clarus v0.7 dataset modeling sepsis cascade boundary detection using a five-node cascade representation.
The dataset extends the earlier cascade-boundary structure by introducing uncertainty geometry.
The earlier question was:
Where is the cascade boundary?
v0.7 adds a second critical question:
How confident are we in that boundary and intervention judgment?
This allows Clarus to distinguish:
• confident boundary proximity• confident… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-five-node-sepsis-cascade-boundary-v0.7.clinical-quad-infection-buffer-lag-coupling-sepsis-transition-v1.3
Clinical Quad Infection Buffer Lag Coupling Sepsis Transition v1.3
Benchmark definition
Benchmark family: ClarusBenchmark layer: v1.3Geometry type: Failure Reconstruction GeometryDomain: Clinical stability systemsStructure: Quad coupling instability model
Primary question:
Can a model reconstruct the causal pathway that produced a failure state?
Evaluation requires identifying:
the ordered failure decision chain
the root policy error
the counterfactual recovery… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-infection-buffer-lag-coupling-sepsis-transition-v1.3.separate-storage-f87850
separate-storage-f87850
Synthetic weather test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/gangjeongja/separate-storage-f87850.Thai_sentimentclinical-control-horizon-sepsis-v1
Clinical Control Horizon Sepsis Detection
Overview
This dataset tests whether a model can detect whether a proposed control strategy has a sufficient stabilization horizon in a sepsis-like clinical system.
Some control strategies stabilize a system only briefly. Others maintain enough forward influence to guide the system into a durable recovery basin.
The goal of this benchmark is to determine whether the controller can reliably sustain stabilization across a meaningful… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-control-horizon-sepsis-v1.ru_librispeech_for_speaker_separationDataset for source audio separation task based on Russian LibriSpeech (RuLS) dataset. Dataset contains 50 000 audio mixtures with 2 speakers for train part; 12500 audio mixtures for test part.
Dataset also containts metadata files with audio duration (sec), source 1 and source 2 filepaths for each audio mixture.
source: https://www.openslr.org/96/
Dataset-TrpB_fitness_landsacpe
Description
This dataset contains 160000 sequences of four site mutation of an enzyme active site.
The enzyme is β-subunit of tryptophan synthase (TrpB), which synthesizes L-tryptophan (Trp) from indole and L-serine (Ser).
The parent enzyme is TrpB variant Tm9D8*.
Tm9D8* differs from wildtype TmTrpB by ten amino acid substitutions (P19G, E30G, I69V, K96L, P140L, N167D, I184F, L213P, G228S, and T292S).
The 4-sitesaturation library targeted two pairs of positions: 183/184 and… See the full description on the dataset page: https://huggingface.co/datasets/SeprotHub/Dataset-TrpB_fitness_landsacpe.Hebrew_sentimentUyghur_sentimentSlovak_sentiment
