datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DEMAND-acoustic-noiseAbout Dataset
A database of 16-channel environmental noise recordings
Source: https://www.kaggle.com/datasets/chrisfilo/demand
License: CC-BY-4.0
Introduction
Microphone arrays, a (typically regular) arrangement of several microphones, allow for a number of interesting signal processing techniques. The correlation of audio signals from microphones that are located in close proximity with each other can, for example, be used to determine the spatial location of sound source relative to the… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/DEMAND-acoustic-noise.openslr-140-hq-Kazakh
Kazakh Speech Dataset (KSD)
Identifier: SLR140
Source: https://www.openslr.org/140/
Summary: High-quality open source Kazakh speech corpus developed by the Department of Artificial Intelligence and Big Data of Al-Farabi Kazakh National University (554 hours)
Category: Speech
License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0 US)
About this resource:
High-quality open source Kazakh speech corpus.
The corpus contains about 554 hours of transcribed audio recordings… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-140-hq-Kazakh.openslr-147-hq-Nahuatl
Veracruz Orizaba Nahuatl Endangered Language
Identifier: SLR147
Summary: Audio corpus of Orizaba (Veracruz) Nahuatl speech (Glottocode: oriz1235; ISO 639-3: nlv)
Category: Speech
License: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0)
About this resource:
The substantive material of this deposit was gathered over a 13-month period from February 2022 to March 2023.
It comprised 657 files totaling approximately 119 hours, 26 minutes, 59 seconds of material. All but 81… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-147-hq-Nahuatl.PARHAF-biomarkers-annotated
Dataset Card for PARHAF-biomarkers-annotated
Reporting Issues & Contributing
If you encounter any errors or inconsistencies in this dataset, please report them in the discussion section of the "Community" tab on Hugging Face.
For more substantial contributions or collaboration opportunities, feel free to contact us directly.
Dataset Summary
PARHAF-biomarkers-annotated is a subpart of the PARHAF corpus, an open French corpus of human-authored clinical… See the full description on the dataset page: https://huggingface.co/datasets/HealthDataHub/PARHAF-biomarkers-annotated.openslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.lsec-biomarkers-splicing-viz-mt1g-plot-v1
lsec-biomarkers-splicing-viz-mt1g-plot-v1
VALERIE v2.1.2 PlotPSI output for the MT1G event (851 split LSEC cells, cell.types=Healthy,F0,F2-3,F4, method=kw). Per-group split cells with >=2 region reads (coverage proxy): Healthy=410,F0=211,F2-3=219,F4=5.
Dataset Info
Rows: 2
Columns: 6
Columns
Column
Type
Description
image
Image(mode=None, decode=True)
PNG plot from PlotPSI: per-cell coverage-ratio PSI heatmap, mean PSI +/- bootstrap CI… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/lsec-biomarkers-splicing-viz-mt1g-plot-v1.lsec-biomarkers-splicing-viz-stab2-plot-v1
lsec-biomarkers-splicing-viz-stab2-plot-v1
VALERIE v2.1.2 PlotPSI output for the STAB2 event (1929 split LSEC cells, cell.types=Healthy,F0,F2-3,F4, method=kw). Per-group split cells with >=2 region reads (coverage proxy): Healthy=779,F0=426,F2-3=478,F4=166.
Dataset Info
Rows: 2
Columns: 6
Columns
Column
Type
Description
image
Image(mode=None, decode=True)
PNG plot from PlotPSI: per-cell coverage-ratio PSI heatmap, mean PSI +/- bootstrap… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/lsec-biomarkers-splicing-viz-stab2-plot-v1.lsec-biomarkers-splicing-viz-stab2-percell-counts-v1
lsec-biomarkers-splicing-viz-stab2-percell-counts-v1
Per-cell region-read coverage table for the STAB2 event, all 17 samples. Per-cell coverage-proxy table (region_reads only -- RI PSI is computed by VALERIE internally from per-base read-coverage ratios over the intron span, not from junction counts we compute ourselves; see EXPERIMENT_README.md section 4).
Dataset Info
Rows: 1929
Columns: 4
Columns
Column
Type
Description
cb… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/lsec-biomarkers-splicing-viz-stab2-percell-counts-v1.openslr-32-hq-SA-languages-Setswana
High quality TTS data for four South African languages - Setswana
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Setswana
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.lsec-biomarkers-splicing-viz-mt1g-percell-counts-v1
lsec-biomarkers-splicing-viz-mt1g-percell-counts-v1
Per-cell region-read coverage table for the MT1G event, all 17 samples. Per-cell coverage-proxy table (region_reads only -- A3SS PSI is computed by VALERIE internally from per-base read-coverage ratios, not from junction counts we compute ourselves; see EXPERIMENT_README.md section 4).
Dataset Info
Rows: 851
Columns: 4
Columns
Column
Type
Description
cb
Value('string')
cell barcode… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/lsec-biomarkers-splicing-viz-mt1g-percell-counts-v1.clinical-biomarkers
500 Clinical Biomarkers and Longevity Reference Ranges
Curated by PurpleDaisy | Creators of Meridian: Lab Records Tracker
A structured clinical dataset of 500 diagnostic blood and urinary biomarkers covering standard commercial reference intervals (Quest / Labcorp), evidence-based optimal longevity targets, and organ system classifications.
Official Web Verification & Interactive Tools
Every biomarker in this dataset is cross-referenced with full physiological… See the full description on the dataset page: https://huggingface.co/datasets/purpledaisy-studio/clinical-biomarkers.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Sesotho.openslr-32-hq-SA-languages-isiXhosa
High quality TTS data for four South African languages - isiXhosa
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - isiXhosa
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-isiXhosa.sma-bforsma-reported-biomarkers
BforSMA Reported Biomarkers
This repository contains the eight documents released with the Biomarkers for Spinal Muscular Atrophy (BforSMA) clinical study.
The study involved 108 children with SMA and 22 controls and reported candidate proteins, metabolites, and transcripts. The public Figshare package is valuable, but it consists of report/protocol documents and tables—not a ready-to-train participant-by-feature matrix. The repository name says reported biomarkers to make that… See the full description on the dataset page: https://huggingface.co/datasets/YannisTevissen/sma-bforsma-reported-biomarkers.neuroinflammation-biomarkers-depression-2020-2024
Neuroinflammation Biomarkers in Depression (2020-2024)
Author: Juan Moises de la Serna | ORCID: 0000-0002-8401-8018Full dataset: Harvard Dataverse doi.org/10.7910/DVN/QBQXKL
Description
Global meta-analysis of neuroinflammation biomarkers (IL-6, TNF-α, CRP, IL-1β) in depression (MDD), bipolar disorder (BD), and healthy controls. 147 studies, 15,000+ participants, 2020–2024.
Biomarker
MDD
HC
Ratio
IL-6
7.8 pg/mL
2.4 pg/mL
3.25×
TNF-α
18.5 pg/mL
8.2 pg/mL… See the full description on the dataset page: https://huggingface.co/datasets/juanmoisesdelas/neuroinflammation-biomarkers-depression-2020-2024.openslr-148-hq-Nahuatl
