datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mint-1t-pdf-gte6-1.3M
Dataset size: 1.3M entries
Source: Derived from the mint-1t dataset
Criteria: Contains samples with greater than or equal to 6 images (gte6)
mint-1t-pdf-gte6mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
arxiv_20240801_gte-base-en-v1.5_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked arxiv papers from Semantic Scholar. The embedding model used is Alibaba-NLP/gte-large-en-v1.5.
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/arxiv_20240801_gte-base-en-v1.5_qdrant_index.gtex-10M-balanced-tiles
GTEx 10M Balanced Tiles
This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed.
Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.spliceformer-gtex-tissues
Spliceformer — GTEx per-tissue test sets
Per-tissue, per-donor splice test sets used to evaluate
Spliceformer across tissues.
python scripts/download_assets.py --gtex-tissues lung
python scripts/evaluate_splice.py --context 10k --ensemble --tissue lung
Tissues
brain_cortex · lung · testis · whole_blood · haec10
Layout
{tissue}/
├── ml_data_var/
│ ├── individual/test_GTEX-*.h5 per-donor test sets (~650 MB each)
│ └── combined_test.h5… See the full description on the dataset page: https://huggingface.co/datasets/SaumyaGupta-99/spliceformer-gtex-tissues.trst3_20260902_150541This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/gtext/trst3_20260902_150541.Voetbal1_20260904_120709This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/gtext/Voetbal1_20260904_120709.arxiv_embeddings_Alibaba-NLP_gte-base-en-v1.5msmarco-v2.1-gte-large-en-v1.5
Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.bright-passage-index-gte_qwen2-1_5bfineweb-gte-multilingual-base-shard-0hoike_normal_expression_GTEx_Analysis_v10_log2tpmplus1
Hōʻike - Normal GTEx Gene Expression Data in Log2(TPM+1) Format
These data are for use in the Hoike gene expression data generation models as the normal_profile input.
These were obtained from the GTEx_Analysis_v10_RNASeQCv2.4.2_gene_tpm.gct.gz file from the GTEx Portal on 6/10/2026.
gteaGTEA is composed of 50 recorded videos of 25 participants making two different mixed salads. The videos are captured by a camera with a top-down view onto the work-surface. The participants are provided with recipe steps which are randomly sampled from a statistical recipe model.gtex-single-cell-rnaseq
GTEx Single-Cell RNA-seq Dataset
This repository provides tools to create a Hugging Face dataset from GTEx single-nucleus RNA-seq data, transforming the hierarchical H5AD format into a flat, ML-ready structure.
Overview
Data Source
The data comes from GTEx's snRNA-seq atlas:
Source: GTEx Portal
Publication: Eraslan et al., Science 2022 - "Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function"
Content: 209… See the full description on the dataset page: https://huggingface.co/datasets/ai-department-lpnu/gtex-single-cell-rnaseq.Alibaba-NLP__gte-Qwen2-7B-instruct-details
Dataset Card for Evaluation run of Alibaba-NLP/gte-Qwen2-7B-instruct
Dataset automatically created during the evaluation run of model Alibaba-NLP/gte-Qwen2-7B-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Alibaba-NLP__gte-Qwen2-7B-instruct-details.paws-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the google-research-datasets/paws dataset
This is the google-research-datasets/paws dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/paws-gte-modernbert-pooled.book_dataset_no_mem_token_gte_largev1_5_M512_C1024_1Bpubmedqa-query-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the qiaojin/PubMedQA dataset
This is the qiaojin/PubMedQA dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/pubmedqa-query-gte-modernbert-pooled.mimic_clinical_notes_gte2_from_mimic_note_preprocessed
MIMIC-IV Clinical Notes - Patients with 2+ Notes (From Original Dataset)
Dataset Description
This dataset contains discharge notes for all patients from the MIMIC-IV Note dataset who have 2 or more clinical notes. The data has been merged with admissions information to provide comprehensive patient context.
Dataset Statistics
Total Clinical Notes: 246,158
Unique Patients: 60,298
Unique Hospital Admissions (hadm_id): 246,158
Total Columns: 15
Source… See the full description on the dataset page: https://huggingface.co/datasets/mimic-capstone/mimic_clinical_notes_gte2_from_mimic_note_preprocessed.medical-v002-fineweb-10bt-gte-1msmarco-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/msmarco-corpus dataset
This is the sentence-transformers/msmarco-corpus dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-gte-modernbert-pooled.MedRag-textbooks-gte-large-en-v1.5local-emoji-search-gte
local emoji semantic search
Emoji, their text descriptions and precomputed text embeddings with Alibaba-NLP/gte-large-en-v1.5 for use in emoji semantic search.
This work is largely inspired by the original emoji-semantic-search repo and aims to provide the data for fully local use, as the demo is not working as of a few days ago.
This repo only contains a precomputed embedding "database", equivalent to server/emoji-embeddings.jsonl.gz in the original repo, to be used as the… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/local-emoji-search-gte.catalog-fusion-gtex-v8TrASPr_GTEx_datadataset_gtell_psiq_ptgte-multilingual-product-ads-1MGTEx-WSI-CloseQA-Balancedmdlr-query-gte-modernbert-pooled
Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/mldr dataset
This is the sentence-transformers/mldr dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend truncating to the… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/mdlr-query-gte-modernbert-pooled.
