CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01paralym /mint-1t-pdf-gte6-1.3M Dataset size: 1.3M entries Source: Derived from the mint-1t dataset Criteria: Contains samples with greater than or equal to 6 images (gte6) 0 likes2.1k downloads2y agoHugging Face02paralym /mint-1t-pdf-gte60 likes1.5k downloads2y agoHugging Face03paralym /mint-1t-html-images-gte6-sample Size: 6769158 images sampled from Mint-1t-html Criteria: Data entries with greater than or equal to 6 images (gte6) image1M<n<10M0 likes1.1k downloads2y agoHugging Face04stanford-oval /arxiv_20240801_gte-base-en-v1.5_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked arxiv papers from Semantic Scholar. The embedding model used is Alibaba-NLP/gte-large-en-v1.5. This index is compatible with WikiChat v2.0. Refer to the following for more information: GitHub repository: https://github.com/stanford-oval/WikiChat Papers: WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia SPAGHETTI: Open-Domain Question Answering from Heterogeneous… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/arxiv_20240801_gte-base-en-v1.5_qdrant_index.text-retrieval100M<n<1B0 likes935 downloads2y agoHugging Face05medarc /gtex-10M-balanced-tiles GTEx 10M Balanced Tiles This dataset contains 10,000,000 JPEG-encoded 224x224 pathology tiles from GTEx SVS slides in s3://path-datasets/gtex/svs_by_tissue, balanced at 250,000 tiles for each of the 40 tissue prefixes. Source objects under 80,000,000 bytes are ignored because the GTEx prefix contains tiny unsupported SVS objects. Tiles are sampled from virtual OpenSlide levels 0, 1, and 2, so the effective field of view varies while the emitted JPEG size stays fixed. Each… See the full description on the dataset page: https://huggingface.co/datasets/medarc/gtex-10M-balanced-tiles.tabularimage-feature-extraction10M<n<100M0 likes368 downloads3mo agoHugging Face06SaumyaGupta-99 /spliceformer-gtex-tissues Spliceformer — GTEx per-tissue test sets Per-tissue, per-donor splice test sets used to evaluate Spliceformer across tissues. python scripts/download_assets.py --gtex-tissues lung python scripts/evaluate_splice.py --context 10k --ensemble --tissue lung Tissues brain_cortex · lung · testis · whole_blood · haec10 Layout {tissue}/ ├── ml_data_var/ │ ├── individual/test_GTEX-*.h5 per-donor test sets (~650 MB each) │ └── combined_test.h5… See the full description on the dataset page: https://huggingface.co/datasets/SaumyaGupta-99/spliceformer-gtex-tissues.token-classification0 likes274 downloads16d agoHugging Face07gtext /trst3_20260902_150541This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/gtext/trst3_20260902_150541.tabularrobotics1K<n<10K0 likes204 downloads19d agoHugging Face08gtext /Voetbal1_20260904_120709This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/gtext/Voetbal1_20260904_120709.tabularrobotics1K<n<10K0 likes204 downloads18d agoHugging Face09bluuebunny /arxiv_embeddings_Alibaba-NLP_gte-base-en-v1.5text1M<n<10M1 likes150 downloads2y agoHugging Face10spacemanidol /msmarco-v2.1-gte-large-en-v1.5 Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods. Note, that the embeddings are not normalized so you will need to normalize them before usage. Retrieval Performance Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.textquestion-answering10M<n<100M0 likes109 downloads1y agoHugging Face11icedpanda /bright-passage-index-gte_qwen2-1_5btext1M<n<10M0 likes106 downloads1y agoHugging Face12lsb /fineweb-gte-multilingual-base-shard-0text1M<n<10M0 likes73 downloads1y agoHugging Face13silicobio /hoike_normal_expression_GTEx_Analysis_v10_log2tpmplus1 Hōʻike - Normal GTEx Gene Expression Data in Log2(TPM+1) Format These data are for use in the Hoike gene expression data generation models as the normal_profile input. These were obtained from the GTEx_Analysis_v10_RNASeQCv2.4.2_gene_tpm.gct.gz file from the GTEx Portal on 6/10/2026. tabular10K<n<100K0 likes66 downloads2mo agoHugging Face14dinggd /gteaGTEA is composed of 50 recorded videos of 25 participants making two different mixed salads. The videos are captured by a camera with a top-down view onto the work-surface. The participants are provided with recipe steps which are randomly sampled from a statistical recipe model.0 likes61 downloads2y agoHugging Face15ai-department-lpnu /gtex-single-cell-rnaseq GTEx Single-Cell RNA-seq Dataset This repository provides tools to create a Hugging Face dataset from GTEx single-nucleus RNA-seq data, transforming the hierarchical H5AD format into a flat, ML-ready structure. Overview Data Source The data comes from GTEx's snRNA-seq atlas: Source: GTEx Portal Publication: Eraslan et al., Science 2022 - "Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function" Content: 209… See the full description on the dataset page: https://huggingface.co/datasets/ai-department-lpnu/gtex-single-cell-rnaseq.tabularfeature-extraction100K<n<1M1 likes59 downloads11mo agoHugging Face16open-llm-leaderboard /Alibaba-NLP__gte-Qwen2-7B-instruct-detailsgated Dataset Card for Evaluation run of Alibaba-NLP/gte-Qwen2-7B-instruct Dataset automatically created during the evaluation run of model Alibaba-NLP/gte-Qwen2-7B-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Alibaba-NLP__gte-Qwen2-7B-instruct-details.tabular10K<n<100K1 likes50 downloads2y agoHugging Face17stephantulkens /paws-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the google-research-datasets/paws dataset This is the google-research-datasets/paws dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/paws-gte-modernbert-pooled.text1M<n<10M0 likes42 downloads11mo agoHugging Face18infinitylogesh /book_dataset_no_mem_token_gte_largev1_5_M512_C1024_1Btext100K<n<1M0 likes42 downloads9mo agoHugging Face19stephantulkens /pubmedqa-query-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the qiaojin/PubMedQA dataset This is the qiaojin/PubMedQA dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to the model’s maximum token… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/pubmedqa-query-gte-modernbert-pooled.text100K<n<1M0 likes41 downloads11mo agoHugging Face20mimic-capstone /mimic_clinical_notes_gte2_from_mimic_note_preprocessedgated MIMIC-IV Clinical Notes - Patients with 2+ Notes (From Original Dataset) Dataset Description This dataset contains discharge notes for all patients from the MIMIC-IV Note dataset who have 2 or more clinical notes. The data has been merged with admissions information to provide comprehensive patient context. Dataset Statistics Total Clinical Notes: 246,158 Unique Patients: 60,298 Unique Hospital Admissions (hadm_id): 246,158 Total Columns: 15 Source… See the full description on the dataset page: https://huggingface.co/datasets/mimic-capstone/mimic_clinical_notes_gte2_from_mimic_note_preprocessed.tabular100K<n<1M2 likes41 downloads8mo agoHugging Face21ashercn97 /medical-v002-fineweb-10bt-gte-1tabular100K<n<1M0 likes38 downloads1y agoHugging Face22stephantulkens /msmarco-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/msmarco-corpus dataset This is the sentence-transformers/msmarco-corpus dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/msmarco-gte-modernbert-pooled.text1M<n<10M0 likes33 downloads11mo agoHugging Face23AtAndDev /MedRag-textbooks-gte-large-en-v1.5text100K<n<1M0 likes29 downloads2y agoHugging Face24pszemraj /local-emoji-search-gte local emoji semantic search Emoji, their text descriptions and precomputed text embeddings with Alibaba-NLP/gte-large-en-v1.5 for use in emoji semantic search. This work is largely inspired by the original emoji-semantic-search repo and aims to provide the data for fully local use, as the demo is not working as of a few days ago. This repo only contains a precomputed embedding "database", equivalent to server/emoji-embeddings.jsonl.gz in the original repo, to be used as the… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/local-emoji-search-gte.textfeature-extraction1K<n<10K2 likes28 downloads9mo agoHugging Face25wjixiang /catalog-fusion-gtex-v80 likes28 downloads21h agoHugging Face26MagicSign /TrASPr_GTEx_datatabular10K<n<100K0 likes26 downloads9mo agoHugging Face27drguilhermeapolinario /dataset_gtell_psiq_pttabular1K<n<10K0 likes24 downloads2y agoHugging Face28Qdrant /gte-multilingual-product-ads-1Mtabular1M<n<10M1 likes22 downloads5mo agoHugging Face29CNX-PathLLM /GTEx-WSI-CloseQA-Balancedtabular100K<n<1M0 likes21 downloads1y agoHugging Face30stephantulkens /mdlr-query-gte-modernbert-pooled Embedpress: Alibaba-NLP/gte-modernbert-base on the sentence-transformers/mldr dataset This is the sentence-transformers/mldr dataset, embedded with Alibaba-NLP/gte-modernbert-base. For each example, we embed the text directly (no additional instruction prompt). Embeddings have dimensionality 768. These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search. Because the raw text may exceed the model’s limit, we recommend truncating to the… See the full description on the dataset page: https://huggingface.co/datasets/stephantulkens/mdlr-query-gte-modernbert-pooled.text10K<n<100K0 likes20 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.