CoolFace
Datasetpublic

vandijklab/immune-c2s

Overview Cell2Sentence is a novel method for adapting large language models to single-cell transcriptomics. We transform single-cell RNA sequencing data into sequences of gene names ordered by expression level, termed "cell sentences". This dataset was constructed from the immune tissue dataset in Domínguez et al., and it was used to train the Pythia-160m model capable of generating complete cells described in our paper. Details about the Cell2Sentence transformation and… See the full description on the dataset page: https://huggingface.co/datasets/vandijklab/immune-c2s.

sourceHugging Facecc-by-nc-nd-4.0updated 3y agoView on Hugging Face
3likes144downloads
Dataset Card

Overview

Cell2Sentence is a novel method for adapting large language models to single-cell transcriptomics. We transform single-cell RNA sequencing data into sequences of gene names ordered by expression level, termed "cell sentences". This dataset was constructed from the immune tissue dataset in Domínguez et al., and it was used to train the Pythia-160m model capable of generating complete cells described in our paper. Details about the Cell2Sentence transformation and preprocessing pipeline can be found in our paper and GitHub repo linked below.

GitHub: <https://github.com/vandijklab/cell2sentence-ft> Paper: <https://www.biorxiv.org/content/10.1101/2023.09.11.557287v3> Model Card: <https://huggingface.co/vandijklab/pythia-160m-c2s>