CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FishCaduceus /FishCaduceus-Functional-Site-Benchmark FishCaduceus Functional Site Benchmark Dataset description This dataset contains sequence-based benchmarks for four gene-annotation tasks used to evaluate FishCaduceus: translation initiation site (TIS) prediction translation termination site (TTS) prediction splice donor site prediction splice acceptor site prediction The benchmark was designed to evaluate both within-species performance and cross-species transfer. Models are trained and selected only with… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Functional-Site-Benchmark.text1M<n<10M0 likes103 downloads2mo agoHugging Face02FishCaduceus /FishCaduceus-Evolutionary-Constraint-Benchmark FishCaduceus Evolutionary Constraint Benchmark Dataset description This dataset contains sequence-based benchmarks for evaluating whether FishCaduceus representations capture evolutionary constraint in fish genomes. Constraint labels were derived from a 26-fish whole-genome alignment generated with Progressive Cactus. Grass carp (Ctenopharyngodon idella) was used as the primary reference genome for defining aligned and conserved positions. The resulting labeled… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Evolutionary-Constraint-Benchmark.0 likes56 downloads2mo agoHugging Face03FishCaduceus /FishCaduceus-Pretraining-1024 FishCaduceus Pretraining Dataset 1024 Dataset description This dataset contains fixed-length 1,024-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling. Dataset summary The dataset contains 2,709,306 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-1024.tabular1M<n<10M0 likes25 downloads2mo agoHugging Face04FishCaduceus /FishCaduceus-Pretraining-512 FishCaduceus Pretraining Dataset 512 Dataset description This dataset contains fixed-length 512-bp genomic sequence windows prepared for pretraining FishCaduceus DNA language models. It uses single-nucleotide tokenization and is intended for masked language modeling. Dataset summary The dataset contains 6,087,221 JSON Lines records. Each record contains genomic coordinates and a seq field. The six files were audited by streaming line-by-line… See the full description on the dataset page: https://huggingface.co/datasets/FishCaduceus/FishCaduceus-Pretraining-512.tabular1M<n<10M0 likes18 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.