CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01StringFellow /fusion-dwtext1M<n<10M11 likes119k downloads1h agoHugging Face02ariG23498 /coco-detection-stringsProcessed the bounding boxes from coco to paligemma like. Reference dataset -> detection-datasets/coco image100K<n<1M3 likes320 downloads1y agoHugging Face03ohjoonhee /DFFT-Video-Test-Stringtext10K<n<100K0 likes308 downloads10mo agoHugging Face04macwiatrak /bacbench-ppi-stringdb-protein-sequences Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.tabular1K<n<10K0 likes261 downloads1y agoHugging Face05ardauzunoglu /string-opstext100K<n<1M0 likes203 downloads6mo agoHugging Face06LiteFold /STRING STRING v12.0 STRING is a protein association network database that integrates experimental, computational, text-mined, and curated evidence for functional and physical protein interactions. Configs Config Raw source Description protein_links protein.links.full.v12.0.txt.gz Protein-protein association edges with all STRING evidence channels and combined_score. protein_info protein.info.v12.0.txt.gz Protein identifiers, preferred names, sizes, and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/STRING.tabular100M<n<1B1 likes196 downloads4mo agoHugging Face07Synthyra /StringDBSeqsv12All the IDs and sequences in StringDB version 12 https://string-db.org/cgi/download text10M<n<100M1 likes167 downloads2y agoHugging Face08allenai /code-meta-reasoning-cleaned-final-string-idtext100K<n<1M5 likes165 downloads1y agoHugging Face09Lots-of-LoRAs /task079_conala_concat_strings Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task079_conala_concat_strings Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task079_conala_concat_strings.texttext-generation1K<n<10K0 likes154 downloads2y agoHugging Face10ashvardanian /StringWars StringKilla - Small Datasets for String Algorithms Benchmarking The goal of this dataset is to provide a fairly diverse set of strings to evalute the performance of various string-processing algorithms in StringZilla and beyond. English Texts English Leipzig Corpora Collection 124 MB uncompressed 1'000'000 lines of ASCII 8'388'608 tokens of mean length 5 The dataset was originally pulled from Princeton's website: wget --no-clobber -O leipzig1M.txt… See the full description on the dataset page: https://huggingface.co/datasets/ashvardanian/StringWars.textfeature-extraction100K<n<1M1 likes136 downloads1y agoHugging Face11Lots-of-LoRAs /task1189_check_char_in_string Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1189_check_char_in_string Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1189_check_char_in_string.texttext-generationn<1K0 likes93 downloads2y agoHugging Face12Lots-of-LoRAs /task600_find_the_longest_common_substring_in_two_strings Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task600_find_the_longest_common_substring_in_two_strings Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task600_find_the_longest_common_substring_in_two_strings.texttext-generation1K<n<10K0 likes87 downloads2y agoHugging Face13guangyangmusic /OpenScore-StringQuartetsgated OpenScore String Quartets (OMR Evaluation) This dataset is derived from the OpenScore String Quartets corpus (Gotham et al., 2023), a collection of string quartets by "long 19th century" composers. It is designed for evaluating Optical Music Recognition (OMR) systems. We extract a subset of the OpenScore String Quartets that contains both scanned images of real scores and the corresponding MusicXML ground truth. We also render clean images from the MusicXML files using MuseScore.… See the full description on the dataset page: https://huggingface.co/datasets/guangyangmusic/OpenScore-StringQuartets.imageimage-to-textn<1K3 likes87 downloads6mo agoHugging Face14qfq /genminiall_no_na_no_weird_stringtext10K<n<100K0 likes85 downloads2y agoHugging Face15jsonifize /riddles_v1_stringified-jsonifizetextn<1K0 likes80 downloads3y agoHugging Face16gutsy-gambit /chess-time-control-string-parsing Chess Time-Control String Parsing Real-world chess time-control strings, in two forms: .txt files — the source of truth. Every unique time-control string, one per line, with a frequency count. These are the raw, real strings (messy, multilingual, sometimes junk) as scraped/collected. No interpretation. .jsonl files — a tagged, partially-correct derived artifact. Each unique string with an auto-derived (category, stages) parse. The tags are heuristics, not verified ground truth… See the full description on the dataset page: https://huggingface.co/datasets/gutsy-gambit/chess-time-control-string-parsing.texttext-generation1K<n<10K0 likes73 downloads2mo agoHugging Face17Lots-of-LoRAs /task1316_remove_duplicates_string Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1316_remove_duplicates_string Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1316_remove_duplicates_string.texttext-generationn<1K0 likes66 downloads2y agoHugging Face18graphUQ-ls-hxy /amc22-24_stop_stringsAMC-12(2022-2024) textn<1K0 likes54 downloads9mo agoHugging Face19hamishivi /tulu_3_rewritten_400k_string_f1_only_v2_nocode_all_filtered_qwen2_5_openthoughts2text10K<n<100K0 likes47 downloads1y agoHugging Face20macwiatrak /bacbench-ppi-stringdb-protein-sequences-small Dataset for protein-protein interaction prediction across bacteria (Protein sequences) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genome protein sequences and PPI scores have been extracted from STRING DB. Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores. The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.textn<1K0 likes47 downloads5mo agoHugging Face21graphUQ-ls-hxy /gsm8k_stop_stringstext1K<n<10K0 likes47 downloads10mo agoHugging Face22macwiatrak /bacbench-ppi-stringdb-dna-small Dataset for protein-protein interaction prediction across bacteria (DNA) A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome. The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/). Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores. The PPI scores have been extracted using the combined score… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-dna-small.textn<1K0 likes43 downloads5mo agoHugging Face23danliu1226 /STRING_V12_TrainingSet**Repository: https://stringdb-downloads.org/download/protein.physical.links.v12.0.txt.gz **Reference: Szklarczyk, D. et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research 51, D638–D646 (2023). text100K<n<1M0 likes42 downloads1y agoHugging Face24Synthyra /Stringv12ModelOrgSeqstext100K<n<1M0 likes41 downloads1y agoHugging Face25vladak /string_ppi_human_5Mtabular1M<n<10M1 likes40 downloads1y agoHugging Face26YesaOuO /TEKGEN-Strings-100Ktext100K<n<1M0 likes36 downloads2y agoHugging Face27graphUQ-ls-hxy /aime2025_stop_stringstextn<1K0 likes35 downloads10mo agoHugging Face28graphUQ-ls-hxy /mmlu-pro_stop_stringstext1K<n<10K0 likes35 downloads6mo agoHugging Face29welfarefit /custom_lerobot_dataset_with_string_feature_0722_1050This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 3, "total_frames": 30, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/welfarefit/custom_lerobot_dataset_with_string_feature_0722_1050.tabularroboticsn<1K0 likes34 downloads1y agoHugging Face30graphUQ-ls-hxy /math500_stop_stringstext1K<n<10K0 likes33 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.