datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-dwcoco-detection-stringsProcessed the bounding boxes from coco to paligemma like.
Reference dataset -> detection-datasets/coco
DFFT-Video-Test-Stringbacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.string-opsSTRING
STRING v12.0
STRING is a protein association network database that integrates experimental, computational, text-mined, and curated evidence for functional and physical protein interactions.
Configs
Config
Raw source
Description
protein_links
protein.links.full.v12.0.txt.gz
Protein-protein association edges with all STRING evidence channels and combined_score.
protein_info
protein.info.v12.0.txt.gz
Protein identifiers, preferred names, sizes, and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/STRING.StringDBSeqsv12All the IDs and sequences in StringDB version 12
https://string-db.org/cgi/download
code-meta-reasoning-cleaned-final-string-idtask079_conala_concat_strings
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task079_conala_concat_strings
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task079_conala_concat_strings.StringWars
StringKilla - Small Datasets for String Algorithms Benchmarking
The goal of this dataset is to provide a fairly diverse set of strings to evalute the performance of various string-processing algorithms in StringZilla and beyond.
English Texts
English Leipzig Corpora Collection
124 MB uncompressed
1'000'000 lines of ASCII
8'388'608 tokens of mean length 5
The dataset was originally pulled from Princeton's website:
wget --no-clobber -O leipzig1M.txt… See the full description on the dataset page: https://huggingface.co/datasets/ashvardanian/StringWars.task1189_check_char_in_string
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1189_check_char_in_string
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1189_check_char_in_string.task600_find_the_longest_common_substring_in_two_strings
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task600_find_the_longest_common_substring_in_two_strings
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task600_find_the_longest_common_substring_in_two_strings.OpenScore-StringQuartets
OpenScore String Quartets (OMR Evaluation)
This dataset is derived from the OpenScore String Quartets corpus (Gotham et al., 2023), a collection of string quartets by "long 19th century" composers. It is designed for evaluating Optical Music Recognition (OMR) systems.
We extract a subset of the OpenScore String Quartets that contains both scanned images of real scores and the corresponding MusicXML ground truth. We also render clean images from the MusicXML files using MuseScore.… See the full description on the dataset page: https://huggingface.co/datasets/guangyangmusic/OpenScore-StringQuartets.genminiall_no_na_no_weird_stringriddles_v1_stringified-jsonifizechess-time-control-string-parsing
Chess Time-Control String Parsing
Real-world chess time-control strings, in two forms:
.txt files — the source of truth. Every unique time-control string, one per line, with a frequency count. These are the raw, real strings (messy, multilingual, sometimes junk) as scraped/collected. No interpretation.
.jsonl files — a tagged, partially-correct derived artifact. Each unique string with an auto-derived (category, stages) parse. The tags are heuristics, not verified ground truth… See the full description on the dataset page: https://huggingface.co/datasets/gutsy-gambit/chess-time-control-string-parsing.task1316_remove_duplicates_string
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1316_remove_duplicates_string
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1316_remove_duplicates_string.amc22-24_stop_stringsAMC-12(2022-2024)
tulu_3_rewritten_400k_string_f1_only_v2_nocode_all_filtered_qwen2_5_openthoughts2bacbench-ppi-stringdb-protein-sequences-small
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.gsm8k_stop_stringsbacbench-ppi-stringdb-dna-small
Dataset for protein-protein interaction prediction across bacteria (DNA)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genomes' PPI scores have been extracted from STRING DB and their associated DNA from GenBank (https://www.ncbi.nlm.nih.gov/genbank/).
Each row contains a set of DNA sequences from a genome, and a set of associated PPI scores.
The PPI scores have been extracted using the combined score… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-dna-small.STRING_V12_TrainingSet**Repository: https://stringdb-downloads.org/download/protein.physical.links.v12.0.txt.gz
**Reference: Szklarczyk, D. et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research 51, D638–D646 (2023).
Stringv12ModelOrgSeqsstring_ppi_human_5MTEKGEN-Strings-100Kaime2025_stop_stringsmmlu-pro_stop_stringscustom_lerobot_dataset_with_string_feature_0722_1050This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 3,
"total_frames": 30,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/welfarefit/custom_lerobot_dataset_with_string_feature_0722_1050.math500_stop_strings
