datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-dwcoco-detection-stringsProcessed the bounding boxes from coco to paligemma like.
Reference dataset -> detection-datasets/coco
DFFT-Video-Test-Stringtrossen_place_bead_on_string_10_gr00t_crop_02trossen_ai_stationary_place_bead_on_string_15trossen_place_bead_on_string_10_gr00t_01_2bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.trossen_place_bead_on_string_10_gr00t_clip_01trossen_place_bead_on_string_10_gr00t_01string-opsSTRING
STRING v12.0
STRING is a protein association network database that integrates experimental, computational, text-mined, and curated evidence for functional and physical protein interactions.
Configs
Config
Raw source
Description
protein_links
protein.links.full.v12.0.txt.gz
Protein-protein association edges with all STRING evidence channels and combined_score.
protein_info
protein.info.v12.0.txt.gz
Protein identifiers, preferred names, sizes, and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/STRING.StringDBSeqsv12All the IDs and sequences in StringDB version 12
https://string-db.org/cgi/download
code-meta-reasoning-cleaned-final-string-idtask079_conala_concat_strings
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task079_conala_concat_strings
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task079_conala_concat_strings.trossen_place_bead_on_string_10_gr00t_02StringWars
StringKilla - Small Datasets for String Algorithms Benchmarking
The goal of this dataset is to provide a fairly diverse set of strings to evalute the performance of various string-processing algorithms in StringZilla and beyond.
English Texts
English Leipzig Corpora Collection
124 MB uncompressed
1'000'000 lines of ASCII
8'388'608 tokens of mean length 5
The dataset was originally pulled from Princeton's website:
wget --no-clobber -O leipzig1M.txt… See the full description on the dataset page: https://huggingface.co/datasets/ashvardanian/StringWars.task1189_check_char_in_string
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1189_check_char_in_string
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1189_check_char_in_string.task600_find_the_longest_common_substring_in_two_strings
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task600_find_the_longest_common_substring_in_two_strings
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task600_find_the_longest_common_substring_in_two_strings.OpenScore-StringQuartets
OpenScore String Quartets (OMR Evaluation)
This dataset is derived from the OpenScore String Quartets corpus (Gotham et al., 2023), a collection of string quartets by "long 19th century" composers. It is designed for evaluating Optical Music Recognition (OMR) systems.
We extract a subset of the OpenScore String Quartets that contains both scanned images of real scores and the corresponding MusicXML ground truth. We also render clean images from the MusicXML files using MuseScore.… See the full description on the dataset page: https://huggingface.co/datasets/guangyangmusic/OpenScore-StringQuartets.genminiall_no_na_no_weird_stringriddles_v1_stringified-jsonifizetrossen_ai_stationary_place_bead_on_string_10chess-time-control-string-parsing
Chess Time-Control String Parsing
Real-world chess time-control strings, in two forms:
.txt files — the source of truth. Every unique time-control string, one per line, with a frequency count. These are the raw, real strings (messy, multilingual, sometimes junk) as scraped/collected. No interpretation.
.jsonl files — a tagged, partially-correct derived artifact. Each unique string with an auto-derived (category, stages) parse. The tags are heuristics, not verified ground truth… See the full description on the dataset page: https://huggingface.co/datasets/gutsy-gambit/chess-time-control-string-parsing.task1316_remove_duplicates_string
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1316_remove_duplicates_string
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1316_remove_duplicates_string.leandojo-lean4-formal-informal-stringsamc22-24_stop_stringsAMC-12(2022-2024)
tulu_3_rewritten_400k_string_f1_only_v2_nocode_all_filtered_qwen2_5_openthoughts2bacbench-ppi-stringdb-protein-sequences-small
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 261 bacterial genomes across 215 genera with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences-small.gsm8k_stop_stringsStringTheorySolutions
Introduction
This dataset card contains typed solutions and relevant notes to a multitude of problems in string theory that I authored during my study. More notes and solutions are set to be typed as times allows.
Find the solutions and notes available in the Files and Versions section under the file name Reynolds_string_solutions.pdf. The PDF will not render in the UI if using Safari - you will have to
download the file to view it. However, Google Chrome (and likely other browsers)… See the full description on the dataset page: https://huggingface.co/datasets/MarioBarbeque/StringTheorySolutions.
