datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bacbench-ppi-stringdb-protein-sequences
Dataset for protein-protein interaction prediction across bacteria (Protein sequences)
A dataset of 10,533 bacterial genomes across 6,956 species with protein-protein interaction (PPI) scores for each genome.
The genome protein sequences and PPI scores have been extracted from STRING DB.
Each row contains a set of protein sequences from a genome, ordered by their location on the chromosome and plasmids and a set of associated PPI scores.
The PPI scores have been extracted using the… See the full description on the dataset page: https://huggingface.co/datasets/macwiatrak/bacbench-ppi-stringdb-protein-sequences.STRING
STRING v12.0
STRING is a protein association network database that integrates experimental, computational, text-mined, and curated evidence for functional and physical protein interactions.
Configs
Config
Raw source
Description
protein_links
protein.links.full.v12.0.txt.gz
Protein-protein association edges with all STRING evidence channels and combined_score.
protein_info
protein.info.v12.0.txt.gz
Protein identifiers, preferred names, sizes, and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/STRING.string_ppi_human_5Mstringman-socks-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "stringman",
"total_episodes": 42,
"total_frames": 27325,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:42"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/naavox/stringman-socks-2.custom_lerobot_dataset_with_string_feature_0722_1050This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 3,
"total_frames": 30,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/welfarefit/custom_lerobot_dataset_with_string_feature_0722_1050.lerobot_dataset_custom_string_feature_0722_1051This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 3,
"total_frames": 30,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/welfarefit/lerobot_dataset_custom_string_feature_0722_1051.eval_test_string_v1_run05This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1490,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sminjade/eval_test_string_v1_run05.lerobot_dataset_custom_string_feature_0722_1100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 3,
"total_frames": 30,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/welfarefit/lerobot_dataset_custom_string_feature_0722_1100.stringman-socks-3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "stringman",
"total_episodes": 30,
"total_frames": 20840,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/naavox/stringman-socks-3.hub_datasets_string_statisticseval_test_string_v1_run03This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1483,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sminjade/eval_test_string_v1_run03.go_pgn_string_v2
go_pgn_string_v2
It is parsed from https://dl.fbaipublicfiles.com/elfopengo/analysis/data/gogod_commentary_sgfs.gzip.It contains professional game of Go ever played (~100k games drawn from GoGoD), evaluated by Meta AI's ELF OpenGo.The leftmost variation of the game tree is taken in SGF format and translate it into PGN like format.Due to autoregressive nature of decoder, a special token '>' is used to denote the move by the winner of the game.
eval_test_string_v1_run02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1457,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sminjade/eval_test_string_v1_run02.eval_test_string_v1_run04This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1483,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sminjade/eval_test_string_v1_run04.string_ppi_human_1M
STRING PPI Human 1M
This dataset contains 1 million human protein–protein interactions (PPIs) derived from STRING v11.5.
Columns:
seq_a, seq_b: Amino acid sequences of the interacting proteins (≤2048 AA).
seq_name_a, seq_name_b: Protein names from STRING.
score: Combined score from STRING (0–1000, normalized to 0–1). This score integrates various evidence channels (experimental data, text mining, co-expression, etc.) into a single confidence metric.
label: Binary interaction label.… See the full description on the dataset page: https://huggingface.co/datasets/vladak/string_ppi_human_1M.test_string_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 15,
"total_frames": 11568,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sminjade/test_string_v1.test_string_to_dict
Dataset Card for "test_string_to_dict"
More Information needed
count_letters_in_random_string_variedstringleveldigits.252qwen3_0.6b-rlvr_task600_find_the_longest_common_substring_in_two_stringseval_test_string_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 1487,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sminjade/eval_test_string_v1.tulu_3_rewritten_only_string_f1_v2_10k_samplet-eval-reason-stringpermuted-letter-string-analogies
Permuted Letter-String Analogies
This repository contains datasets introduced in Hellwig et al. (2026). The datasets are an extension of the letter-string analogies introduced in Lewis & Mitchell (2025).
Each folder contains a dataset with different data attributes, and contains a train, validation, and test set.
The naming convention is:
all_transformations_<copy>_study<N>_perm<N>
all_transformations:
Below are illustrations for each transformation on the standard alphabet.… See the full description on the dataset page: https://huggingface.co/datasets/philipp-hellwig/permuted-letter-string-analogies.
