datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
protein-ligand-design
🧪 Protein-Ligand Design Gym — Team JAMMY
poolside Laguna Hackathon submission. A tool-use reinforcement-learning
environment that teaches an LLM to reason like a bench computational chemist /
protein engineer — by measuring, not guessing.
The problem
Proteins are the molecular machines inside living cells, each built from a long
string of amino-acid "letters". Ligands are the small molecules — most drugs
among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.Laguna-S-2.1-trajectories
Laguna S 2.1 Trajectories
ShareGPT string conversion of
Poolside's public Laguna S 2.1 trajectory archive.
Format
{
"id": "...",
"conversations": [
{"from": "human", "value": "..."},
{"from": "gpt", "value": "..."}
],
"source": "poolside-laguna-s-2.1/<benchmark>/<variant>"
}
Assistant reasoning, messages, and tool calls are stored in gpt turns.
Tool results are stored in human turns.
Contents
10,287 trajectories across thinking and… See the full description on the dataset page: https://huggingface.co/datasets/mgoin/Laguna-S-2.1-trajectories.processrl-terminal-environments
ProcessRL Terminal Environments
ProcessRL is a collection of behavior-conditioned terminal environments for training and evaluating agent process control. The tasks are designed around failures that appear in interactive terminal work: stopping after a misleading successful command, repeating an unproductive action, failing to pivot after a dead end, losing track of migrated state, and leaving partial progress unfinished.
This release contains the first public train/heldout… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/processrl-terminal-environments.laguna-xs-ultrachat-responsespy-bug-trace-laguna-xs-2-l1-rolloutslaguna-xs-ultrachat-conversationspy-bug-trace-laguna-xs-2-l1-rolloutsumpalumpas
Compaction Dataset Based on SWE-Smith
This is a sample of a larger dataset. Work in progress.
Parquet Layout
This repository contains a single Parquet file:
File
Rows
What it contains
data/oracle_trajectories.parquet
10
One row per oracle trajectory, preprocessed with linked offline checkpoints, oracle memories, and oracle continuations.
The compaction data is represented by linked offline H2 fixed-milestone
checkpoints:
offline_checkpoint_count… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/umpalumpas.Laguna-S-2.1-REAP-Mixed-Observer-v1py-bug-trace-laguna-m-1-free-l1-rolloutslaguna-xs-magpie-300k-responsespy-bug-trace-laguna-xs-2-l2-rolloutslaguna-xs-magpie-300k-conversationspy-bug-trace-laguna-xs-2-l2-rolloutspy-bug-trace-laguna-xs-2-l3-rolloutspy-bug-trace-laguna-xs-2-l3-rollouts
