datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-alphaseq
Open Protein–Protein Interaction Affinity Datasets with AlphaSeq
Protein-protein interactions (PPIs) are fundamental to countless biological processes. One of the most informative biophysical properties of a PPI is the binding affinity: the strength of how two proteins interact. Yet, despite its importance, publicly available affinity data remains limited, constraining the development and benchmarking of protein modeling methods.
Our high-throughput yeast mating assay, AlphaSeq… See the full description on the dataset page: https://huggingface.co/datasets/aalphabio/open-alphaseq.worldquant-swarm-alphas
🐟 Alpha Factory — WorldQuant BRAIN Alpha Discovery Pipeline
One command. Full pipeline. Only generates BRAIN-valid alphas.
git clone https://huggingface.co/datasets/anky2002/worldquant-swarm-alphas
cd worldquant-swarm-alphas
pip install -r requirements.txt
python app.py
Open http://127.0.0.1:7860 — done.
What It Does
Tab
What Happens
Cost
🚀 Run Pipeline
Generate → Lint → Simulate → Mutate → Store. Full DAG.
FREE (no BRAIN credits)
🔍 Test Expression… See the full description on the dataset page: https://huggingface.co/datasets/anky2002/worldquant-swarm-alphas.alphastack-cost-sensitivityalphastack-breakeven-costalphastack-backtest-resultsalphastack-backtest-execution-logalphastack-backtest-trade-logalpha_splitalphas-mistral-basesnli-alphastreet2
Custom SNLI Dataset
A custom SNLI-style dataset generated using GPT-3.5
Dataset Structure
Data Instances
Each instance contains:
premise: The initial statement
hypothesis: A generated statement
label: The relationship between premise and hypothesis (0: entailment, 1: contradiction, 2: neutral)
Data Fields
premise: string
hypothesis: string
label: int
Data Splits
train: 134 examples
validation: 16 examples
test: 18 examples
TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/alphass123/TinyStories.alphas2pdfalphas3alphasalphastar
