datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.DataComp_large_pool_BLIP2_captions
Dataset Card for DataComp_large_pool_BLIP2_captions
Dataset Summary
Supported Tasks and Leaderboards
We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp.
Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work.
Languages
Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_large_pool_BLIP2_captions.lotte_pooled_colbertv2DataComp_medium_pool_BLIP2_captions
Dataset Card for DataComp_medium_pool_BLIP2_captions
Dataset Summary
Supported Tasks and Leaderboards
We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp.
Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work.
Languages
Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_medium_pool_BLIP2_captions.permutation-pools
Permutation Pools
Permutation-augmentation pools: demonstration-order permutations of the label-transfer and evidence-tracking sets, used to enlarge filtered test sets to 1000 rows with balanced golds.
Part of the icl-heads collection: the datasets behind a study of which
Llama-3.1-8B-Instruct attention heads mediate in-context evidence accumulation
and in-context label mapping, discovered with Differentiable Circuit Masking
(a learned per-head mask over K/V activations patched… See the full description on the dataset page: https://huggingface.co/datasets/icl-heads/permutation-pools.minor-pool-9b57ec
minor-pool-9b57ec
Synthetic sensors test data: 55 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/MaryGonzalez/minor-pool-9b57ec.liquidity-pool-performance-benchmarks
