datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PAMIPTDC_pampa_ncats
TDC PAMPA NCATS
PAMPA NCATS dataset [1], part of TDC [2] benchmark. It is intended to be used through
scikit-fingerprints library.
NCATS subset of PAMPA dataset.
PAMPA (parallel artificial membrane permeability assay) is an assay to evaluate drug permeability across the cellular membrane. The task models only the passive membrane diffusion. This is “NCATS” subset of the dataset created at National Center for Advancing Translational Sciences (NCATS).
This dataset is a part of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_pampa_ncats.pamela
PAM∃LA
Personalizing Text-to-Image Generation to Individual Taste
Anonymous submission — author and affiliation details withheld during review.
PAM∃LA is a dataset of AI-generated images rated by human participants for aesthetic quality, built specifically for personalization research. It pairs each rating with rich participant demographics and image metadata, enabling research on personalized aesthetic prediction, demographic variation in visual preference, and reward modelling for… See the full description on the dataset page: https://huggingface.co/datasets/pamela-dataset/pamela.PAM-data
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Perceive Anything Model (PAM) is a conceptually simple and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations… See the full description on the dataset page: https://huggingface.co/datasets/Perceive-Anything/PAM-data.pa-minimal-green-team-SFT-500m
pa-minimal-green-team-SFT-500m
The PA green-team minimal 500M SFT mix — math + science + light agentic, short-reasoning-first.
Two training configs (same 216,668 documents, same order):
config
tokens
what
math_sci_agentic_500m
500.0M
reasoning traces intact (think)
math_sci_agentic_500m_nothink
131.8M
every trace replaced by a literal empty <think></think> (no-think)
Composition: 37.7% math_reasoning · 36.3% science_mcq · 16.7% science_research · 9.4%… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-minimal-green-team-SFT-500m.pampa_ncats-multimodalswe-gym-trainswe-gym-test-getmotoswe-gym-testolivers-mtor-atlas
Oliver's mTOR Atlas
The mTOR pathway, mapped by what the evidence can actually carry. This dataset is the curated corpus behind mtor-atlas.org: 411 hand-selected studies on mTOR (mechanistic target of rapamycin) signalling, each labelled by the kind of study behind it, and a list of 149 pathway entities (genes and proteins, complexes, drugs, interventions, biological processes, diseases, outcomes, organelles, nutrients and conditions) that the studies refer to.
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/pampalini1/olivers-mtor-atlas.redred-gemma-4-E2B-it-lora-summariespampa_ncats
Dataset Details
Dataset Description
PAMPA (parallel artificial membrane permeability assay) is a commonly
employed assay to evaluate drug permeability across the cellular membrane.
PAMPA is a non-cell-based, low-cost and high-throughput alternative to cellular models.
Although PAMPA does not model active and efflux transporters, it still provides permeability values
that are useful for absorption prediction because the majority of drugs are absorbed
by passive diffusion… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/pampa_ncats.analise_sentimento_ptbr_rewiews_ecommercemerged_final1_cleanedADME-Property-Prediction-PAMPA_NCATSUZ-STThis is the official dataset repo for UZ-ST in Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance.
OR-PAM-Reg-4K
📢 当前版本:V3(Latest) / Current Version: V3 (Latest)
本数据集当前为 V3 版本,所有先前版本(V1、V2 等)已全部废弃,不再具有任何合法授权。若您持有旧版本数据,请立即停止使用并删除。
All previous versions (V1, V2, etc.) have been deprecated and revoked. If you hold any earlier version, please cease use and delete it immediately.
🔒 访问与使用规范 / Access & Usage Policy
本数据集欢迎学术研究人员申请使用。为防止违规行为,请注意以下规定:
禁止再分发:未经作者书面授权,不得将本数据集(或其任何部分)以任何形式重新分发、上传至其他平台或共享给第三方。
禁止二次生成与衍生数据集发布:不得利用本数据集生成新的数据集并公开发布,除非获得作者明确书面许可。
仅限学术研究用途:本数据集仅授权用于非商业学术研究,任何商业用途均需另行授权。… See the full description on the dataset page: https://huggingface.co/datasets/chengliuyan/OR-PAM-Reg-4K.duckad-data
DuckAD Driving Dataset
Expert driving demonstrations for DuckAD, an end-to-end vision-based driving model,
collected in CARLA on custom Duckietown-style maps. A rule-aware expert
driver was rolled out under six traffic/obstacle scenarios; every frame pairs a front camera
image with the expert's future trajectory, a high-level navigation command, and ground-truth
bird's-eye-view (BEV) semantics.
214,200 training frames from two maps (duckietown_04, duckietown_05), six scenarios… See the full description on the dataset page: https://huggingface.co/datasets/pamasan/duckad-data.Wikihow_sinhalaOpus_100_en_ptAnroidwadakaroyoinside-out-replication-v2-pamqfix-metrics-v1
Inside-Out Replication V2 — A.8.1 in-context P(a|q) re-score (FULL)
3 models × 4 relations × test split, all 12 cells. P(a|q)/P_norm rebuilt with
paper-faithful A.8.1 in-context tokenization (the answer is tokenized jointly
with the generation-prompt context, not naive-concatenated). P(True) and
verifier variants unchanged (their code path is unaffected). Probe is
DEFERRED in this phase — externals only.
Headline finding: the A.8.1 fix moves K by ≤ 0.001 everywhere… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-pamqfix-metrics-v1.arcanum-cross-platform-queries-syntheticScrape_finalinside-out-replication-v2-pamqfix-scores-full
Inside-Out Replication V2 — full corrected raw scores (all 12 cells)
This is the dataset to recompute K and K* for P(True), P(a|q),
P_norm(a|q) (and the sensitivity variants) across all 3 models × 4
relations (12 "cells"), test split.
One row per unique (question, answer) candidate — the greedy answer + the
1,000 temperature-1 samples are deduplicated to unique strings, then scored
once each. ~1.57M rows. Pipeline: judge-bug-fixed labels (pure unique-verdict
"Scheme A") + A.8.1… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-pamqfix-scores-full.pambanudesh_sinhalaPamela_Skylinepa-moduli-italiabbc_sinhala_csv
