sapiens
Datasets
All datasets matching “sapiens”Perturb-Sapiens
Perturb Sapiens: A Human Whole-Organism Atlas of Perturbed Cells
Dataset Description
Perturb Sapiens is an evolving database of AI-predicted single-cell perturbation responses, representing the first human whole-organism atlas of perturbed cells.
Perturb Sapiens is generated using the post-trained Stack model (Stack-Large-Aligned), an in-context learning foundation model for single-cell biology.
Data Sources:
Prompt Data: Parse/OpenProblems PBMC perturbation data
Query… See the full description on the dataset page: https://huggingface.co/datasets/arcinstitute/Perturb-Sapiens.PartNetMobility
PartNet-Mobility Dataset
PartNet-Mobility dataset is a collection of 2K articulated objects with motion annotations and rendering material. The dataset powers research for generalizable computer vision and manipulation. The dataset is a continuation of ShapeNet and PartNet.
The dataset is compatible with the SAPIEN simulator, a realistic and physics-rich simulated environment that hosts a large-scale set for articulated objects. SAPIEN enables various robotic vision and… See the full description on the dataset page: https://huggingface.co/datasets/sapien-sim/PartNetMobility.CF-MS_Homo_sapiens_PPI
CF-MS Elution Profile PPI Dataset
Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.gpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
context_length_benchmarking
🧠 Context Length - Benchmarking
A Mathematical Framework for Long-Context Attention Evaluation
The Context Length Benchmarking, developed by Sapiens Technology®, is a deterministic and scalable framework designed to evaluate how effectively large language models retain and retrieve information across extremely long contexts, isolating pure attention capability by removing semantic complexity and focusing on distributed anomaly detection; the methodology involves normalizing the… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/context_length_benchmarking.global_mmlu_lite_pt
🌎 Global-MMLU Lite (Portuguese)
A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models
Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.
