datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AMPBench-MT
AMPBench-MT
AMPBench-MT is a homology-controlled benchmark for antimicrobial peptide endpoint prediction. The release is dated 2026-07-08.
Repository: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT
The benchmark is organized around endpoint-aware prediction rather than binary AMP recognition alone. It contains processed task tables for AMP/non-AMP classification, species-conditioned MIC regression, activity spectrum positive-evidence audits, low-toxicity classification… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.Historical_Nifty_50_Constituent_Weights_20YSUMMARY & CONTEXT:
This dataset aims to provide a comprehensive, rolling 20-year history of the constituent stocks and their corresponding weights in India's Nifty 50 index. The data begins on January 31, 2008, and is actively maintained with monthly updates. After hitting the 20-year mark, as new monthly data is added, the oldest month's data will be removed to maintain a consistent 20-year window. This dataset was developed as a foundational feature for a graph-based model analyzing the… See the full description on the dataset page: https://huggingface.co/datasets/AMP4010/Historical_Nifty_50_Constituent_Weights_20Y.ample-math
AMPLE-Math
5,319 mathematics problems, each with a verified final answer and six references to that same
answer. The references differ only in how much of the reasoning they show, which makes them useful
for studying what a teacher's reference content contributes during distillation.
Problems and original reasoning come from the metadata configuration of
OpenThoughts-114k, and keep its
Apache-2.0 attribution. A question was kept only if all six references exist, every generated… See the full description on the dataset page: https://huggingface.co/datasets/xiuyuz/ample-math.owm-rm-3.2mannual_reports_us.ko.jaowm-rm-645klyu-wang-balius-singh-2019-ampc
Ultra-large docking data: AmpC 96M compounds
These data are from John J. Irwin, Bryan L. Roth, and Brian K. Shoichet's labs. They published it as:
[!NOTE]Lyu J, Wang S, Balius TE, Singh I, Levit A, Moroz YS, O'Meara MJ, Che T, Algaa E, Tolmachova K, Tolmachev AA, Shoichet BK, Roth BL, Irwin JJ.
Ultra-large library docking for discovering new chemotypes. Nature. 2019 Feb;566(7743):224-229. doi: 10.1038/s41586-019-0917-9.
Epub 2019 Feb 6. PMID: 30728502; PMCID: PMC6383769.… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/lyu-wang-balius-singh-2019-ampc.math-intuition-reasoning-traces
math-intuition reasoning traces
Full chain-of-thought traces from 7 reasoning models on the same 4,020 problems, graded
by each problem family's own verifier.
Questions come from
amphora/math-intuition-20260908-402-easy-10
— 402 arXiv-derived problem families x 10 seeds, easy preset. Every row here refers to an id
in that dataset, so prompts and the instance cache can be joined from it.
Generation settings
Identical for every model, so the traces are directly… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-reasoning-traces.Ultrafeedback-Mistral-Instruct-AMPOmath-intuition-20260906-403-demo-10
math-intuition-20260906-403-demo-10
3,936 mathematics problems drawn from 403 problem families, each derived from a
distinct arXiv paper. Every problem is generated answer-first, so the answer is known by
construction and is checked by the family's own verify() before the row is written.
No row in this file is ungraded.
This is the demo rung — read this before using it
Each family exposes a four-rung ladder: demo, easy, medium, hard. This file samples
demo, which… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-demo-10.math-intuition-20260906-403-easy-30
math-intuition-20260906-403-easy-30
12,090 synthetic mathematics problems drawn from 403 problem families, each family
derived from a distinct arXiv paper. Every problem is generated answer-first, so the
answer is known by construction and is checked by the family's own verify() before
the row is written. No row in this file is ungraded.
This is the easy slice: 30 instances per family at each family's easiest difficulty
preset. It is not the hard benchmark — see Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/amphora/math-intuition-20260906-403-easy-30.LSD_AmpC_Liu_2025owm-rm-1mUltrafeedback-Llama3-8B-Instruct-AMPOko-r1-v3Ultrafeedback-Gemma-Instruct-AMPOquantization-cache-amplification
Quantization as Cache Amplification
Trillion-Parameter Mixture-of-Experts Inference on a Commodity Laptop
Kavin Kumar, Neural Metrics
📄 Read the paper — 11 pages
What this is
Weight quantization is usually justified as footprint reduction. This work argues
that for offloaded mixture-of-experts inference that framing misses the leverage.
The binding resource is not storage capacity but the fraction of expert slots
resident in DRAM — and storage traffic depends on… See the full description on the dataset page: https://huggingface.co/datasets/mkvn/quantization-cache-amplification.Ultrafeedback-Gemma-Instruct-AMPO-finafrica-synth-disability-amputation-prosthetics-all
Amputation & Prosthetics (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-disability-amputation-prosthetics-all.AMPO-OPT-Selectionamp_dataset_viri
Dataset Card for "amp_dataset_viri"
More Information needed
eval_so101_act_pick_green_cube_ampThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/annyi/eval_so101_act_pick_green_cube_amp.lmsys-finance
Dataset Card for "lmsys-finance"
This dataset is a curated version of the lmsys-chat-1m dataset,
focusing solely on finance-related conversations. The refinement process encompassed:
Removing non-English conversations.
Selecting conversations from models: "vicuna-33b", "wizardlm-13b", "gpt-4", "gpt-3.5-turbo", "claude-2", "palm-2", and "claude-instant-1".
Excluding conversations with responses under 30 characters.
Using 100 financial keywords, choosing conversations with at least… See the full description on the dataset page: https://huggingface.co/datasets/amphora/lmsys-finance.m-math500jacobian-research-trajectory-batch-001
Jacobian research trajectory — batch 001
A flattened archival extraction of research attempts concerning polynomial maps in two variables with nonzero constant Jacobian determinant, using Research Trajectory Schema v1.0.0.
Contents
research_trajectory_flat.csv contains 207 records and 145 columns. Each row represents one canonical record: 1 problem, 24 branches, 55 attempts, 56 evaluations, 19 relations, 3 decisions, or 49 artifacts.
Nested object fields use… See the full description on the dataset page: https://huggingface.co/datasets/amphora/jacobian-research-trajectory-batch-001.owm-train-200kclinical-diagnostic-inference-error-amplification-mapping-v0.1What this dataset tests
How small inference errors introduced at a decision nodeamplify into downstream diagnostic distortion.
Required outputs
error entry node
inference error type
amplification factor
downstream distortion map
delay and misdiagnosis probabilities
self-correction points
prevention guardrails
qwen-sc-it2instructkr-ko-arena-0407-abit-cleanedafrica-mauritius-location-of-hypermarkets-amp-supermarkets-in-mauritius-0069f8ea
Location of Hypermarkets Amp Supermarkets in Mauritius | Africa (MDPA)
142 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 142 rows from MDPA, covering Location of Hypermarkets Amp Supermarkets in Mauritius. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-location-of-hypermarkets-amp-supermarkets-in-mauritius-0069f8ea.
