datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HiPHI
HiPHI: A large-scale benchmark for high-precision human motion and object interaction.
Project page: https://noitom-robotics.github.io/hiphi/
Online viewer: https://hiphi-viewer.modalitynet.com/
Paper: http://arxiv.org/abs/2608.16222
GitHub: https://github.com/noitom-robotics/hiphi/
HiPHI is an optical motion-capture dataset for humanoid learning and
whole-body motion modeling. It provides standardized BVH motion and, for
human-object interaction… See the full description on the dataset page: https://huggingface.co/datasets/noitomrobotics/HiPHI.HippoCamp
HippoCamp: Benchmarking Contextual Agents on Personal Computers
📖 Paper |
🏠 Project Page |
🛠️ GitHub |
🤗 Dataset |
🎬 Demo
Overview
HippoCamp is a benchmark for evaluating contextual agents in realistic, device-resident personal computing environments. Unlike agent benchmarks centered on web interaction, tool use, or generic software automation, HippoCamp focuses on multimodal file management over large personal file systems: agents must… See the full description on the dataset page: https://huggingface.co/datasets/MMMem-org/HippoCamp.hippoFakeParts
FakeParts: A New Family of AI-Generated DeepFakes
Abstract
We introduce FakeParts, a new class of deepfakes characterized by subtle, localized manipulations to specific spatial regions or temporal segments of otherwise authentic videos. Unlike fully synthetic content, these partial manipulations—ranging from altered facial expressions to object substitutions and background modifications—blend seamlessly with real elements, making them particularly deceptive… See the full description on the dataset page: https://huggingface.co/datasets/hi-paris/FakeParts.HippoRAG_2
HippoRAG 2 is a powerful memory framework for LLMs that enhances their ability to recognize and utilize connections in new knowledge—mirroring a key function of human long-term memory.HIP-UMA-OMol25
HIP-UMA-OMol25
Snapshot of completed UMA-S-1.2 (uma-s-1p2) predictions on a uniform
sample of OMol-1 train geometries, using the omol task.
This snapshot contains only validated, complete HDF5 shards. Checkpoint
files (*.partial.h5) are intentionally excluded.
Current snapshot (2026-08-21): 10846 complete shards, about 25 GB,
4.3 million samples. The full 10M campaign is still running (41911 shards planned).
Shard files are grouped to stay under Hugging Face's 10… See the full description on the dataset page: https://huggingface.co/datasets/andreasburger/HIP-UMA-OMol25.HiPhO
🥇 HiPhO: High School Physics Olympiad Benchmark
[🏆 Leaderboard]
[📊 Dataset]
[✨ GitHub]
[📄 Paper]
📊 New (Dec. 8): Results from Gemini-3-Pro, DeepSeek-V3.2-Speciale, and Kimi-K2-Thinking have been added to the HiPhO Leaderboard. Notably, Gemini-3-Pro achieved gold-medal performance across all 13 Olympiads in HiPhO.
🧩 New (Nov. 5):We added CPhO 2025 (Chinese Physics Olympiad) — the national final theoretical exam to the HiPhO benchmark.
🏆 New (Sep. 16): We launched "PhyArena", a… See the full description on the dataset page: https://huggingface.co/datasets/SciYu/HiPhO.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.hi-proverbs-cpt
Hindi Proverbs — Idiom-Tagged Continued-Pretraining Corpus
A 338K-document Hindi corpus (~1.4B tokens) for continued pretraining on
cultural knowledge in figurative language, plus a structured dataset of
16,617 Hindi proverbs (लोकोक्तियाँ) with meanings, recovered via
OCR-repair from a classic proverb dictionary. Each corpus document is natural
Hindi text containing at least one proverb (matched including common surface
variants), with an appended knowledge block listing every… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/hi-proverbs-cpt.HippoVlogdetails_tushar310__Hippy-AAI-7B
Dataset Card for Evaluation run of tushar310/Hippy-AAI-7B
Dataset automatically created during the evaluation run of model tushar310/Hippy-AAI-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tushar310__Hippy-AAI-7B.hippo-lora-story-suite
Hippocampus LoRA story suite
Five controlled analyses of a video-conditioned, generated LoRA adapter (the C2L "context-to-LoRA"
mechanism), run against real OVO-Bench video clips. This is not a main-table benchmark
reproduction. It is a set of curated, small-N probes designed to ask one question each:
Kind
Question
Cases in this run
paired_history
Same current input, different history: does the answer track the intended history?
3
retention
Does support for a fact… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/hippo-lora-story-suite.airr_hip
HIP — TCRβ repertoires with CMV serostatus and HLA typing
Bulk TCRβ immunosequencing of 786 subjects with known cytomegalovirus (CMV) serostatus
and HLA-A/HLA-B typing — the cohort of Emerson et al. 2017 (666 discovery + 120 validation
subjects). Provided as a benchmark for reproducing HLA- and CMV/HLA-associated TCR biomarker
discovery: learning public TCRβ sequences that predict CMV exposure and infer HLA alleles.
According to PubMed, the source study is Emerson RO, DeWitt WS… See the full description on the dataset page: https://huggingface.co/datasets/isalgo/airr_hip.autotrain-flair-hipe2022-de-hmbert
NER Fine-Tuning
We use Flair for fine-tuning NER models on
HIPE-2022 datasets from
HIPE-2022 Shared Task.
All models are fine-tuned on A10 (24GB) and A100 (40GB) instances from
Lambda Cloud using Flair:
$ git clone https://github.com/flairNLP/flair.git
$ cd flair && git checkout 419f13a05d6b36b2a42dd73a551dc3ba679f820c
$ pip3 install -e .
$ cd ..
Clone this repo for fine-tuning NER models:
$ git clone https://github.com/stefan-it/hmTEAMS.git
$ cd hmTEAMS/bench
Authorize via Hugging… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/autotrain-flair-hipe2022-de-hmbert.HiPhO
🥇 HiPhO: High School Physics Olympiad Benchmark
[🏆 Leaderboard]
[📊 Dataset]
[✨ GitHub]
[📄 Paper]
🏆 New (Sep. 16): We launched "PhyArena", a physics reasoning leaderboard incorporating the HiPhO benchmark.
🌐 Introduction
HiPhO (High School Physics Olympiad Benchmark) is the first benchmark specifically designed to evaluate the physical reasoning abilities of (M)LLMs on real-world Physics Olympiads from 2024–2025.
✨ Key Features
Up-to-date Coverage:… See the full description on the dataset page: https://huggingface.co/datasets/HY-Wan/HiPhO.hippocorpusTo examine the cognitive processes of remembering and imagining and their traces in language, we introduce Hippocorpus, a dataset of 6,854 English diary-like short stories about recalled and imagined events. Using a crowdsourcing framework, we first collect recalled stories and summaries from workers, then provide these summaries to other workers who write imagined stories. Finally, months later, we collect a retold version of the recalled stories from a subset of recalled authors. Our dataset comes paired with author demographics (age, gender, race), their openness to experience, as well as some variables regarding the author's relationship to the event (e.g., how personal the event is, how often they tell its story, etc.).autotrain-flair-hipe2022-fr-hmbert
NER Fine-Tuning
We use Flair for fine-tuning NER models on
HIPE-2022 datasets from
HIPE-2022 Shared Task.
All models are fine-tuned on A10 (24GB) and A100 (40GB) instances from
Lambda Cloud using Flair:
$ git clone https://github.com/flairNLP/flair.git
$ cd flair && git checkout 419f13a05d6b36b2a42dd73a551dc3ba679f820c
$ pip3 install -e .
$ cd ..
Clone this repo for fine-tuning NER models:
$ git clone https://github.com/stefan-it/hmTEAMS.git
$ cd hmTEAMS/bench
Authorize via Hugging… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/autotrain-flair-hipe2022-fr-hmbert.hippo-v1_4_7_reasoning-qwen35-122bhipsc_2dhippo-paper-lora-visualization
Hippocampus paper LoRA visualization: v1_4_7_2xlr on VideoMME
This repo holds every output of evaluation/lora_visualization on the paper checkpoint v1_4_7_2xlr/checkpoint-2800, in two runs. The thumbnails are frames from VideoMME videos. Follow the VideoMME terms of use when you reuse them.
Folder
Run
20260916T225013Z/
all 594 eligible videos x 3 questions = 1782 questions, 16 GPU shards
20260916T203455Z/
the 2-video smoke, the fixed 24-video pilot, the failed… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/hippo-paper-lora-visualization.hippo-v1_4_7-cot-check
Is the v1.4.7 reasoning true to the video?
70 timelens_streaming turns for eyeballing against the clip. The text
judge that produced judge_* never saw the video, so it cannot catch reasoning
that invents visual detail and still lands on the right answer. That is what this
sample is for.
Four strata in bucket, sampled at random (seed 0):
bucket
population
what to look for
keep_withGT
177,353
the judge kept these. Does the reasoning describe what is actually on screen?… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/hippo-v1_4_7-cot-check.hiphiFakeParts_Legacy
FakeParts: A New Family of AI-Generated DeepFakes
Abstract
We introduce FakeParts, a new class of deepfakes characterized by subtle, localized manipulations to specific spatial regions or temporal segments of otherwise authentic videos. Unlike fully synthetic content, these partial manipulations, ranging from altered facial expressions to object substitutions and background modifications, blend seamlessly with real elements, making them particularly… See the full description on the dataset page: https://huggingface.co/datasets/hi-paris/FakeParts_Legacy.msd-hippocampus
Medical Segmentation Decathlon: Hippocampus
Dataset Description
This is the Hippocampus dataset from the Medical Segmentation Decathlon (MSD) challenge. The dataset contains MRI scans with segmentation annotations for hippocampus segmentation.
Dataset Details
Modality: MRI
Task: Task04_Hippocampus
Target: anterior and posterior hippocampus
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/msd-hippocampus.hipct_3dHIPE2020_sent-splitTODOhipe2020TODOiclr2026-lm-logprobs
LM Log-Probabilities for Value Bias Analysis
Next-token log-probability distributions from 12 language models across 54 prompts, used in the paper:
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A.F. Thompson, Elle, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska (ICLR 2026)
Part of the Oxford-HIPlab collection for this paper.
Dataset description
Each CSV contains the full next-token log-probability… See the full description on the dataset page: https://huggingface.co/datasets/Oxford-HIPlab/iclr2026-lm-logprobs.HiPaS_datasetThis dataset contains created nifti files for the dataset as mentioned in the paper "Deep learning-driven pulmonary artery and vein segmentation reveals demography-associated vasculature anatomical differences". They also released the code and datasets at https://github.com/Arturia-Pendragon-Iris/HiPaS_AV_Segmentation.
If you find this dataset be helpful to you, please city their paper
"Chu, Y., Luo, G., Zhou, L. et al. Deep learning-driven pulmonary artery and vein segmentation reveals… See the full description on the dataset page: https://huggingface.co/datasets/JasperEppink/HiPaS_dataset.audio-diffusion-instrumental-hiphop-256256x256 mel spectrograms of 5 second samples of instrumental Hip Hop. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 256
y_res = 256
sample_rate = 22050
n_fft = 2048
hop_length = 512
