CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01noitomrobotics /HiPHIgated HiPHI: A large-scale benchmark for high-precision human motion and object interaction. Project page: https://noitom-robotics.github.io/hiphi/ Online viewer: https://hiphi-viewer.modalitynet.com/ Paper: http://arxiv.org/abs/2608.16222 GitHub: https://github.com/noitom-robotics/hiphi/ HiPHI is an optical motion-capture dataset for humanoid learning and whole-body motion modeling. It provides standardized BVH motion and, for human-object interaction… See the full description on the dataset page: https://huggingface.co/datasets/noitomrobotics/HiPHI.10K<n<100K42 likes4.9k downloads11d agoHugging Face02MMMem-org /HippoCamp HippoCamp: Benchmarking Contextual Agents on Personal Computers 📖 Paper | 🏠 Project Page | 🛠️ GitHub | 🤗 Dataset | 🎬 Demo Overview HippoCamp is a benchmark for evaluating contextual agents in realistic, device-resident personal computing environments. Unlike agent benchmarks centered on web interaction, tool use, or generic software automation, HippoCamp focuses on multimodal file management over large personal file systems: agents must… See the full description on the dataset page: https://huggingface.co/datasets/MMMem-org/HippoCamp.documentquestion-answeringn<1K6 likes4.8k downloads6mo agoHugging Face03rogerwallace22463 /hippo0 likes2k downloads11m agoHugging Face04hi-paris /FakeParts FakeParts: A New Family of AI-Generated DeepFakes Abstract We introduce FakeParts, a new class of deepfakes characterized by subtle, localized manipulations to specific spatial regions or temporal segments of otherwise authentic videos. Unlike fully synthetic content, these partial manipulations—ranging from altered facial expressions to object substitutions and background modifications—blend seamlessly with real elements, making them particularly deceptive… See the full description on the dataset page: https://huggingface.co/datasets/hi-paris/FakeParts.textvideo-classification10K<n<100K2 likes1.1k downloads9mo agoHugging Face05osunlp /HippoRAG_2 HippoRAG 2 is a powerful memory framework for LLMs that enhances their ability to recognize and utilize connections in new knowledge—mirroring a key function of human long-term memory.11 likes945 downloads1y agoHugging Face06andreasburger /HIP-UMA-OMol25 HIP-UMA-OMol25 Snapshot of completed UMA-S-1.2 (uma-s-1p2) predictions on a uniform sample of OMol-1 train geometries, using the omol task. This snapshot contains only validated, complete HDF5 shards. Checkpoint files (*.partial.h5) are intentionally excluded. Current snapshot (2026-08-21): 10846 complete shards, about 25 GB, 4.3 million samples. The full 10M campaign is still running (41911 shards planned). Shard files are grouped to stay under Hugging Face's 10… See the full description on the dataset page: https://huggingface.co/datasets/andreasburger/HIP-UMA-OMol25.graph-ml0 likes627 downloads1mo agoHugging Face07SciYu /HiPhO 🥇 HiPhO: High School Physics Olympiad Benchmark [🏆 Leaderboard] [📊 Dataset] [✨ GitHub] [📄 Paper] 📊 New (Dec. 8): Results from Gemini-3-Pro, DeepSeek-V3.2-Speciale, and Kimi-K2-Thinking have been added to the HiPhO Leaderboard. Notably, Gemini-3-Pro achieved gold-medal performance across all 13 Olympiads in HiPhO. 🧩 New (Nov. 5):We added CPhO 2025 (Chinese Physics Olympiad) — the national final theoretical exam to the HiPhO benchmark. 🏆 New (Sep. 16): We launched "PhyArena", a… See the full description on the dataset page: https://huggingface.co/datasets/SciYu/HiPhO.imagequestion-answeringn<1K5 likes494 downloads10mo agoHugging Face08YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes387 downloads4mo agoHugging Face09jiviteshjn /hi-proverbs-cpt Hindi Proverbs — Idiom-Tagged Continued-Pretraining Corpus A 338K-document Hindi corpus (~1.4B tokens) for continued pretraining on cultural knowledge in figurative language, plus a structured dataset of 16,617 Hindi proverbs (लोकोक्तियाँ) with meanings, recovered via OCR-repair from a classic proverb dictionary. Each corpus document is natural Hindi text containing at least one proverb (matched including common surface variants), with an appended knowledge block listing every… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/hi-proverbs-cpt.text-generation100K<n<1M0 likes334 downloads2mo agoHugging Face10JakeTian /HippoVlogimage10K<n<100K0 likes299 downloads6mo agoHugging Face11open-llm-leaderboard-old /details_tushar310__Hippy-AAI-7B Dataset Card for Evaluation run of tushar310/Hippy-AAI-7B Dataset automatically created during the evaluation run of model tushar310/Hippy-AAI-7B on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tushar310__Hippy-AAI-7B.0 likes257 downloads3y agoHugging Face12dipta007 /hippo-lora-story-suite Hippocampus LoRA story suite Five controlled analyses of a video-conditioned, generated LoRA adapter (the C2L "context-to-LoRA" mechanism), run against real OVO-Bench video clips. This is not a main-table benchmark reproduction. It is a set of curated, small-N probes designed to ask one question each: Kind Question Cases in this run paired_history Same current input, different history: does the answer track the intended history? 3 retention Does support for a fact… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/hippo-lora-story-suite.video-text-to-text0 likes252 downloads4d agoHugging Face13isalgo /airr_hip HIP — TCRβ repertoires with CMV serostatus and HLA typing Bulk TCRβ immunosequencing of 786 subjects with known cytomegalovirus (CMV) serostatus and HLA-A/HLA-B typing — the cohort of Emerson et al. 2017 (666 discovery + 120 validation subjects). Provided as a benchmark for reproducing HLA- and CMV/HLA-associated TCR biomarker discovery: learning public TCRβ sequences that predict CMV exposure and infer HLA alleles. According to PubMed, the source study is Emerson RO, DeWitt WS… See the full description on the dataset page: https://huggingface.co/datasets/isalgo/airr_hip.texttabular-classification100M<n<1B1 likes242 downloads2mo agoHugging Face14stefan-it /autotrain-flair-hipe2022-de-hmbert NER Fine-Tuning We use Flair for fine-tuning NER models on HIPE-2022 datasets from HIPE-2022 Shared Task. All models are fine-tuned on A10 (24GB) and A100 (40GB) instances from Lambda Cloud using Flair: $ git clone https://github.com/flairNLP/flair.git $ cd flair && git checkout 419f13a05d6b36b2a42dd73a551dc3ba679f820c $ pip3 install -e . $ cd .. Clone this repo for fine-tuning NER models: $ git clone https://github.com/stefan-it/hmTEAMS.git $ cd hmTEAMS/bench Authorize via Hugging… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/autotrain-flair-hipe2022-de-hmbert.0 likes210 downloads3y agoHugging Face15HY-Wan /HiPhO 🥇 HiPhO: High School Physics Olympiad Benchmark [🏆 Leaderboard] [📊 Dataset] [✨ GitHub] [📄 Paper] 🏆 New (Sep. 16): We launched "PhyArena", a physics reasoning leaderboard incorporating the HiPhO benchmark. 🌐 Introduction HiPhO (High School Physics Olympiad Benchmark) is the first benchmark specifically designed to evaluate the physical reasoning abilities of (M)LLMs on real-world Physics Olympiads from 2024–2025. ✨ Key Features Up-to-date Coverage:… See the full description on the dataset page: https://huggingface.co/datasets/HY-Wan/HiPhO.imagequestion-answeringn<1K1 likes199 downloads10mo agoHugging Face16allenai /hippocorpusTo examine the cognitive processes of remembering and imagining and their traces in language, we introduce Hippocorpus, a dataset of 6,854 English diary-like short stories about recalled and imagined events. Using a crowdsourcing framework, we first collect recalled stories and summaries from workers, then provide these summaries to other workers who write imagined stories. Finally, months later, we collect a retold version of the recalled stories from a subset of recalled authors. Our dataset comes paired with author demographics (age, gender, race), their openness to experience, as well as some variables regarding the author's relationship to the event (e.g., how personal the event is, how often they tell its story, etc.).text-classification1K<n<10K5 likes189 downloads3y agoHugging Face17stefan-it /autotrain-flair-hipe2022-fr-hmbert NER Fine-Tuning We use Flair for fine-tuning NER models on HIPE-2022 datasets from HIPE-2022 Shared Task. All models are fine-tuned on A10 (24GB) and A100 (40GB) instances from Lambda Cloud using Flair: $ git clone https://github.com/flairNLP/flair.git $ cd flair && git checkout 419f13a05d6b36b2a42dd73a551dc3ba679f820c $ pip3 install -e . $ cd .. Clone this repo for fine-tuning NER models: $ git clone https://github.com/stefan-it/hmTEAMS.git $ cd hmTEAMS/bench Authorize via Hugging… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/autotrain-flair-hipe2022-fr-hmbert.0 likes164 downloads3y agoHugging Face18JonesLin /hippo-v1_4_7_reasoning-qwen35-122b0 likes161 downloads15d agoHugging Face19eternalaudrey /hipsc_2dimage10K<n<100K0 likes158 downloads1y agoHugging Face20dipta007 /hippo-paper-lora-visualization Hippocampus paper LoRA visualization: v1_4_7_2xlr on VideoMME This repo holds every output of evaluation/lora_visualization on the paper checkpoint v1_4_7_2xlr/checkpoint-2800, in two runs. The thumbnails are frames from VideoMME videos. Follow the VideoMME terms of use when you reuse them. Folder Run 20260916T225013Z/ all 594 eligible videos x 3 questions = 1782 questions, 16 GPU shards 20260916T203455Z/ the 2-video smoke, the fixed 24-video pilot, the failed… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/hippo-paper-lora-visualization.0 likes138 downloads5d agoHugging Face21JonesLin /hippo-v1_4_7-cot-check Is the v1.4.7 reasoning true to the video? 70 timelens_streaming turns for eyeballing against the clip. The text judge that produced judge_* never saw the video, so it cannot catch reasoning that invents visual detail and still lands on the right answer. That is what this sample is for. Four strata in bucket, sampled at random (seed 0): bucket population what to look for keep_withGT 177,353 the judge kept these. Does the reasoning describe what is actually on screen?… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/hippo-v1_4_7-cot-check.videon<1K0 likes135 downloads15d agoHugging Face22introvoyz042 /hiphi3dn<1K0 likes135 downloads12d agoHugging Face23hi-paris /FakeParts_Legacy FakeParts: A New Family of AI-Generated DeepFakes Abstract We introduce FakeParts, a new class of deepfakes characterized by subtle, localized manipulations to specific spatial regions or temporal segments of otherwise authentic videos. Unlike fully synthetic content, these partial manipulations, ranging from altered facial expressions to object substitutions and background modifications, blend seamlessly with real elements, making them particularly… See the full description on the dataset page: https://huggingface.co/datasets/hi-paris/FakeParts_Legacy.videovideo-classificationn<1K8 likes132 downloads9mo agoHugging Face24MedOtter /msd-hippocampus Medical Segmentation Decathlon: Hippocampus Dataset Description This is the Hippocampus dataset from the Medical Segmentation Decathlon (MSD) challenge. The dataset contains MRI scans with segmentation annotations for hippocampus segmentation. Dataset Details Modality: MRI Task: Task04_Hippocampus Target: anterior and posterior hippocampus Format: NIfTI (.nii.gz) Dataset Structure Each sample in the JSONL file contains: { "image":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/msd-hippocampus.textimage-segmentationn<1K0 likes110 downloads11mo agoHugging Face25eternalaudrey /hipct_3dimage10K<n<100K0 likes106 downloads11mo agoHugging Face26bigscience-historical-texts /HIPE2020_sent-splitTODO0 likes105 downloads4y agoHugging Face27bigscience-historical-texts /hipe2020TODO3 likes94 downloads4y agoHugging Face28Oxford-HIPlab /iclr2026-lm-logprobs LM Log-Probabilities for Value Bias Analysis Next-token log-probability distributions from 12 language models across 54 prompts, used in the paper: Reward Models Inherit Value Biases from Pretraining Brian Christian, Jessica A.F. Thompson, Elle, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, Tsvetomira Dumbalska (ICLR 2026) Part of the Oxford-HIPlab collection for this paper. Dataset description Each CSV contains the full next-token log-probability… See the full description on the dataset page: https://huggingface.co/datasets/Oxford-HIPlab/iclr2026-lm-logprobs.tabulartext-generation1M<n<10M0 likes93 downloads7mo agoHugging Face29JasperEppink /HiPaS_datasetThis dataset contains created nifti files for the dataset as mentioned in the paper "Deep learning-driven pulmonary artery and vein segmentation reveals demography-associated vasculature anatomical differences". They also released the code and datasets at https://github.com/Arturia-Pendragon-Iris/HiPaS_AV_Segmentation. If you find this dataset be helpful to you, please city their paper "Chu, Y., Luo, G., Zhou, L. et al. Deep learning-driven pulmonary artery and vein segmentation reveals… See the full description on the dataset page: https://huggingface.co/datasets/JasperEppink/HiPaS_dataset.0 likes89 downloads1y agoHugging Face30teticio /audio-diffusion-instrumental-hiphop-256256x256 mel spectrograms of 5 second samples of instrumental Hip Hop. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 256 y_res = 256 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K7 likes84 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.