datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AlgoTune
Website |
Paper |
Code
How good are language models at coming up with new algorithms? To try to answer this, we built a benchmark, AlgoTune, comprised of 154 widely used math, physics, and computer science functions. For each function, the goal is to write code that produces the same outputs as the original function, while being faster. In addition to the benchmark, we also provide an agent, AlgoTuner, which allows language models to easily optimize code.… See the full description on the dataset page: https://huggingface.co/datasets/oripress/AlgoTune.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.repro-organic-data-72Bbiomnibench-organized
BioMniBench DA — Reorganized
A clean, manifest-driven reorganization of the
BioMniBench DA (Data Analysis) task suite,
shaped for use with the
biomnibench-adapter evaluation
harness and the native skill-learning loop that ships with it.
This Hugging Face repository hosts the metadata, evaluation rubric and data manifest for
all 50 tasks. The raw input data files (which total ~77 GB and originate upstream from
GEO/TCGA/cBioPortal/etc.) are not redistributed here — see
Getting… See the full description on the dataset page: https://huggingface.co/datasets/starpacker52/biomnibench-organized.zoonomia-v1-v4_ccre_non_promoter-order
bolinas-dna/zoonomia-v1-v4_ccre_non_promoter-order
The bolinas-dna/zoonomia-v1-v4_ccre_non_promoter cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_non_promoter-order.swipe.futo.org
Dataset Card for swipe.futo.org
This dataset is presented in the paper FUTO Swipe: Layout-Agnostic Neural Swipe Decoding.
It contains multiple collection runs from the swipe.futo.org website. The QWERTY layout definition is provided here
Collection process
Users were able to volunteer to contribute to our dataset. After visiting the site on a mobile device, they were given words to swipe as part of a pre-defined sentence set.
Users were allowed to go back to retry… See the full description on the dataset page: https://huggingface.co/datasets/futo-org/swipe.futo.org.Orbit_Planner
Orbit-Planner Orbital Evasion Dataset
Orbit-Planner is a simulated multimodal trajectory dataset for vision-based
spacecraft navigation and obstacle avoidance. It contains synchronized
first-person RGB images, depth maps, spacecraft states, thruster commands, and
event labels collected in the Orbital Evasion task from Space Robotics Bench
and NVIDIA Isaac Sim.
The dataset is intended for learning latent world models, spacecraft dynamics,
visual representations… See the full description on the dataset page: https://huggingface.co/datasets/warriorLZJ/Orbit_Planner.zoonomia-v1-v4_cds-order
bolinas-dna/zoonomia-v1-v4_cds-order
The bolinas-dna/zoonomia-v1-v4_cds cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same per-window… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_cds-order.zoonomia-v1-v4_ccre_noexon_enhancer-order
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order
The bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order.compare_oracleORORAiThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ororai/ORORAi.LFM-Orbit-SatData
LFM Orbit SatData
Retagged Earth-observation training data produced by LFM Orbit for the Liquid AI x DPhi Space Hackathon.
The default viewer config is training_assets.jsonl, which contains single-image SFT rows with image, messages, and metadata. Temporal sequence rows live in the temporal_sft config so the Hugging Face Dataset Viewer does not try to cast sequence rows into the single-image schema.
Configs
Config
File
Purpose
default
training_assets.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Shoozes/LFM-Orbit-SatData.OR-Clarify
OR-Clarify
📄 Paper: Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
OR-Clarify is a benchmark for testing whether an agent asks the right questions before formulating an optimization model.
Most optimization benchmarks give an agent a complete problem statement. OR-Clarify instead starts with an incomplete business brief. The agent must identify missing requirements that could change the optimization formulation, ask for the relevant… See the full description on the dataset page: https://huggingface.co/datasets/AIOR-Research/OR-Clarify.zoonomia-v1-v4_ncrna_exon-order
bolinas-dna/zoonomia-v1-v4_ncrna_exon-order
The bolinas-dna/zoonomia-v1-v4_ncrna_exon cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ncrna_exon-order.bottleneck-oracle-graphsorc-bench
ORC-bench
Task 1: Topological Path Finding
Task 2: Topological Connectivity
Task 3: Linear Power Flow
Task 4: Contingency Analysis
Task 5: Power Grid ControlTask 6: Power Flow Optimization
Task 1: Topological Path Finding
Problem Formulation
This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.orbit-mt-eval
ORBIT-MT-Eval: Expert Annotations and Prompts for Machine Translation Meta-Evaluation
ORBIT-MT-Eval is an English-to-Russian reference dataset for machine translation meta-evaluation: comparing evaluator outputs against professional expert annotations. Each of the 600 source–translation segments has three independent expert annotations following the RATE protocol. In Beyond Prompts: A Systematic Study of LLM-Based Machine Translation Evaluators, two annotations serve as gold… See the full description on the dataset page: https://huggingface.co/datasets/foksly/orbit-mt-eval.zoonomia-v1-v4_utr3-order
bolinas-dna/zoonomia-v1-v4_utr3-order
The bolinas-dna/zoonomia-v1-v4_utr3 cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human anchors and same per-window… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_utr3-order.OrionEditBench
🧩 OrionEditBench
OrionEditBench is a large-scale dataset for cross-image editing, where each sample is structured as:
(reference image(s), source image) → synthesis image
It is designed to support multi-image conditioned generation and editing, allowing models to integrate visual information across inputs.
The dataset contains approximately 50K high-quality samples, covering key editing scenarios including attribute transfer, style alignment, and multi-image fusion.
We first… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuJiang1/OrionEditBench.ovos-wake-word-bench-community-ey-ordenador
OVOS wake_word bench — community-ey-ordenador
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/ovos-community-wakewords-dataset.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-community-ey-ordenador.AISE-Bench
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
🌐 Project Page •
💻 GitHub •
📖 KDD 2026 Paper
AISE-Bench is a real-world benchmark for information seeking on academic knowledge graphs. It is built from authentic AMiner user search queries and provides human-verified academic question-answering data with executable multi-step API trajectories, standardized tool inputs, API execution outputs, and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/AISE-Bench.beyond_accept_or_deny
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
orpr
Open Retail Process Reference (ORPR)
ORPR is an open, versioned reference architecture and notation for retail business processes. It
organizes retail domains, processes, cross-functional flows, responsibilities, data, and handoffs
between organizational areas.
This Hugging Face repository is the machine-readable distribution of ORPR, prepared for search,
analysis, retrieval-augmented generation, knowledge applications, and AI agents. The source language
of release v0.1.0 is… See the full description on the dataset page: https://huggingface.co/datasets/rafmyr/orpr.zoonomia-v1-v4_ccre_enhancer_centered-order
bolinas-dna/zoonomia-v1-v4_ccre_enhancer_centered-order
An enhancer-CENTERED training set for issue
#351, built by the
snakemake/zoonomia_projection_dataset pipeline
(workflow/rules/centered.smk) at commit
8127acfea5aa.
Provenance
Each training window is defined directly from an ENCODE cCRE V4 enhancer
(dELS + pELS): one 255 bp window centered on the cCRE midpoint
(make_enhancer_anchors, keep-all — clustered enhancers each keep their own
window), rather than a… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_enhancer_centered-order.schema-dot-org
Geolocated text from the Web Data Commons schema.org GeoCoordinates subset
12,427,530 geolocated text records, extracted from the class-specific
GeoCoordinates subset of the Web Data Commons schema.org data set series
(release 2024-12). Each record pairs one coordinate pair published on a web page
with the text published next to it on that same page.
Each source stream is deduplicated by host-local runs: a coordinate-and-name
pair is kept once per contiguous host run. A host… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/schema-dot-org.zoonomia-v1-v4_tss_region_and_utr5-order
bolinas-dna/zoonomia-v1-v4_tss_region_and_utr5-order
The bolinas-dna/zoonomia-v1-v4_tss_region_and_utr5 cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_tss_region_and_utr5-order.polymarket-crypto-5min-orderbooks
Polymarket Crypto 5-Minute L2 Order Books
Full level-2 order book snapshots at 1-second resolution for 15 complete Polymarket
crypto up/down markets — every resting bid and ask, not just top of book.
Seven assets: BTC, ETH, SOL, XRP, BNB, DOGE, HYPE. Each market is covered end to end
across its entire five-minute life, so it can be replayed rather than sampled.
from datasets import load_dataset
ds = load_dataset("polyorderbooks/polymarket-crypto-5min-orderbooks", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/polyorderbooks/polymarket-crypto-5min-orderbooks.HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details.ecml_arena_dataset
ARENA: A Cognitive Multi-Agent Framework for Modeling Conflict-Driven Multi-party Conversation
⚠️ Code under internal review. The generation / simulation code is
currently under internal code review — the GitHub repository is
coming soon. This repository already provides the dataset (a sample
subset) so it can be referenced from the paper.
📄 Paper. ARENA: A Cognitive Multi-Agent Framework for Modeling
Conflict-Driven Multi-party Conversation — ECML-PKDD.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Orange/ecml_arena_dataset.lm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private
Dataset Card for Evaluation run of chihoonlee10/T3Q-Mistral-Orca-Math-DPO
Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-Mistral-Orca-Math-DPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private.
