datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OctoCodingBench
OctoCodingBench: Instruction-Following Benchmark for Coding Agents
English | 中文
🌟 Overview
OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Why OctoCodingBench?
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/OctoCodingBench.spanish-diccionary-lexicon
Dataset de Léxico y Definiciones en Español
Este conjunto de datos recopila palabras en español, sus categorías gramaticales, definiciones y metadatos regionales o de origen. Fue diseñado para facilitar tareas de procesamiento del lenguaje natural (NLP), modelos de lenguaje, lexicografía y análisis lingüístico del español.
Estructura del Dataset
El dataset cuenta con los siguientes campos por cada registro:
Campo
Tipo
Descripción
term
String
La palabra… See the full description on the dataset page: https://huggingface.co/datasets/octaviomartinez/spanish-diccionary-lexicon.DataSet_mix_duck_oct_caboasst-octopackThis is a filtered version of OASST to focus only on high-quality conversation trees as used in the OctoPack paper.
from datasets import load_dataset
d = load_dataset("bigcode/oasst-octopack")["train"]
prithivMLmods__Llama-3.2-3B-Math-Oct-details
Dataset Card for Evaluation run of prithivMLmods/Llama-3.2-3B-Math-Oct
Dataset automatically created during the evaluation run of model prithivMLmods/Llama-3.2-3B-Math-Oct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Llama-3.2-3B-Math-Oct-details.DataSet_mix_duck_octharvard-oct-glaucoma-200-bilateral
Harvard-GF 200^3 — bilateral-filtered volumes
Bilateral-filtered counterpart of
tqhuyen/harvard-oct-glaucoma-200.
Volumes remain at the canonical 200^3 resolution, stored as uint8 .npy.
Provenance
Raw source: tqhuyen/harvard-oct-glaucoma-200, revision
939a38876b7b9313162842ef2d44b7edc2b57020
(itself derived from harvardairobotics/Harvard-GF, IEEE TMI 2024, Luo et al.).
Filter: grayscale bilateral, applied slice-by-slice along source axis 0.
Parameters:… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-200-bilateral.langtech-lab2mosla
Overview
The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one
of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction,
and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human
annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.empathy-finetune-datasetxp3x-octopacksuggestedTitleModelData_20_Octqueries-oct-2025-updatedOctoCodingBench
OctoCodingBench: Instruction-Following Benchmark for Coding Agents
English | 中文
🌟 Overview
OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Why OctoCodingBench?
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding… See the full description on the dataset page: https://huggingface.co/datasets/Dodgeblurd/OctoCodingBench.databird-oct25-collection
Databird October Collection 2025
This is the majority of data from the databird collection, as it looked mid-October 2025, put into a single data set.
harvard-oct-glaucoma-96
Harvard-GF 96^3 resized dataset
Resized (antialiased) version of Harvard-GF OCT volumes at 96^3 uint8.
Source: harvardairobotics/Harvard-GF (raw 200^3 B-scans), IEEE TMI 2024 (Luo et al.)
store_shape: [1, 96, 96, 96] | source_shape: [200, 200, 200] | antialias: True
Splits (volumes, pos=glaucoma / neg):
Training: 2100 (pos 1083 / neg 1017)
Validation: 300 (pos 176 / neg 124)
Test: 900 (pos 489 / neg 411)
Layout: {Training,Validation,Test}_{volumes,labels}.npy
(volumes (N,1,96… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-96.trainjOctoCodingBench
OctoCodingBench: Instruction-Following Benchmark for Coding Agents
English | 中文
🌟 Overview
OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Why OctoCodingBench?
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/rico2512/OctoCodingBench.OctoCodingBench
OctoCodingBench: Instruction-Following Benchmark for Coding Agents
English | 中文
🌟 Overview
OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Why OctoCodingBench?
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/yuan909815/OctoCodingBench.OctoCodingBench
OctoCodingBench: Instruction-Following Benchmark for Coding Agents
English | 中文
🌟 Overview
OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Why OctoCodingBench?
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/Okok109/OctoCodingBench.OctoCodingBench
OctoCodingBench: Instruction-Following Benchmark for Coding Agents
English | 中文
🌟 Overview
OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Why OctoCodingBench?
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/itsPrerna202/OctoCodingBench.c-tutor-master-datasetharvard-oct-glaucoma-128
Harvard-GF 128^3 resized dataset
Resized (antialiased) version of Harvard-GF OCT volumes at 128^3 uint8.
Source: harvardairobotics/Harvard-GF (raw 200^3 B-scans), IEEE TMI 2024 (Luo et al.)
store_shape: [1, 128, 128, 128] | source_shape: [200, 200, 200] | antialias: True
Splits (volumes, pos=glaucoma / neg):
Training: 2100 (pos 1083 / neg 1017)
Validation: 300 (pos 176 / neg 124)
Test: 900 (pos 489 / neg 411)
Layout: {Training,Validation,Test}_{volumes,labels}.npy
(volumes (N… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-128.oasst_octopack_enqueries-oct-2025-cohereoasst-octopack-enoct-runs-low-conscientiousness-full-v1eigenbench-oct-dpo-vs-introspection
EigenBench OCT: DPO vs Introspection — Scenario-Level Wins
This dataset contains the scenarios on which a DPO-trained persona model (DPO-final)
is judged to be more aligned with a target persona constitution than an Introspection-trained
persona model (Introspection-final), aggregated across multiple judges and orderings.
The ten persona constitutions are taken from the OCT (Open Constitution Taxonomy)
set shipped with EigenBench (data/constitutions/oct_*.json): goodness, humor… See the full description on the dataset page: https://huggingface.co/datasets/sdananya/eigenbench-oct-dpo-vs-introspection.final_filter_datasetqueries-oct-2024-updated
