CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MiniMaxAI /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/OctoCodingBench.texttext-generationn<1K368 likes426 downloads8mo agoHugging Face02octaviomartinez /spanish-diccionary-lexicon Dataset de Léxico y Definiciones en Español Este conjunto de datos recopila palabras en español, sus categorías gramaticales, definiciones y metadatos regionales o de origen. Fue diseñado para facilitar tareas de procesamiento del lenguaje natural (NLP), modelos de lenguaje, lexicografía y análisis lingüístico del español. Estructura del Dataset El dataset cuenta con los siguientes campos por cada registro: Campo Tipo Descripción term String La palabra… See the full description on the dataset page: https://huggingface.co/datasets/octaviomartinez/spanish-diccionary-lexicon.text1M<n<10M2 likes134 downloads16d agoHugging Face03Koki-Kurita /DataSet_mix_duck_oct_cabtabularn<1K1 likes71 downloads14d agoHugging Face04bigcode /oasst-octopackThis is a filtered version of OASST to focus only on high-quality conversation trees as used in the OctoPack paper. from datasets import load_dataset d = load_dataset("bigcode/oasst-octopack")["train"] text1K<n<10K6 likes61 downloads3y agoHugging Face05open-llm-leaderboard /prithivMLmods__Llama-3.2-3B-Math-Oct-detailsgated Dataset Card for Evaluation run of prithivMLmods/Llama-3.2-3B-Math-Oct Dataset automatically created during the evaluation run of model prithivMLmods/Llama-3.2-3B-Math-Oct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Llama-3.2-3B-Math-Oct-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face06Koki-Kurita /DataSet_mix_duck_octtabularn<1K0 likes39 downloads15d agoHugging Face07tqhuyen /harvard-oct-glaucoma-200-bilateral Harvard-GF 200^3 — bilateral-filtered volumes Bilateral-filtered counterpart of tqhuyen/harvard-oct-glaucoma-200. Volumes remain at the canonical 200^3 resolution, stored as uint8 .npy. Provenance Raw source: tqhuyen/harvard-oct-glaucoma-200, revision 939a38876b7b9313162842ef2d44b7edc2b57020 (itself derived from harvardairobotics/Harvard-GF, IEEE TMI 2024, Luo et al.). Filter: grayscale bilateral, applied slice-by-slice along source axis 0. Parameters:… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-200-bilateral.textn<1K0 likes30 downloads13d agoHugging Face08Octatom /langtech-lab2text1M<n<10M0 likes21 downloads3d agoHugging Face09octanove /moslagated Overview The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction, and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.tabularautomatic-speech-recognition100K<n<1M5 likes19 downloads2y agoHugging Face10OctopusMode /empathy-finetune-datasettext10K<n<100K0 likes19 downloads1y agoHugging Face11bigcode /xp3x-octopacktext1K<n<10K5 likes14 downloads3y agoHugging Face12madhavsinghabcde /suggestedTitleModelData_20_Octtextn<1K0 likes14 downloads2y agoHugging Face13ngkuissi /queries-oct-2025-updatedtextn<1K0 likes14 downloads8mo agoHugging Face14Dodgeblurd /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding… See the full description on the dataset page: https://huggingface.co/datasets/Dodgeblurd/OctoCodingBench.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face15theprint /databird-oct25-collection Databird October Collection 2025 This is the majority of data from the databird collection, as it looked mid-October 2025, put into a single data set. texttext-generation10K<n<100K0 likes11 downloads1y agoHugging Face16tqhuyen /harvard-oct-glaucoma-96 Harvard-GF 96^3 resized dataset Resized (antialiased) version of Harvard-GF OCT volumes at 96^3 uint8. Source: harvardairobotics/Harvard-GF (raw 200^3 B-scans), IEEE TMI 2024 (Luo et al.) store_shape: [1, 96, 96, 96] | source_shape: [200, 200, 200] | antialias: True Splits (volumes, pos=glaucoma / neg): Training: 2100 (pos 1083 / neg 1017) Validation: 300 (pos 176 / neg 124) Test: 900 (pos 489 / neg 411) Layout: {Training,Validation,Test}_{volumes,labels}.npy (volumes (N,1,96… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-96.textn<1K0 likes11 downloads14d agoHugging Face17octavio-gutierrez /trainjtext10K<n<100K0 likes10 downloads3y agoHugging Face18rico2512 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/rico2512/OctoCodingBench.texttext-generationn<1K0 likes10 downloads8mo agoHugging Face19yuan909815 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/yuan909815/OctoCodingBench.texttext-generationn<1K0 likes9 downloads8mo agoHugging Face20Okok109 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/Okok109/OctoCodingBench.texttext-generationn<1K0 likes9 downloads8mo agoHugging Face21itsPrerna202 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/itsPrerna202/OctoCodingBench.texttext-generationn<1K0 likes9 downloads5mo agoHugging Face22Octavian10 /c-tutor-master-datasettextn<1K0 likes9 downloads4mo agoHugging Face23tqhuyen /harvard-oct-glaucoma-128 Harvard-GF 128^3 resized dataset Resized (antialiased) version of Harvard-GF OCT volumes at 128^3 uint8. Source: harvardairobotics/Harvard-GF (raw 200^3 B-scans), IEEE TMI 2024 (Luo et al.) store_shape: [1, 128, 128, 128] | source_shape: [200, 200, 200] | antialias: True Splits (volumes, pos=glaucoma / neg): Training: 2100 (pos 1083 / neg 1017) Validation: 300 (pos 176 / neg 124) Test: 900 (pos 489 / neg 411) Layout: {Training,Validation,Test}_{volumes,labels}.npy (volumes (N… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-128.textn<1K0 likes9 downloads14d agoHugging Face24GritLM /oasst_octopack_entext1K<n<10K0 likes8 downloads3y agoHugging Face25ngkuissi /queries-oct-2025-coheretextn<1K0 likes8 downloads8mo agoHugging Face26openaccess-ai-collective /oasst-octopack-entext1K<n<10K1 likes7 downloads3y agoHugging Face27persona-shattering-lasr /oct-runs-low-conscientiousness-full-v1textn<1K0 likes5 downloads6mo agoHugging Face28sdananya /eigenbench-oct-dpo-vs-introspection EigenBench OCT: DPO vs Introspection — Scenario-Level Wins This dataset contains the scenarios on which a DPO-trained persona model (DPO-final) is judged to be more aligned with a target persona constitution than an Introspection-trained persona model (Introspection-final), aggregated across multiple judges and orderings. The ten persona constitutions are taken from the OCT (Open Constitution Taxonomy) set shipped with EigenBench (data/constitutions/oct_*.json): goodness, humor… See the full description on the dataset page: https://huggingface.co/datasets/sdananya/eigenbench-oct-dpo-vs-introspection.tabulartext-classificationn<1K0 likes5 downloads5mo agoHugging Face29Octo26 /final_filter_datasettext1K<n<10K0 likes3 downloads3y agoHugging Face30ngkuissi /queries-oct-2024-updatedtextn<1K0 likes3 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.