CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MiniMaxAI /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/OctoCodingBench.texttext-generationn<1K365 likes431 downloads8mo agoHugging Face02octaviomartinez /spanish-diccionary-lexicon Dataset de Léxico y Definiciones en Español Este conjunto de datos recopila palabras en español, sus categorías gramaticales, definiciones y metadatos regionales o de origen. Fue diseñado para facilitar tareas de procesamiento del lenguaje natural (NLP), modelos de lenguaje, lexicografía y análisis lingüístico del español. Estructura del Dataset El dataset cuenta con los siguientes campos por cada registro: Campo Tipo Descripción term String La palabra… See the full description on the dataset page: https://huggingface.co/datasets/octaviomartinez/spanish-diccionary-lexicon.text1M<n<10M2 likes134 downloads15d agoHugging Face03Koki-Kurita /DataSet_mix_duck_oct_cabtabularn<1K1 likes70 downloads13d agoHugging Face04bigcode /oasst-octopackThis is a filtered version of OASST to focus only on high-quality conversation trees as used in the OctoPack paper. from datasets import load_dataset d = load_dataset("bigcode/oasst-octopack")["train"] text1K<n<10K6 likes65 downloads3y agoHugging Face05open-llm-leaderboard /prithivMLmods__Llama-3.2-3B-Math-Oct-detailsgated Dataset Card for Evaluation run of prithivMLmods/Llama-3.2-3B-Math-Oct Dataset automatically created during the evaluation run of model prithivMLmods/Llama-3.2-3B-Math-Oct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Llama-3.2-3B-Math-Oct-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face06Koki-Kurita /DataSet_mix_duck_octtabularn<1K0 likes37 downloads14d agoHugging Face07AronDaron /OctoBench-2.2k OctoBench-2.2k — Coding Assistant Dataset Synthetic dataset for fine-tuning coding-focused LLMs. Generated with Dataset Generator — an open-source pipeline for building high-quality training data. Overview 2,248 multi-turn conversations across 8 categories: Category Examples Focus Model Gen Model Judge Refactor & Code Review 1 72 Performance refactors, behavior preservation qwen/qwen3-coder openai/gpt-oss-120b Refactor & Code Review 2 69 Performance… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/OctoBench-2.2k.text-generation1K<n<10K0 likes34 downloads5mo agoHugging Face08tqhuyen /harvard-oct-glaucoma-200-bilateral Harvard-GF 200^3 — bilateral-filtered volumes Bilateral-filtered counterpart of tqhuyen/harvard-oct-glaucoma-200. Volumes remain at the canonical 200^3 resolution, stored as uint8 .npy. Provenance Raw source: tqhuyen/harvard-oct-glaucoma-200, revision 939a38876b7b9313162842ef2d44b7edc2b57020 (itself derived from harvardairobotics/Harvard-GF, IEEE TMI 2024, Luo et al.). Filter: grayscale bilateral, applied slice-by-slice along source axis 0. Parameters:… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-200-bilateral.textn<1K0 likes30 downloads12d agoHugging Face09OctopusMode /empathy-finetune-datasettext10K<n<100K0 likes19 downloads1y agoHugging Face10octanove /moslagated Overview The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction, and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.tabularautomatic-speech-recognition100K<n<1M5 likes16 downloads2y agoHugging Face11bigcode /xp3x-octopacktext1K<n<10K5 likes15 downloads3y agoHugging Face12madhavsinghabcde /suggestedTitleModelData_20_Octtextn<1K0 likes14 downloads2y agoHugging Face13ngkuissi /queries-oct-2025-updatedtextn<1K0 likes14 downloads8mo agoHugging Face14Santarabantoosoo /Long_Covid_word_frequency_TFIDF_21_Jul_Octtabularn<1K0 likes12 downloads4y agoHugging Face15rico2512 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/rico2512/OctoCodingBench.texttext-generationn<1K0 likes12 downloads8mo agoHugging Face16Dodgeblurd /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding… See the full description on the dataset page: https://huggingface.co/datasets/Dodgeblurd/OctoCodingBench.texttext-generationn<1K0 likes12 downloads3mo agoHugging Face17Santarabantoosoo /whole_text_TF_21_Jul_Octtabularn<1K0 likes11 downloads4y agoHugging Face18yuan909815 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/yuan909815/OctoCodingBench.texttext-generationn<1K0 likes11 downloads8mo agoHugging Face19Okok109 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/Okok109/OctoCodingBench.texttext-generationn<1K0 likes11 downloads8mo agoHugging Face20itsPrerna202 /OctoCodingBench OctoCodingBench: Instruction-Following Benchmark for Coding Agents English | 中文 🌟 Overview OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding. Why OctoCodingBench? Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task? In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/itsPrerna202/OctoCodingBench.texttext-generationn<1K0 likes11 downloads5mo agoHugging Face21tqhuyen /harvard-oct-glaucoma-96 Harvard-GF 96^3 resized dataset Resized (antialiased) version of Harvard-GF OCT volumes at 96^3 uint8. Source: harvardairobotics/Harvard-GF (raw 200^3 B-scans), IEEE TMI 2024 (Luo et al.) store_shape: [1, 96, 96, 96] | source_shape: [200, 200, 200] | antialias: True Splits (volumes, pos=glaucoma / neg): Training: 2100 (pos 1083 / neg 1017) Validation: 300 (pos 176 / neg 124) Test: 900 (pos 489 / neg 411) Layout: {Training,Validation,Test}_{volumes,labels}.npy (volumes (N,1,96… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-96.textn<1K0 likes11 downloads13d agoHugging Face22Octatom /langtech-lab2text1M<n<10M0 likes11 downloads1d agoHugging Face23octavio-gutierrez /trainjtext10K<n<100K0 likes10 downloads3y agoHugging Face24GritLM /oasst_octopack_entext1K<n<10K0 likes9 downloads3y agoHugging Face25theprint /databird-oct25-collection Databird October Collection 2025 This is the majority of data from the databird collection, as it looked mid-October 2025, put into a single data set. texttext-generation10K<n<100K0 likes9 downloads1y agoHugging Face26ngkuissi /qrels_oct_2024 qrels_oct_2024 QRELS dataset generated from 2024 experimentation assessment results. Dataset Structure The dataset follows the QRELS (Query Relevance) JSON format: { "qrels_nuggets": { "query_id": { "doc_id": score, ... }, ... } } query_id: Unique identifier for the query. doc_id: Unique identifier for the document chunk. score: Relevance score (derived from nugget-level judgment). Usage import json from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/ngkuissi/qrels_oct_2024.n<1K0 likes9 downloads8mo agoHugging Face27Octavian10 /c-tutor-master-datasettextn<1K0 likes9 downloads4mo agoHugging Face28tqhuyen /harvard-oct-glaucoma-128 Harvard-GF 128^3 resized dataset Resized (antialiased) version of Harvard-GF OCT volumes at 128^3 uint8. Source: harvardairobotics/Harvard-GF (raw 200^3 B-scans), IEEE TMI 2024 (Luo et al.) store_shape: [1, 128, 128, 128] | source_shape: [200, 200, 200] | antialias: True Splits (volumes, pos=glaucoma / neg): Training: 2100 (pos 1083 / neg 1017) Validation: 300 (pos 176 / neg 124) Test: 900 (pos 489 / neg 411) Layout: {Training,Validation,Test}_{volumes,labels}.npy (volumes (N… See the full description on the dataset page: https://huggingface.co/datasets/tqhuyen/harvard-oct-glaucoma-128.textn<1K0 likes9 downloads13d agoHugging Face29ngkuissi /queries-oct-2025-coheretextn<1K0 likes8 downloads8mo agoHugging Face30openaccess-ai-collective /oasst-octopack-entext1K<n<10K1 likes7 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.