datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenCodeReasoning-2
OpenCodeReasoning-2: A Large-scale Dataset for Reasoning in Code Generation and Critique
Dataset Description
OpenCodeReasoning-2 is the largest reasoning-based synthetic dataset to date for coding, comprising 1.4M samples in Python and 1.1M samples in C++ across 34,799 unique competitive programming questions.
OpenCodeReasoning-2 is designed for supervised fine-tuning (SFT) tasks of code completion and code critique.
Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/cublya/OpenCodeReasoning-2.cub-druid
Dataset Card for DRUID
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the DRUID dataset by Hagström et al. (2024). For this version, we have sampled 4,500 DRUID entries for which a "true target" (the factcheck verdict) and a "new target" (the stance of the context) could be found.
Dataset Structure
Thus far, we use two versions of the dataset: gpt2-xl and pythia-6.9b with corresponding validation (200 samples) and test splits… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-druid.cub-counterfact
Dataset Card for CounterFact
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the popular CounterFact dataset, originally proposed by Meng et al. (2022) and re-used in different variants by e.g. Ortu et al. (2024). For this version, the 899 CounterFact samples have been sampled based on the parametric memory of Pythia 6.9B, such that it contains samples for which the top model prediction without context is correct. We note that 546 samples in the… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-counterfact.cub-nq
Dataset Card for NQ
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the popular NQ dataset, originally proposed by Kwiatkowski et al. (2019). For this version, NQ samples have been obtained based on whether we can recover the gold passage from the original Wikipedia page and for which there is one short answer (less than 5 words in length). The context used in these samples is the correct gold context annotated by the original NQ annotators… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-nq.AXL-DATASET-1
AXL Training Data
Training datasets for the AXL multi-scale transformer model family by Cubic.
Dataset Structure
code/ # Raw Python code for language modeling
├── axl_hf_evol_codealpaca_v1.txt (239 MB) Real Python from HuggingFace
├── axl_hf_Evol_Instruct_Code_80k_v1.txt (113 MB) Evol-Instruct code data
├── axl_python_code_5gb.txt (58 MB) Python code corpus
├── axl_python_code_hf.txt (5.5 MB) Python code
├──… See the full description on the dataset page: https://huggingface.co/datasets/CubicLabs/AXL-DATASET-1.Agda-Cubical
Agda-Cubical
Structured dataset from agda/cubical — Cubical Agda library for HoTT and univalent mathematics.
Source
Repository: https://github.com/agda/cubical
Commit: d4a2af62de40a6ca9a0b51981e41f804d879a1b9
Files: 1192
License: mit
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body, empty… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Agda-Cubical.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad cultural a través de las provincias de Cuba, permitiendo a investigadores y profesionales desarrollar y evaluar modelos que comprendan y respeten los matices culturales cubanos.
Estadísticas del… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-cultural-appropriateness-prompts.patriae-cuban-cultural-appropriateness-prompts
Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana como parte de su participación en el reto #HackathonSomosNLP 2026: Preferencias.
Descripción General
Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuban-cultural-appropriateness-prompts.patriae-cuban-literature-dataset
Patriae Cuban Literature Dataset (40k)
Dataset de literatura cubana con 40k registros en formato CSV y Parquet, diseñado para el entrenamiento y evaluación de Modelos de Lenguaje (LLMs) y sistemas de Síntesis de Voz (TTS) orientados al dialecto y la cultura cubana.
Autores
Carlos Luis Barnés Infante (https://huggingface.co/blacknoize404)
Yisel Clavel Quintero (https://huggingface.co/clavel)
Curado por: Carlos Luis Barnés Infante
Descripción del… See the full description on the dataset page: https://huggingface.co/datasets/blacknoize404/patriae-cuban-literature-dataset.cubicaltt
cubicaltt
Declarations from cubicaltt, an experimental cubical type theory.
Source
Repository: https://github.com/mortberg/cubicaltt
Commit: 9baa6f2491cc61dbd4fd81d58323c04100381451
Files: 78
License: other
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body, empty if the declaration has none… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/cubicaltt.patriae-cuba-literature-dataset
Patriae Cuban Literature Dataset (31k)
Dataset de literatura cubana curado por el equipo de Patriae como parte de su participación en el evento SomosNLP 2026, con el objetivo de emplearse por el mismo en la realización de tareas de reproducción del dialecto cubano.
📌 Nota de procedencia: Este repositorio es un espejo (mirror) oficial para el evento. El desarrollo activo, las actualizaciones del dataset y la autoría principal pertenecen a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuba-literature-dataset.cuban-spanish-sample
que Cuban Spanish Conversational Sample (v0.3)
A synthetic sample that demonstrates the schema of the que conversational dataset for Cuban Spanish (es-CU). It accompanies the que white paper and shows prospective research partners what a que record looks like. This is synthetic demonstration data, not a collected corpus.
What this is
These records were constructed to the production schema to illustrate its shape. They stand in for data that has yet to be collected… See the full description on the dataset page: https://huggingface.co/datasets/que-app/cuban-spanish-sample.voice-agent-sft-v1
Voice-Agent SFT v1
Intended use: Downstream SFT fine-tuning for phone AI voice agents. NOT for pretraining.
This dataset targets models that must handle natural conversation + tool calling + function
calls (CRM lookups, web search, RAG DB queries) in a voice-agent context.
Built as a downstream specialization dataset for the
Cubix-AI diffusion-research project,
specifically for fine-tuning the Qwen3.5-4B-Base masked-diffusion LLM produced in Milestone 1.
Published under HF org… See the full description on the dataset page: https://huggingface.co/datasets/CubixAI/voice-agent-sft-v1.patriae-cuban-literature-dataset
Patriae Cuban Literature Dataset
Dataset de literatura cubana con +31k registros en formato CSV y Parquet, diseñado para el entrenamiento y evaluación de Modelos de Lenguaje (LLMs) y sistemas de Síntesis de Voz (TTS) orientados al dialecto y la cultura cubana.
Autores
Carlos Luis Barnés Infante (https://huggingface.co/blacknoize404)
Yisel Clavel Quintero (https://huggingface.co/clavel)
Curado por: Carlos Luis Barnés Infante
Descripción
Patriae… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-literature-dataset.
