CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cublya /OpenCodeReasoning-2 OpenCodeReasoning-2: A Large-scale Dataset for Reasoning in Code Generation and Critique Dataset Description OpenCodeReasoning-2 is the largest reasoning-based synthetic dataset to date for coding, comprising 1.4M samples in Python and 1.1M samples in C++ across 34,799 unique competitive programming questions. OpenCodeReasoning-2 is designed for supervised fine-tuning (SFT) tasks of code completion and code critique. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/cublya/OpenCodeReasoning-2.texttext-generation1M<n<10M0 likes589 downloads8mo agoHugging Face02copenlu /cub-druid Dataset Card for DRUID Of the cmt-benchmark project. Dataset Details This dataset is a version of the DRUID dataset by Hagström et al. (2024). For this version, we have sampled 4,500 DRUID entries for which a "true target" (the factcheck verdict) and a "new target" (the stance of the context) could be found. Dataset Structure Thus far, we use two versions of the dataset: gpt2-xl and pythia-6.9b with corresponding validation (200 samples) and test splits… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-druid.tabularquestion-answering10K<n<100K0 likes353 downloads1y agoHugging Face03copenlu /cub-counterfact Dataset Card for CounterFact Of the cmt-benchmark project. Dataset Details This dataset is a version of the popular CounterFact dataset, originally proposed by Meng et al. (2022) and re-used in different variants by e.g. Ortu et al. (2024). For this version, the 899 CounterFact samples have been sampled based on the parametric memory of Pythia 6.9B, such that it contains samples for which the top model prediction without context is correct. We note that 546 samples in the… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-counterfact.tabularquestion-answering10K<n<100K0 likes330 downloads1y agoHugging Face04copenlu /cub-nq Dataset Card for NQ Of the cmt-benchmark project. Dataset Details This dataset is a version of the popular NQ dataset, originally proposed by Kwiatkowski et al. (2019). For this version, NQ samples have been obtained based on whether we can recover the gold passage from the original Wikipedia page and for which there is one short answer (less than 5 words in length). The context used in these samples is the correct gold context annotated by the original NQ annotators… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-nq.tabularquestion-answering10K<n<100K0 likes309 downloads1y agoHugging Face05CubicLabs /AXL-DATASET-1 AXL Training Data Training datasets for the AXL multi-scale transformer model family by Cubic. Dataset Structure code/ # Raw Python code for language modeling ├── axl_hf_evol_codealpaca_v1.txt (239 MB) Real Python from HuggingFace ├── axl_hf_Evol_Instruct_Code_80k_v1.txt (113 MB) Evol-Instruct code data ├── axl_python_code_5gb.txt (58 MB) Python code corpus ├── axl_python_code_hf.txt (5.5 MB) Python code ├──… See the full description on the dataset page: https://huggingface.co/datasets/CubicLabs/AXL-DATASET-1.texttext-generation10M<n<100M0 likes59 downloads3mo agoHugging Face06phanerozoic /Agda-Cubical Agda-Cubical Structured dataset from agda/cubical — Cubical Agda library for HoTT and univalent mathematics. Source Repository: https://github.com/agda/cubical Commit: d4a2af62de40a6ca9a0b51981e41f804d879a1b9 Files: 1192 License: mit Schema Column Type Description statement string Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof proof string Verbatim proof/body, empty… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Agda-Cubical.texttext-generation1K<n<10K0 likes19 downloads3mo agoHugging Face07Patriae /patriae-cuban-cultural-appropriateness-prompts Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana Descripción General Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad cultural a través de las provincias de Cuba, permitiendo a investigadores y profesionales desarrollar y evaluar modelos que comprendan y respeten los matices culturales cubanos. Estadísticas del… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-cultural-appropriateness-prompts.texttext-classification1K<n<10K0 likes16 downloads4mo agoHugging Face08somosnlp-hackathon-2026 /patriae-cuban-cultural-appropriateness-prompts Patriae - Dataset de Prompts para evaluar Apropiación Cultural Cubana Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana como parte de su participación en el reto #HackathonSomosNLP 2026: Preferencias. Descripción General Este dataset contiene 1,000 prompts diseñados para evaluar la apropiación cultural en el contexto de la cultura regional cubana. El dataset captura la rica diversidad… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuban-cultural-appropriateness-prompts.texttext-classification1K<n<10K0 likes16 downloads4mo agoHugging Face09blacknoize404 /patriae-cuban-literature-datasetgated Patriae Cuban Literature Dataset (40k) Dataset de literatura cubana con 40k registros en formato CSV y Parquet, diseñado para el entrenamiento y evaluación de Modelos de Lenguaje (LLMs) y sistemas de Síntesis de Voz (TTS) orientados al dialecto y la cultura cubana. Autores Carlos Luis Barnés Infante (https://huggingface.co/blacknoize404) Yisel Clavel Quintero (https://huggingface.co/clavel) Curado por: Carlos Luis Barnés Infante Descripción del… See the full description on the dataset page: https://huggingface.co/datasets/blacknoize404/patriae-cuban-literature-dataset.texttext-generation10K<n<100K1 likes15 downloads4mo agoHugging Face10phanerozoic /cubicaltt cubicaltt Declarations from cubicaltt, an experimental cubical type theory. Source Repository: https://github.com/mortberg/cubicaltt Commit: 9baa6f2491cc61dbd4fd81d58323c04100381451 Files: 78 License: other Schema Column Type Description statement string Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof proof string Verbatim proof/body, empty if the declaration has none… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/cubicaltt.texttext-generation1K<n<10K0 likes13 downloads4mo agoHugging Face11somosnlp-hackathon-2026 /patriae-cuba-literature-datasetgated Patriae Cuban Literature Dataset (31k) Dataset de literatura cubana curado por el equipo de Patriae como parte de su participación en el evento SomosNLP 2026, con el objetivo de emplearse por el mismo en la realización de tareas de reproducción del dialecto cubano. 📌 Nota de procedencia: Este repositorio es un espejo (mirror) oficial para el evento. El desarrollo activo, las actualizaciones del dataset y la autoría principal pertenecen a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/patriae-cuba-literature-dataset.tabulartext-generation10K<n<100K1 likes8 downloads3mo agoHugging Face12que-app /cuban-spanish-sample que Cuban Spanish Conversational Sample (v0.3) A synthetic sample that demonstrates the schema of the que conversational dataset for Cuban Spanish (es-CU). It accompanies the que white paper and shows prospective research partners what a que record looks like. This is synthetic demonstration data, not a collected corpus. What this is These records were constructed to the production schema to illustrate its shape. They stand in for data that has yet to be collected… See the full description on the dataset page: https://huggingface.co/datasets/que-app/cuban-spanish-sample.texttext-generationn<1K0 likes7 downloads3mo agoHugging Face13CubixAI /voice-agent-sft-v1gated Voice-Agent SFT v1 Intended use: Downstream SFT fine-tuning for phone AI voice agents. NOT for pretraining. This dataset targets models that must handle natural conversation + tool calling + function calls (CRM lookups, web search, RAG DB queries) in a voice-agent context. Built as a downstream specialization dataset for the Cubix-AI diffusion-research project, specifically for fine-tuning the Qwen3.5-4B-Base masked-diffusion LLM produced in Milestone 1. Published under HF org… See the full description on the dataset page: https://huggingface.co/datasets/CubixAI/voice-agent-sft-v1.texttext-generation1M<n<10M0 likes6 downloads5mo agoHugging Face14Patriae /patriae-cuban-literature-datasetgated Patriae Cuban Literature Dataset Dataset de literatura cubana con +31k registros en formato CSV y Parquet, diseñado para el entrenamiento y evaluación de Modelos de Lenguaje (LLMs) y sistemas de Síntesis de Voz (TTS) orientados al dialecto y la cultura cubana. Autores Carlos Luis Barnés Infante (https://huggingface.co/blacknoize404) Yisel Clavel Quintero (https://huggingface.co/clavel) Curado por: Carlos Luis Barnés Infante Descripción Patriae… See the full description on the dataset page: https://huggingface.co/datasets/Patriae/patriae-cuban-literature-dataset.tabulartext-generation10K<n<100K1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.