CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01somosnlp /RecetasDeLaAbuela Motivación inicial Este corpus ha sido creado durante el Hackathon SomosNLP Marzo 2024: #Somos600M (https://somosnlp.org/hackathon). Responde a una de las propuestas somosnlp sobre 'Recetas típicas por país/zona geográfica'. Nombre del Proyecto Este corpus o dataset se llama 'RecetasDeLaAbuel@' y es un homenaje a todas nuestr@s abuel@s que nos han enseñado a cocinar. Se trata de la mayor y más completa colección de recetas open-source en español de países… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/RecetasDeLaAbuela.tabularquestion-answering10K<n<100K7 likes701 downloads2y agoHugging Face02somosnlp /SMC Dataset Card for Spanish Medical Corpus (SMC) This dataset groups and organizes several datasets present in hugginface (e.g.: PlanTL-GOB-ES/cantemist-ner, PlanTL-GOB-ES/pharmaconer) and other public resources created by researchers with different formats (e.g.; MedLexSp ) to allow it to be a source of knowledge of large language models in Spanish for the medical domain. Dataset Card in Spanish Dataset Details Dataset Description Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/SMC.textquestion-answering1M<n<10M10 likes207 downloads2y agoHugging Face03somosnlp /instruct-legal-refugiados-es Dataset Card for AsistenciaRefugiados README in Spanish Spain is the third country with the highest number of asylum applications, receiving each year approximately more than 100,000 applications, and the third with the lowest number of approvals within the EU. The main objective of this project is to facilitate the tasks of NGOs in this field and other institutions and help them to obtain answers to questions (QA) related to refugee legislation in Spanish. With its… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/instruct-legal-refugiados-es.textquestion-answering10K<n<100K3 likes192 downloads2y agoHugging Face04somasekhar-dev /nexttoken-model-1-dataset-sft NextToken Model 1 SFT dataset (v4) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from 846 scheme/product source documents across 57 schemes/products, chunked into 1,445 passages. v4 vs v3: v3 merged in a second batch (6,664 rows) without… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-1-dataset-sft.question-answering0 likes76 downloads6d agoHugging Face05AngelGabrielTroncoso /dataset-aeroespacial-cultural-somosnlp LATAM Aerospace History QA Descripción General LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica. El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos. La colección está especializada en: historia aeroespacial, programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.textquestion-answering100K<n<1M0 likes69 downloads4mo agoHugging Face06NeilFarmer /soma-skincare-qatextquestion-answering10K<n<100K0 likes59 downloads2mo agoHugging Face07somosnlp /recetasdelaabuela_genstruct_it Descripción Dataset creado para la hackathon #Somos600M con el objetivo de entrenar un modelo que pueda recomendar recetas de paises hispanohablantes. Este conjunto de datos consiste en pregunta-respuesta y fue elaborado a partir de un contexto usando Genstruct-7B y distilabel. Elaborado a partir del dataset en crudo somosnlp/RecetasDeLaAbuela elaborado por el equipo recetasdelaabuela mediante web scraping. Origen del Dataset El dataset se obtuvo mediante web… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/recetasdelaabuela_genstruct_it.textquestion-answering10K<n<100K3 likes56 downloads2y agoHugging Face08somosnlp /constitucion-politica-del-peru-1993-qa Dataset Summary Compuesto por unos 2075 registros que contienen los campos: pregunta: pregunta que sirve como una instrucción o consulta sobre algún aspecto de la Constitución Política del Perú de 1993. respuesta: La respuesta proporcionada para cada pregunta es un contexto relevante que ayuda a resolver la consulta. Este contexto es un extracto de la Constitución. fuente: Para cada respuesta, se indica el capítulo y/o artículo de la Constitución Política del Perú de 1993… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/constitucion-politica-del-peru-1993-qa.textsummarization1K<n<10K2 likes54 downloads2y agoHugging Face09somasekhar-dev /nexttoken-pmkisan-domain-sft-data NextToken pmkisan domain SFT data (v1) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from ~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA, banking products, insurance, savings instruments, etc.). Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.tabularquestion-answering1K<n<10K0 likes53 downloads7d agoHugging Face10somosnlp /LingComp_QA Dataset Card for LingComp_QA, un corpus educativo de lingüística computacional en español Dataset Details Dataset Description Curated by: Jorge Zamora Rey, Isabel Moyano Moreno, Mario Crespo Miguel Funded by: SomosNLP, HuggingFace, Argilla, Instituto de Lingüística Aplicada de la Universidad de Cádiz Language(s) (NLP): es-ES License: apache-2.0 Dataset Sources Repository: https://github.com/reddrex/lingcomp_QA/tree/main Paper:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/LingComp_QA.textquestion-answering1K<n<10K1 likes51 downloads8mo agoHugging Face11somosnlp /LLM_SQL_BaseDatosEspanol Usos Usos directos El objetivo principal de este dataset es proporcionar ejemplos simples para el fine-tuning de modelos de procesamiento de lenguaje natural (NLP) en el contexto de consultas SQL. Usos fuera de mira Podria usarse para el entrenamiento de una IA que sirva como creadora de base de datos artificiales Estructura del conjunto de datos Question: Es la pegunta que el usuario le dara al chatbot Answer: La respuesta el que chatbot le… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/LLM_SQL_BaseDatosEspanol.textquestion-answeringn<1K10 likes51 downloads2y agoHugging Face12somosnlp /wikihow_es Citation @misc{quispe2024wikihowes, author = {Quispe Castillo, David Alonso}, title = {WikiHowES}, month = March, year = 2024, url = {https://huggingface.co/datasets/somosnlp/wikihow_es} } textquestion-answering100K<n<1M0 likes48 downloads2y agoHugging Face13sommify /sommbench SommBench SommBench is a multilingual benchmark for assessing sommelier expertise in large language models. It comprises 3,024 expert-curated examples across eight languages (da, de, en, es, fi, it, sk, sv), designed by professional sommeliers to evaluate sensory grounding, factual wine knowledge, and practical pairing skills through three tasks: WTQA, WFC, and FWP. Configs wtqa — Wine Theory Question Answering (1,024 examples) Multiple-choice questions about… See the full description on the dataset page: https://huggingface.co/datasets/sommify/sommbench.tabularquestion-answering1K<n<10K0 likes46 downloads6mo agoHugging Face14somosnlp /justicio_evaluacion_ideonidad_preguntas_legalesEste dataset nos permitirá evaluar la idoneidad de las preguntas generadas para su uso dentro de la plataforma Justicio, un archivero que permite consultar desde una interfaz chat las distintas legislaciones, tanto a nivel nacional derivadas del Boletín Oficial del Estado, así como de las derivadas de las distintas Comunidades Autónomas. Internamente, Justicio utiliza un esquema de tipo RAG (Retrieval-Augmented Generation) en el que se localizan aquellos fragmentos almacenados más similares a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/justicio_evaluacion_ideonidad_preguntas_legales.textquestion-answeringn<1K1 likes45 downloads3y agoHugging Face15Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes43 downloads15d agoHugging Face16somosnlp /recetasdelaabuela_it Nombre del dataset Este dataset se llama 'RecetasDeLaAbuel@' y es un homenaje a todas nuestr@s abuel@s que nos han enseñado a cocinar. Se trata de la mayor y más completa colección de recetas open-source en español de países hispanoamericanos. Descripción Dataset creado on el objetivo de entrenar un modelo que pueda recomendar recetas de paises hispanohablantes. Nuestra IA responderá a cuestiones de los sigientes tipos: 'Qué puedo cocinar con 3 ingredientes?', 'Dime una… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/recetasdelaabuela_it.textquestion-answering10K<n<100K4 likes42 downloads2y agoHugging Face17Somtharu181coder /Annex_7_Educational_Indicators_2024-2025 Grounded Educational Indicators — Nepali SFT Dataset File: grounded_education_indicators_nepali_sft.jsonl Records: 10,091 Language: Nepali (ne / ISO 639-3 npi), Devanagari script License: Apache-2.0 (permissive) Format: JSON Lines, ShareGPT-style conversational (human / gpt turns) Task type: Grounded question answering over structured (tabular) educational statistics Size on disk: ~15.5 MB 1. Overview This dataset is a supervised fine-tuning (SFT) corpus of 10… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Annex_7_Educational_Indicators_2024-2025.textquestion-answering10K<n<100K0 likes39 downloads25d agoHugging Face18farihashifa /SOMAJGYAAN SomajGyaan (সমাজজ্ঞান) - Bangla MCQ Dataset 📊 Dataset Description SomajGyaan (সমাজজ্ঞান) is a comprehensive Bangla multiple-choice question dataset featuring 4,234 questions across 7 academic categories with ~12,000 unique answer options. Dataset Summary Total Questions: 4,234 Unique Answer Options: ~12,000 Answer Diversity: 70.8% Language: Bangla (Bengali) Categories: 7 (History, Economics, Geography, Politics, Social Studies, Law… See the full description on the dataset page: https://huggingface.co/datasets/farihashifa/SOMAJGYAAN.tabularquestion-answering1K<n<10K0 likes36 downloads2mo agoHugging Face19Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes36 downloads1mo agoHugging Face20anonymous-somebody /CHIMERA CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation. Total: 9,225 problems Subjects: 8 Topics: 1,179 Overview Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.texttext-generation1K<n<10K0 likes35 downloads5mo agoHugging Face21Somtharu181coder /educational_domain_dataset Nepali Grounded Education QA (OpenHermes-format) A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about student enrollment statistics from Nepal's Ministry of Education. Every answer is anchored to a real numeric value pulled from government open data — nothing in the answers is model-hallucinated. Dataset Summary Rows 611 Language Nepali (Devanagari script) Format ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.textquestion-answeringn<1K0 likes30 downloads1mo agoHugging Face22michsethowusu /Code-170k-somali Dataset Description Code-170k-somali is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Somali, making coding education accessible to Somali speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Somali language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-somali.texttext-generation100K<n<1M1 likes29 downloads11mo agoHugging Face23somosnlp /SMC-instruct Dataset Card for Spanish Medical Corpus (SMC) This dataset groups and organizes several datasets present in hugginface (e.g.: PlanTL-GOB-ES/cantemist-ner, PlanTL-GOB-ES/pharmaconer) and other public resources created by researchers with different formats (e.g.; MedLexSp ) to allow it to be a source of knowledge of large language models in Spanish for the medical domain. Dataset Card in Spanish Dataset Details Dataset Description Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/SMC-instruct.question-answering0 likes27 downloads2y agoHugging Face24somosnlp-hackathon-2025 /ec-prompts-refranestextquestion-answeringn<1K0 likes26 downloads1y agoHugging Face25somosnlp-hackathon-2025 /cenia-team-sabiduriapopulartextquestion-answeringn<1K0 likes25 downloads1y agoHugging Face26somosnlp-hackathon-2025 /ibero-tales-es Conjunto de datos de historias sintéticas de mitos y leyendas iberoamericanos. ⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y refinar el proceso de generación para mejorar la diversidad y calidad narrativa. 📚 Descripción Dataset de historias sintéticas generadas a partir de mitos y leyendas de Iberoamérica, curado y estructurado para el entrenamiento y alineamiento de modelos de lenguaje en narrativa… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-tales-es.textquestion-answering1K<n<10K0 likes25 downloads1y agoHugging Face27Somtharu181coder /SFT_Dataset_domain_social Nepali Social Studies MCQ — SFT Dataset A cleaned, deduplicated, bias-corrected instruction-tuning dataset of Nepali-language multiple-choice questions on social studies topics, derived from the Aya Dataset. Dataset Summary Rows 27,891 Language Nepali (ne / npi), Devanagari script Task type Instruction-following (single-turn MCQ Q&A) Domain Social studies (सामाजिक) — MCQ only License Apache-2.0 (permissive) Source CohereLabs/aya_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/SFT_Dataset_domain_social.textquestion-answering10K<n<100K0 likes24 downloads1mo agoHugging Face28somosnlp-hackathon-2026 /Onexe-QA-Dataset Dataset Card: Canarian Linguistic Evaluation Dataset (QA without Answers) Dataset Summary This dataset has been designed specifically for evaluating the dialectal, linguistic, and cultural understanding of Large Language Models (LLMs) within the context of Canarian Spanish. It contains 4,683 evaluation questions based on the official lexicon of the Academy of Canarian Language (Academia Canaria de la Lengua - ACL). Each record presents a linguistic query phrased… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/Onexe-QA-Dataset.textquestion-answering1K<n<10K0 likes22 downloads3mo agoHugging Face29Zyroxx66 /Somali-Reasoning-Dataset Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴 This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language. 🌟 What makes this unique? This is a Hybrid Dataset that combines two powerful sources: The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.texttext-generation10K<n<100K0 likes21 downloads6mo agoHugging Face30somosnlp-hackathon-2025 /exam_zh_multitopic_dialect_culture exam_zh_multitopic_dialect_culture This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge. 📚 Description The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories: 🗣️ Regional Dialect Tests These assess language understanding across major Chinese dialects and topolects: Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.tabularmultiple-choicen<1K0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.