datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RecetasDeLaAbuela
Motivación inicial
Este corpus ha sido creado durante el Hackathon SomosNLP Marzo 2024: #Somos600M (https://somosnlp.org/hackathon).
Responde a una de las propuestas somosnlp sobre 'Recetas típicas por país/zona geográfica'.
Nombre del Proyecto
Este corpus o dataset se llama 'RecetasDeLaAbuel@' y es un homenaje a todas nuestr@s abuel@s que nos han enseñado a cocinar. Se trata de la mayor y más completa colección de recetas open-source en español de países… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/RecetasDeLaAbuela.SMC
Dataset Card for Spanish Medical Corpus (SMC)
This dataset groups and organizes several datasets present in hugginface (e.g.: PlanTL-GOB-ES/cantemist-ner, PlanTL-GOB-ES/pharmaconer)
and other public resources created by researchers with different formats (e.g.; MedLexSp )
to allow it to be a source of knowledge of large language models in Spanish for the medical domain.
Dataset Card in Spanish
Dataset Details
Dataset Description
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/SMC.instruct-legal-refugiados-es
Dataset Card for AsistenciaRefugiados
README in Spanish
Spain is the third country with the highest number of asylum applications, receiving each year approximately more than 100,000 applications, and the third with the lowest number of approvals within the EU.
The main objective of this project is to facilitate the tasks of NGOs in this field and other institutions and help them to obtain answers to questions (QA) related to refugee legislation in Spanish. With its… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/instruct-legal-refugiados-es.nexttoken-model-1-dataset-sft
NextToken Model 1 SFT dataset (v4)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
846 scheme/product source documents across 57 schemes/products, chunked
into 1,445 passages.
v4 vs v3: v3 merged in a second batch (6,664 rows) without… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-1-dataset-sft.dataset-aeroespacial-cultural-somosnlp
LATAM Aerospace History QA
Descripción General
LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica.
El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos.
La colección está especializada en:
historia aeroespacial,
programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.soma-skincare-qarecetasdelaabuela_genstruct_it
Descripción
Dataset creado para la hackathon #Somos600M con el objetivo de entrenar un modelo que pueda recomendar recetas de paises hispanohablantes.
Este conjunto de datos consiste en pregunta-respuesta y fue elaborado a partir de un contexto usando Genstruct-7B y distilabel.
Elaborado a partir del dataset en crudo somosnlp/RecetasDeLaAbuela elaborado por el equipo recetasdelaabuela mediante web scraping.
Origen del Dataset
El dataset se obtuvo mediante web… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/recetasdelaabuela_genstruct_it.constitucion-politica-del-peru-1993-qa
Dataset Summary
Compuesto por unos 2075 registros que contienen los campos:
pregunta: pregunta que sirve como una instrucción o consulta sobre algún aspecto de la Constitución Política del Perú de 1993.
respuesta: La respuesta proporcionada para cada pregunta es un contexto relevante que ayuda a resolver la consulta. Este contexto es un extracto de la Constitución.
fuente: Para cada respuesta, se indica el capítulo y/o artículo de la Constitución Política del Perú de 1993… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/constitucion-politica-del-peru-1993-qa.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.LingComp_QA
Dataset Card for LingComp_QA, un corpus educativo de lingüística computacional en español
Dataset Details
Dataset Description
Curated by: Jorge Zamora Rey, Isabel Moyano Moreno, Mario Crespo Miguel
Funded by: SomosNLP, HuggingFace, Argilla, Instituto de Lingüística Aplicada de la Universidad de Cádiz
Language(s) (NLP): es-ES
License: apache-2.0
Dataset Sources
Repository: https://github.com/reddrex/lingcomp_QA/tree/main
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/LingComp_QA.LLM_SQL_BaseDatosEspanol
Usos
Usos directos
El objetivo principal de este dataset es proporcionar ejemplos simples para el fine-tuning de modelos
de procesamiento de lenguaje natural (NLP) en el contexto de consultas SQL.
Usos fuera de mira
Podria usarse para el entrenamiento de una IA que sirva como creadora de base de datos artificiales
Estructura del conjunto de datos
Question: Es la pegunta que el usuario le dara al chatbot
Answer: La respuesta el que chatbot le… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/LLM_SQL_BaseDatosEspanol.wikihow_es
Citation
@misc{quispe2024wikihowes,
author = {Quispe Castillo, David Alonso},
title = {WikiHowES},
month = March,
year = 2024,
url = {https://huggingface.co/datasets/somosnlp/wikihow_es}
}
sommbench
SommBench
SommBench is a multilingual benchmark for assessing sommelier expertise in large language models. It comprises 3,024 expert-curated examples across eight languages (da, de, en, es, fi, it, sk, sv), designed by professional sommeliers to evaluate sensory grounding, factual wine knowledge, and practical pairing skills through three tasks: WTQA, WFC, and FWP.
Configs
wtqa — Wine Theory Question Answering (1,024 examples)
Multiple-choice questions about… See the full description on the dataset page: https://huggingface.co/datasets/sommify/sommbench.justicio_evaluacion_ideonidad_preguntas_legalesEste dataset nos permitirá evaluar la idoneidad de las preguntas generadas para su uso dentro de la plataforma Justicio, un archivero que permite consultar desde una interfaz chat las distintas legislaciones, tanto a nivel nacional derivadas del Boletín Oficial del Estado, así como de las derivadas de las distintas Comunidades Autónomas.
Internamente, Justicio utiliza un esquema de tipo RAG (Retrieval-Augmented Generation) en el que se localizan aquellos fragmentos almacenados más similares a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/justicio_evaluacion_ideonidad_preguntas_legales.Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.recetasdelaabuela_it
Nombre del dataset
Este dataset se llama 'RecetasDeLaAbuel@' y es un homenaje a todas nuestr@s abuel@s que nos han enseñado a cocinar. Se trata de la mayor y más completa colección de recetas open-source en español de países hispanoamericanos.
Descripción
Dataset creado on el objetivo de entrenar un modelo que pueda recomendar recetas de paises hispanohablantes. Nuestra IA responderá a cuestiones de los sigientes tipos: 'Qué puedo cocinar con 3 ingredientes?', 'Dime una… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/recetasdelaabuela_it.Annex_7_Educational_Indicators_2024-2025
Grounded Educational Indicators — Nepali SFT Dataset
File: grounded_education_indicators_nepali_sft.jsonl
Records: 10,091
Language: Nepali (ne / ISO 639-3 npi), Devanagari script
License: Apache-2.0 (permissive)
Format: JSON Lines, ShareGPT-style conversational (human / gpt turns)
Task type: Grounded question answering over structured (tabular) educational statistics
Size on disk: ~15.5 MB
1. Overview
This dataset is a supervised fine-tuning (SFT) corpus of 10… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Annex_7_Educational_Indicators_2024-2025.SOMAJGYAAN
SomajGyaan (সমাজজ্ঞান) - Bangla MCQ Dataset
📊 Dataset Description
SomajGyaan (সমাজজ্ঞান) is a comprehensive Bangla multiple-choice question dataset featuring 4,234 questions across 7 academic categories with ~12,000 unique answer options.
Dataset Summary
Total Questions: 4,234
Unique Answer Options: ~12,000
Answer Diversity: 70.8%
Language: Bangla (Bengali)
Categories: 7 (History, Economics, Geography, Politics, Social Studies, Law… See the full description on the dataset page: https://huggingface.co/datasets/farihashifa/SOMAJGYAAN.science_behavioral_and_domain_diversity_dataset
Nepali Science SFT Dataset — Clean Candidate
A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script.
This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering.
Dataset Overview
Property
Value
Dataset file
clean_candidate.jsonl
Records
29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
Overview
Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.educational_domain_dataset
Nepali Grounded Education QA (OpenHermes-format)
A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about
student enrollment statistics from Nepal's Ministry of Education. Every answer is
anchored to a real numeric value pulled from government open data — nothing in the
answers is model-hallucinated.
Dataset Summary
Rows
611
Language
Nepali (Devanagari script)
Format
ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.Code-170k-somali
Dataset Description
Code-170k-somali is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Somali, making coding education accessible to Somali speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Somali language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-somali.SMC-instruct
Dataset Card for Spanish Medical Corpus (SMC)
This dataset groups and organizes several datasets present in hugginface (e.g.: PlanTL-GOB-ES/cantemist-ner, PlanTL-GOB-ES/pharmaconer)
and other public resources created by researchers with different formats (e.g.; MedLexSp )
to allow it to be a source of knowledge of large language models in Spanish for the medical domain.
Dataset Card in Spanish
Dataset Details
Dataset Description
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/SMC-instruct.ec-prompts-refranescenia-team-sabiduriapopularibero-tales-es
Conjunto de datos de historias sintéticas de mitos y leyendas iberoamericanos.
⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y refinar el proceso de generación para mejorar la diversidad y calidad narrativa.
📚 Descripción
Dataset de historias sintéticas generadas a partir de mitos y leyendas de Iberoamérica, curado y estructurado para el entrenamiento y alineamiento de modelos de lenguaje en narrativa… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-tales-es.SFT_Dataset_domain_social
Nepali Social Studies MCQ — SFT Dataset
A cleaned, deduplicated, bias-corrected instruction-tuning dataset of Nepali-language
multiple-choice questions on social studies topics, derived from the Aya Dataset.
Dataset Summary
Rows
27,891
Language
Nepali (ne / npi), Devanagari script
Task type
Instruction-following (single-turn MCQ Q&A)
Domain
Social studies (सामाजिक) — MCQ only
License
Apache-2.0 (permissive)
Source
CohereLabs/aya_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/SFT_Dataset_domain_social.Onexe-QA-Dataset
Dataset Card: Canarian Linguistic Evaluation Dataset (QA without Answers)
Dataset Summary
This dataset has been designed specifically for evaluating the dialectal, linguistic, and cultural understanding of Large Language Models (LLMs) within the context of Canarian Spanish.
It contains 4,683 evaluation questions based on the official lexicon of the Academy of Canarian Language (Academia Canaria de la Lengua - ACL). Each record presents a linguistic query phrased… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/Onexe-QA-Dataset.Somali-Reasoning-Dataset
Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴
This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language.
🌟 What makes this unique?
This is a Hybrid Dataset that combines two powerful sources:
The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.exam_zh_multitopic_dialect_culture
exam_zh_multitopic_dialect_culture
This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge.
📚 Description
The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories:
🗣️ Regional Dialect Tests
These assess language understanding across major Chinese dialects and topolects:
Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.
