datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SMC
Dataset Card for Spanish Medical Corpus (SMC)
This dataset groups and organizes several datasets present in hugginface (e.g.: PlanTL-GOB-ES/cantemist-ner, PlanTL-GOB-ES/pharmaconer)
and other public resources created by researchers with different formats (e.g.; MedLexSp )
to allow it to be a source of knowledge of large language models in Spanish for the medical domain.
Dataset Card in Spanish
Dataset Details
Dataset Description
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/SMC.instruct-legal-refugiados-es
Dataset Card for AsistenciaRefugiados
README in Spanish
Spain is the third country with the highest number of asylum applications, receiving each year approximately more than 100,000 applications, and the third with the lowest number of approvals within the EU.
The main objective of this project is to facilitate the tasks of NGOs in this field and other institutions and help them to obtain answers to questions (QA) related to refugee legislation in Spanish. With its… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/instruct-legal-refugiados-es.somaliweb-v1
SomaliWeb v1 — Quality-filtered Somali web corpus
📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus
SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.RoMo-SOMA-77
RoMo-SOMA-77 — RoMo Body+Hand Motion in 933-D Kimodo SOMA-77 Features
RoMo-SOMA-77 is the RoMo body+hand corpus packed in a 933-dimensional Kimodo SOMA-77 motion-feature representation, paired with rich multi-level text descriptions. It is the publication target for the SOMA-based body-and-hand model family.
Scope: paper-core (romo_official = True), matching RoMo-SMPL, RoMo-HML-263, and RoMo-272. A small number of clips are dropped where SOMA conversion produced non-finite… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-SOMA-77.sommelier-xlam-single-call-splits
sommelier xlam single-call splits
Deterministic, deduplicated, single-tool-call train/validation/test
splits derived from
Salesforce/xlam-function-calling-60k
(APIGen, CC-BY-4.0), produced by the
sommelier pipeline for
reproducible tool-calling fine-tuning. These are the exact splits used to
train and evaluate
abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora.
Why single-call
The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.somos-alpaca-es
Dataset Card de "somos-alpaca-es"
Este conjunto de datos es una versión traducida del dataset Alpaca en Español.
Este conjunto de datos sirve como referencia para el esfuerzo colaborativo de limpieza y mejora del dataset durante el hackathon SomosNLP 2023.
Cuantas más personas y equipos participen mayor calidad final se podrá obtener.
➡️ ACTUALIZACION: Contribuye al etiquetado de la traducción de la versión limpia: somos-clean-alpaca-es
El reto
A continuación se… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/somos-alpaca-es.recetas-cocinaalpaca_ccass_motivations_sommaires_titres
Training dataset for summarizing and titling decisions of the French Court of cassation based on motivations
This alpaca-format dataset is designed to train models for summarizing and titling French Supreme Court decisions based on the grounds of them. Created with a view to producing metadata for decisions not published in the bulletin, this dataset aims to simplify the development of annotation and categorization tools, and is positioned as a facilitator for jurisprudential… See the full description on the dataset page: https://huggingface.co/datasets/Cour-de-cassation/alpaca_ccass_motivations_sommaires_titres.somali-master-pretraining-corpus
🇸🇴 Somali Master Pretraining Corpus (176.5k Rows)
The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali).
It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories).
🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.somali-100k-native-conversations
🇸🇴 Somali High-Diversity Multi-Turn Conversational SFT Dataset
A state-of-the-art, 100.00% unique (zero duplicate responses) multi-turn conversational dataset in authentic Somali (Af-Soomaali) across 25 real-world knowledge domains.
🌟 Quality Standards:
100% Unique Assistant Responses: Guaranteed zero template repetition (23,334 / 23,334 unique turns).
Grounded Knowledge: Spanning Python coding, web dev, Git/Linux, cybersecurity, diabetes & health, business… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-native-conversations.dataset-aeroespacial-cultural-somosnlp
LATAM Aerospace History QA
Descripción General
LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica.
El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos.
La colección está especializada en:
historia aeroespacial,
programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.Habilidades_Agente_v1
Description
Español:
Presentamos un conjunto de datos que presenta tres partes principales:
1. Dataset sobre habilidades blandas.
2. Dataset de conversaciones empresariales entre agentes y clientes.
3. Dataset curado de Alpaca en español: Este dataset toma como base el dataset https://huggingface.co/datasets/somosnlp/somos-alpaca-es,
y fue curado con la herramienta Argilla, alcanzando 9400 registros curados.
Los datos están estructurados en torno a un método que se describe… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/Habilidades_Agente_v1.psychoanalysis-dataset-100k
Psychoanalysis Synthetic Instruction Dataset (v1, 100k)
Domain: psychoanalytic reflection / therapy-style dialoguesLocale: English + Hinglish (India context)Size: 100,000 rows; 10 shards × 10k JSONL
Schema
Chat-style messages + instruction/input/output + safety + metadata.Educational only; not clinical advice.
Split
train only (create validation downstream with train_test_split).
Generation Notes
Synthetic templates + slot-filling; no… See the full description on the dataset page: https://huggingface.co/datasets/SomyaSaraswati/psychoanalysis-dataset-100k.constitucion-politica-del-peru-1993-qa
Dataset Summary
Compuesto por unos 2075 registros que contienen los campos:
pregunta: pregunta que sirve como una instrucción o consulta sobre algún aspecto de la Constitución Política del Perú de 1993.
respuesta: La respuesta proporcionada para cada pregunta es un contexto relevante que ayuda a resolver la consulta. Este contexto es un extracto de la Constitución.
fuente: Para cada respuesta, se indica el capítulo y/o artículo de la Constitución Política del Perú de 1993… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/constitucion-politica-del-peru-1993-qa.somalibench-v0
SomaliBench v0
The first native-author-verified Somali safety evaluation benchmark.
100 harmful-intent prompts drawn from HarmBench (Mazeika et al. 2024) and
AdvBench (Zou et al. 2023), translated into Somali by a native speaker
(Khalid Yusuf Dahir, Mogadishu) and released as an evaluation set for
measuring multilingual safety alignment.
Why this exists
Somali has 15–20 million speakers and zero native-verified safety
evaluation resources. SomaliBench fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somalibench-v0.Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.tac-closing-efficiency-sft
TAC closing-efficiency slice
500 synthetic multi-turn tool-use trajectories that teach an agent to close bookings decisively — the welfare-neutral capability piece of the tool-use SFT mix used to train somaxsoma/qwen2.5-7b-tac-recovery-sft.
What it teaches
Built to fix the dominant failure mode observed on the TAC benchmark — the model reformulating search keywords in a loop and never closing a booking. Three patterns:
settle/browse (200): after failed keyword… See the full description on the dataset page: https://huggingface.co/datasets/somaxsoma/tac-closing-efficiency-sft.number_of_death_by_sex_hermes_calling
Nepal Education Enrollment Statistics – Hermes Function-Calling Dataset
1. Overview
This dataset contains 16,760 single-turn function-calling records in Hermes / ShareGPT conversation format. Each record pairs a natural-language request for education enrollment statistics with the exact tool call that satisfies it. Tool-call arguments are grounded in the administrative hierarchy of an Excel source workbook (Annex 4 – Enrollment Details): every province, district… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/number_of_death_by_sex_hermes_calling.some
FinEE Dataset
Dataset Description
A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks.
Languages
English (en) - 86%
Hindi (hi) - 3%
Tamil (ta) - 3%
Telugu (te) - 3%
Bengali (bn) - 3%
Kannada (kn) - 2%
Supported Transaction Types
UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.claude-tools-sft-merged
claude-tools-sft-merged
Merged SFT dataset in ChatML format (<|im_start|> / <|im_end|>),
deduplicated and filtered, ready for instruction fine-tuning.
Covers general instruction following, reasoning (<think> traces),
function calling, coding, and multi-turn conversation.
Statistics
Metric
Value
Total examples
298,979
Duplicates removed
44,928
Min length (chars)
142
Median length (chars)
3,033
Mean length (chars)
4,231
P90 length (chars)
11… See the full description on the dataset page: https://huggingface.co/datasets/someoneatemylastsliceofpizza/claude-tools-sft-merged.science_behavioral_and_domain_diversity_dataset
Nepali Science SFT Dataset — Clean Candidate
A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script.
This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering.
Dataset Overview
Property
Value
Dataset file
clean_candidate.jsonl
Records
29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.CHIMERA
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation.
Total: 9,225 problems
Subjects: 8
Topics: 1,179
Overview
Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.fineweb-somali
FineWeb Somali Dataset
A curated collection of Somali language content from BBC Somali, designed for training and evaluating small language models on low-resource languages.
Dataset Description
This dataset contains 4,910 high-quality Somali language articles scraped from BBC Somali's website. The content covers diverse topics including news, culture, technology, sports, and human interest stories, providing a rich corpus for Somali language model training.… See the full description on the dataset page: https://huggingface.co/datasets/IbraahimLab/fineweb-somali.sommelier-xlam-single-call-splits-fr
sommelier-xlam-single-call-splits-fr
French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring.
How it was built
The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.educational_domain_dataset
Nepali Grounded Education QA (OpenHermes-format)
A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about
student enrollment statistics from Nepal's Ministry of Education. Every answer is
anchored to a real numeric value pulled from government open data — nothing in the
answers is model-hallucinated.
Dataset Summary
Rows
611
Language
Nepali (Devanagari script)
Format
ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.Code-170k-somali
Dataset Description
Code-170k-somali is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Somali, making coding education accessible to Somali speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Somali language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-somali.somali-web-corpus
SOMALI-WEB-CORPUS V1
This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language.
Dataset Details
Language: Somali (so)
Format: JSON lines (.jsonl)
Data Structure: Each record has a single text field containing a cleaned paragraph.
Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.cyber_security
Digital Literacy & Cybersecurity Nepali SFT Dataset
Dataset Overview
This dataset is a Nepali-language Supervised Fine-Tuning (SFT) dataset focused on digital literacy and cybersecurity.
The dataset contains 1,000 valid JSONL records designed for instruction-following tasks. Each record contains a human instruction and a corresponding GPT-generated response.
Dataset Statistics
Property
Value
Total records
1,000
Valid JSONL rows
1,000… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/cyber_security.SFT_Dataset_domain_social
Nepali Social Studies MCQ — SFT Dataset
A cleaned, deduplicated, bias-corrected instruction-tuning dataset of Nepali-language
multiple-choice questions on social studies topics, derived from the Aya Dataset.
Dataset Summary
Rows
27,891
Language
Nepali (ne / npi), Devanagari script
Task type
Instruction-following (single-turn MCQ Q&A)
Domain
Social studies (सामाजिक) — MCQ only
License
Apache-2.0 (permissive)
Source
CohereLabs/aya_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/SFT_Dataset_domain_social.ibero-characters-es
Conjunto de datos de personajes de mitos y leyendas iberoamericanos.
⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y mejorar la cobertura de imágenes.
📚 Descripción
Dataset de personajes míticos y legendarios de Iberoamérica, diseñado para preservar y promover el patrimonio cultural a través de la inteligencia artificial.
🌟 Motivación e Impacto
📱 Preservación Digital: Conservación del… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-characters-es.
