CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01somosnlp /SMC Dataset Card for Spanish Medical Corpus (SMC) This dataset groups and organizes several datasets present in hugginface (e.g.: PlanTL-GOB-ES/cantemist-ner, PlanTL-GOB-ES/pharmaconer) and other public resources created by researchers with different formats (e.g.; MedLexSp ) to allow it to be a source of knowledge of large language models in Spanish for the medical domain. Dataset Card in Spanish Dataset Details Dataset Description Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/SMC.textquestion-answering1M<n<10M10 likes207 downloads2y agoHugging Face02somosnlp /instruct-legal-refugiados-es Dataset Card for AsistenciaRefugiados README in Spanish Spain is the third country with the highest number of asylum applications, receiving each year approximately more than 100,000 applications, and the third with the lowest number of approvals within the EU. The main objective of this project is to facilitate the tasks of NGOs in this field and other institutions and help them to obtain answers to questions (QA) related to refugee legislation in Spanish. With its… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/instruct-legal-refugiados-es.textquestion-answering10K<n<100K3 likes192 downloads2y agoHugging Face03khaledyusuf44 /somaliweb-v1 SomaliWeb v1 — Quality-filtered Somali web corpus 📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark 💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.tabulartext-generation100K<n<1M4 likes182 downloads4mo agoHugging Face04RoMoDataset /RoMo-SOMA-77 RoMo-SOMA-77 — RoMo Body+Hand Motion in 933-D Kimodo SOMA-77 Features RoMo-SOMA-77 is the RoMo body+hand corpus packed in a 933-dimensional Kimodo SOMA-77 motion-feature representation, paired with rich multi-level text descriptions. It is the publication target for the SOMA-based body-and-hand model family. Scope: paper-core (romo_official = True), matching RoMo-SMPL, RoMo-HML-263, and RoMo-272. A small number of clips are dropped where SOMA conversion produced non-finite… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-SOMA-77.texttext-to-3d100K<n<1M0 likes171 downloads4mo agoHugging Face05abdelstark /sommelier-xlam-single-call-splits sommelier xlam single-call splits Deterministic, deduplicated, single-tool-call train/validation/test splits derived from Salesforce/xlam-function-calling-60k (APIGen, CC-BY-4.0), produced by the sommelier pipeline for reproducible tool-calling fine-tuning. These are the exact splits used to train and evaluate abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora. Why single-call The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.texttext-generation10K<n<100K0 likes163 downloads3mo agoHugging Face06somosnlp /somos-alpaca-es Dataset Card de "somos-alpaca-es" Este conjunto de datos es una versión traducida del dataset Alpaca en Español. Este conjunto de datos sirve como referencia para el esfuerzo colaborativo de limpieza y mejora del dataset durante el hackathon SomosNLP 2023. Cuantas más personas y equipos participen mayor calidad final se podrá obtener. ➡️ ACTUALIZACION: Contribuye al etiquetado de la traducción de la versión limpia: somos-clean-alpaca-es El reto A continuación se… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/somos-alpaca-es.texttext-generation10K<n<100K8 likes94 downloads4y agoHugging Face07somosnlp /recetas-cocinatexttable-question-answering10K<n<100K4 likes88 downloads3y agoHugging Face08Cour-de-cassation /alpaca_ccass_motivations_sommaires_titres Training dataset for summarizing and titling decisions of the French Court of cassation based on motivations This alpaca-format dataset is designed to train models for summarizing and titling French Supreme Court decisions based on the grounds of them. Created with a view to producing metadata for decisions not published in the bulletin, this dataset aims to simplify the development of annotation and categorization tools, and is positioned as a facilitator for jurisprudential… See the full description on the dataset page: https://huggingface.co/datasets/Cour-de-cassation/alpaca_ccass_motivations_sommaires_titres.textsummarization10K<n<100K3 likes85 downloads1y agoHugging Face09Zyroxx66 /somali-master-pretraining-corpus 🇸🇴 Somali Master Pretraining Corpus (176.5k Rows) The Somali Master Pretraining Corpus is a curated, balanced dataset designed for Continued Pre-Training (CPT) and foundational pre-training of Large Language Models (LLMs) in the Somali language (Af-Soomaali). It addresses the fundamental challenges of low-resource NLP for Somali by combining quality-filtered web knowledge, structured modern domain knowledge, and synthetic narrative intelligence (TinyStories). 🎯… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/somali-master-pretraining-corpus.texttext-generation100K<n<1M0 likes80 downloads1mo agoHugging Face10yacdev /somali-100k-native-conversations 🇸🇴 Somali High-Diversity Multi-Turn Conversational SFT Dataset A state-of-the-art, 100.00% unique (zero duplicate responses) multi-turn conversational dataset in authentic Somali (Af-Soomaali) across 25 real-world knowledge domains. 🌟 Quality Standards: 100% Unique Assistant Responses: Guaranteed zero template repetition (23,334 / 23,334 unique turns). Grounded Knowledge: Spanning Python coding, web dev, Git/Linux, cybersecurity, diabetes & health, business… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-native-conversations.texttext-generation10K<n<100K0 likes75 downloads22d agoHugging Face11AngelGabrielTroncoso /dataset-aeroespacial-cultural-somosnlp LATAM Aerospace History QA Descripción General LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica. El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos. La colección está especializada en: historia aeroespacial, programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.textquestion-answering100K<n<1M0 likes69 downloads4mo agoHugging Face12somosnlp-hackathon-2023 /Habilidades_Agente_v1 Description Español: Presentamos un conjunto de datos que presenta tres partes principales: 1. Dataset sobre habilidades blandas. 2. Dataset de conversaciones empresariales entre agentes y clientes. 3. Dataset curado de Alpaca en español: Este dataset toma como base el dataset https://huggingface.co/datasets/somosnlp/somos-alpaca-es, y fue curado con la herramienta Argilla, alcanzando 9400 registros curados. Los datos están estructurados en torno a un método que se describe… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/Habilidades_Agente_v1.texttext-generation10K<n<100K22 likes68 downloads3y agoHugging Face13SomyaSaraswati /psychoanalysis-dataset-100k Psychoanalysis Synthetic Instruction Dataset (v1, 100k) Domain: psychoanalytic reflection / therapy-style dialoguesLocale: English + Hinglish (India context)Size: 100,000 rows; 10 shards × 10k JSONL Schema Chat-style messages + instruction/input/output + safety + metadata.Educational only; not clinical advice. Split train only (create validation downstream with train_test_split). Generation Notes Synthetic templates + slot-filling; no… See the full description on the dataset page: https://huggingface.co/datasets/SomyaSaraswati/psychoanalysis-dataset-100k.texttext-generation100K<n<1M0 likes68 downloads1y agoHugging Face14somosnlp /constitucion-politica-del-peru-1993-qa Dataset Summary Compuesto por unos 2075 registros que contienen los campos: pregunta: pregunta que sirve como una instrucción o consulta sobre algún aspecto de la Constitución Política del Perú de 1993. respuesta: La respuesta proporcionada para cada pregunta es un contexto relevante que ayuda a resolver la consulta. Este contexto es un extracto de la Constitución. fuente: Para cada respuesta, se indica el capítulo y/o artículo de la Constitución Política del Perú de 1993… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/constitucion-politica-del-peru-1993-qa.textsummarization1K<n<10K2 likes54 downloads2y agoHugging Face15khaledyusuf44 /somalibench-v0 SomaliBench v0 The first native-author-verified Somali safety evaluation benchmark. 100 harmful-intent prompts drawn from HarmBench (Mazeika et al. 2024) and AdvBench (Zou et al. 2023), translated into Somali by a native speaker (Khalid Yusuf Dahir, Mogadishu) and released as an evaluation set for measuring multilingual safety alignment. Why this exists Somali has 15–20 million speakers and zero native-verified safety evaluation resources. SomaliBench fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somalibench-v0.texttext-classificationn<1K0 likes45 downloads4mo agoHugging Face16Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes43 downloads16d agoHugging Face17somaxsoma /tac-closing-efficiency-sft TAC closing-efficiency slice 500 synthetic multi-turn tool-use trajectories that teach an agent to close bookings decisively — the welfare-neutral capability piece of the tool-use SFT mix used to train somaxsoma/qwen2.5-7b-tac-recovery-sft. What it teaches Built to fix the dominant failure mode observed on the TAC benchmark — the model reformulating search keywords in a loop and never closing a booking. Three patterns: settle/browse (200): after failed keyword… See the full description on the dataset page: https://huggingface.co/datasets/somaxsoma/tac-closing-efficiency-sft.texttext-generationn<1K0 likes42 downloads27d agoHugging Face18Somtharu181coder /number_of_death_by_sex_hermes_calling Nepal Education Enrollment Statistics – Hermes Function-Calling Dataset 1. Overview This dataset contains 16,760 single-turn function-calling records in Hermes / ShareGPT conversation format. Each record pairs a natural-language request for education enrollment statistics with the exact tool call that satisfies it. Tool-call arguments are grounded in the administrative hierarchy of an Excel source workbook (Annex 4 – Enrollment Details): every province, district… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/number_of_death_by_sex_hermes_calling.texttext-generation10K<n<100K0 likes42 downloads4d agoHugging Face19Siddhu077 /some FinEE Dataset Dataset Description A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks. Languages English (en) - 86% Hindi (hi) - 3% Tamil (ta) - 3% Telugu (te) - 3% Bengali (bn) - 3% Kannada (kn) - 2% Supported Transaction Types UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.texttoken-classification100K<n<1M0 likes41 downloads2mo agoHugging Face20someoneatemylastsliceofpizza /claude-tools-sft-merged claude-tools-sft-merged Merged SFT dataset in ChatML format (<|im_start|> / <|im_end|>), deduplicated and filtered, ready for instruction fine-tuning. Covers general instruction following, reasoning (<think> traces), function calling, coding, and multi-turn conversation. Statistics Metric Value Total examples 298,979 Duplicates removed 44,928 Min length (chars) 142 Median length (chars) 3,033 Mean length (chars) 4,231 P90 length (chars) 11… See the full description on the dataset page: https://huggingface.co/datasets/someoneatemylastsliceofpizza/claude-tools-sft-merged.texttext-generation100K<n<1M3 likes40 downloads5mo agoHugging Face21Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes36 downloads1mo agoHugging Face22anonymous-somebody /CHIMERA CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning CHIMERA is a compact, high-difficulty synthetic reasoning dataset with long Chain-of-Thought (CoT) trajectories and broad scientific coverage. It is designed to support reasoning post-training for large language models. All examples are LLM-generated and automatically verified without human annotation. Total: 9,225 problems Subjects: 8 Topics: 1,179 Overview Recent reasoning advances rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-somebody/CHIMERA.texttext-generation1K<n<10K0 likes35 downloads5mo agoHugging Face23IbraahimLab /fineweb-somali FineWeb Somali Dataset A curated collection of Somali language content from BBC Somali, designed for training and evaluating small language models on low-resource languages. Dataset Description This dataset contains 4,910 high-quality Somali language articles scraped from BBC Somali's website. The content covers diverse topics including news, culture, technology, sports, and human interest stories, providing a rich corpus for Somali language model training.… See the full description on the dataset page: https://huggingface.co/datasets/IbraahimLab/fineweb-somali.texttext-generation1K<n<10K0 likes30 downloads8mo agoHugging Face24abdelstark /sommelier-xlam-single-call-splits-fr sommelier-xlam-single-call-splits-fr French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring. How it was built The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.texttext-generation10K<n<100K0 likes30 downloads3mo agoHugging Face25Somtharu181coder /educational_domain_dataset Nepali Grounded Education QA (OpenHermes-format) A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about student enrollment statistics from Nepal's Ministry of Education. Every answer is anchored to a real numeric value pulled from government open data — nothing in the answers is model-hallucinated. Dataset Summary Rows 611 Language Nepali (Devanagari script) Format ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.textquestion-answeringn<1K0 likes30 downloads1mo agoHugging Face26michsethowusu /Code-170k-somali Dataset Description Code-170k-somali is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Somali, making coding education accessible to Somali speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Somali language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-somali.texttext-generation100K<n<1M1 likes29 downloads11mo agoHugging Face27maanka2 /somali-web-corpus SOMALI-WEB-CORPUS V1 This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language. Dataset Details Language: Somali (so) Format: JSON lines (.jsonl) Data Structure: Each record has a single text field containing a cleaned paragraph. Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.texttext-generation100K<n<1M1 likes26 downloads4mo agoHugging Face28Somtharu181coder /cyber_security Digital Literacy & Cybersecurity Nepali SFT Dataset Dataset Overview This dataset is a Nepali-language Supervised Fine-Tuning (SFT) dataset focused on digital literacy and cybersecurity. The dataset contains 1,000 valid JSONL records designed for instruction-following tasks. Each record contains a human instruction and a corresponding GPT-generated response. Dataset Statistics Property Value Total records 1,000 Valid JSONL rows 1,000… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/cyber_security.tabulartext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face29Somtharu181coder /SFT_Dataset_domain_social Nepali Social Studies MCQ — SFT Dataset A cleaned, deduplicated, bias-corrected instruction-tuning dataset of Nepali-language multiple-choice questions on social studies topics, derived from the Aya Dataset. Dataset Summary Rows 27,891 Language Nepali (ne / npi), Devanagari script Task type Instruction-following (single-turn MCQ Q&A) Domain Social studies (सामाजिक) — MCQ only License Apache-2.0 (permissive) Source CohereLabs/aya_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/SFT_Dataset_domain_social.textquestion-answering10K<n<100K0 likes24 downloads1mo agoHugging Face30somosnlp-hackathon-2025 /ibero-characters-es Conjunto de datos de personajes de mitos y leyendas iberoamericanos. ⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y mejorar la cobertura de imágenes. 📚 Descripción Dataset de personajes míticos y legendarios de Iberoamérica, diseñado para preservar y promover el patrimonio cultural a través de la inteligencia artificial. 🌟 Motivación e Impacto 📱 Preservación Digital: Conservación del… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-characters-es.imagevideo-text-to-textn<1K0 likes23 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.