CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abdelstark /sommelier-xlam-single-call-splits sommelier xlam single-call splits Deterministic, deduplicated, single-tool-call train/validation/test splits derived from Salesforce/xlam-function-calling-60k (APIGen, CC-BY-4.0), produced by the sommelier pipeline for reproducible tool-calling fine-tuning. These are the exact splits used to train and evaluate abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora. Why single-call The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.texttext-generation10K<n<100K0 likes163 downloads3mo agoHugging Face02Cour-de-cassation /alpaca_ccass_motivations_sommaires_titres Training dataset for summarizing and titling decisions of the French Court of cassation based on motivations This alpaca-format dataset is designed to train models for summarizing and titling French Supreme Court decisions based on the grounds of them. Created with a view to producing metadata for decisions not published in the bulletin, this dataset aims to simplify the development of annotation and categorization tools, and is positioned as a facilitator for jurisprudential… See the full description on the dataset page: https://huggingface.co/datasets/Cour-de-cassation/alpaca_ccass_motivations_sommaires_titres.textsummarization10K<n<100K3 likes85 downloads1y agoHugging Face03AngelGabrielTroncoso /dataset-aeroespacial-cultural-somosnlp LATAM Aerospace History QA Descripción General LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica. El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos. La colección está especializada en: historia aeroespacial, programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.textquestion-answering100K<n<1M0 likes69 downloads4mo agoHugging Face04SomyaSaraswati /psychoanalysis-dataset-100k Psychoanalysis Synthetic Instruction Dataset (v1, 100k) Domain: psychoanalytic reflection / therapy-style dialoguesLocale: English + Hinglish (India context)Size: 100,000 rows; 10 shards × 10k JSONL Schema Chat-style messages + instruction/input/output + safety + metadata.Educational only; not clinical advice. Split train only (create validation downstream with train_test_split). Generation Notes Synthetic templates + slot-filling; no… See the full description on the dataset page: https://huggingface.co/datasets/SomyaSaraswati/psychoanalysis-dataset-100k.texttext-generation100K<n<1M0 likes68 downloads1y agoHugging Face05Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes43 downloads16d agoHugging Face06somaxsoma /tac-closing-efficiency-sft TAC closing-efficiency slice 500 synthetic multi-turn tool-use trajectories that teach an agent to close bookings decisively — the welfare-neutral capability piece of the tool-use SFT mix used to train somaxsoma/qwen2.5-7b-tac-recovery-sft. What it teaches Built to fix the dominant failure mode observed on the TAC benchmark — the model reformulating search keywords in a loop and never closing a booking. Three patterns: settle/browse (200): after failed keyword… See the full description on the dataset page: https://huggingface.co/datasets/somaxsoma/tac-closing-efficiency-sft.texttext-generationn<1K0 likes42 downloads27d agoHugging Face07Somtharu181coder /number_of_death_by_sex_hermes_calling Nepal Education Enrollment Statistics – Hermes Function-Calling Dataset 1. Overview This dataset contains 16,760 single-turn function-calling records in Hermes / ShareGPT conversation format. Each record pairs a natural-language request for education enrollment statistics with the exact tool call that satisfies it. Tool-call arguments are grounded in the administrative hierarchy of an Excel source workbook (Annex 4 – Enrollment Details): every province, district… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/number_of_death_by_sex_hermes_calling.texttext-generation10K<n<100K0 likes42 downloads4d agoHugging Face08Siddhu077 /some FinEE Dataset Dataset Description A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks. Languages English (en) - 86% Hindi (hi) - 3% Tamil (ta) - 3% Telugu (te) - 3% Bengali (bn) - 3% Kannada (kn) - 2% Supported Transaction Types UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.texttoken-classification100K<n<1M0 likes41 downloads2mo agoHugging Face09Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes36 downloads1mo agoHugging Face10abdelstark /sommelier-xlam-single-call-splits-fr sommelier-xlam-single-call-splits-fr French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring. How it was built The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.texttext-generation10K<n<100K0 likes30 downloads3mo agoHugging Face11Somtharu181coder /educational_domain_dataset Nepali Grounded Education QA (OpenHermes-format) A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about student enrollment statistics from Nepal's Ministry of Education. Every answer is anchored to a real numeric value pulled from government open data — nothing in the answers is model-hallucinated. Dataset Summary Rows 611 Language Nepali (Devanagari script) Format ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.textquestion-answeringn<1K0 likes30 downloads1mo agoHugging Face12maanka2 /somali-web-corpus SOMALI-WEB-CORPUS V1 This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language. Dataset Details Language: Somali (so) Format: JSON lines (.jsonl) Data Structure: Each record has a single text field containing a cleaned paragraph. Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.texttext-generation100K<n<1M1 likes26 downloads4mo agoHugging Face13Somtharu181coder /cyber_security Digital Literacy & Cybersecurity Nepali SFT Dataset Dataset Overview This dataset is a Nepali-language Supervised Fine-Tuning (SFT) dataset focused on digital literacy and cybersecurity. The dataset contains 1,000 valid JSONL records designed for instruction-following tasks. Each record contains a human instruction and a corresponding GPT-generated response. Dataset Statistics Property Value Total records 1,000 Valid JSONL rows 1,000… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/cyber_security.tabulartext-generation1K<n<10K0 likes24 downloads2mo agoHugging Face14Somtharu181coder /SFT_Dataset_domain_social Nepali Social Studies MCQ — SFT Dataset A cleaned, deduplicated, bias-corrected instruction-tuning dataset of Nepali-language multiple-choice questions on social studies topics, derived from the Aya Dataset. Dataset Summary Rows 27,891 Language Nepali (ne / npi), Devanagari script Task type Instruction-following (single-turn MCQ Q&A) Domain Social studies (सामाजिक) — MCQ only License Apache-2.0 (permissive) Source CohereLabs/aya_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/SFT_Dataset_domain_social.textquestion-answering10K<n<100K0 likes24 downloads1mo agoHugging Face15sayurio /somewhereinblog-article Somewhereinblog Article Archive Overview This repository contains a large-scale text dataset scraped from m.somewhereinblog.net, the largest and first-ever Bengali community blogging platform. The primary goal of this archive is to preserve a massive collection of purely human-written blog posts, personal stories, socio-political opinions, and community discussions, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/somewhereinblog-article.imagetext-generation10K<n<100K1 likes23 downloads6mo agoHugging Face16abdelstark /sommelier-xlam-single-call-splits-he-hymt-sanitized Sommelier xLAM single-call Hebrew paired rows (Hy-MT2, sanitized release) This CC-BY-4.0 dataset is derived from Salesforce/xlam-function-calling-60k. Sommelier filters the source corpus to single-tool-call examples, deterministically splits it, and machine-translates only each natural-language query into Hebrew. The exact training snapshot kept tool schemas and gold answers byte-identical to the English root. For public release, 15 GitHub-PAT-shaped substrings inherited from… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-he-hymt-sanitized.texttext-generation10K<n<100K1 likes22 downloads2mo agoHugging Face17Zyroxx66 /Somali-Reasoning-Dataset Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴 This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language. 🌟 What makes this unique? This is a Hybrid Dataset that combines two powerful sources: The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.texttext-generation10K<n<100K0 likes21 downloads6mo agoHugging Face18SOMIL366 /4D4T Four Dataset For Training 4D4T: 4-Domain Training Dataset A curated ~60GB corpus split into four balanced domains for training small language models. 📊 Domain Breakdown Domain Source Format Math openbmb/UltraData-Math data/math/math_train_shard_*.jsonl.gz History allenai/c4 (realnewslike) data/history_news/history_train_shard_*.jsonl.gz Science sentence-transformers/s2orc data/science/science_train_shard_*.jsonl.gz General… See the full description on the dataset page: https://huggingface.co/datasets/SOMIL366/4D4T.texttext-generation10M<n<100M0 likes17 downloads5mo agoHugging Face19Zyroxx66 /Somali-Somlish-Instruct-2K-Dataset Somlish-Tech-Instruct-2K This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast. 🌟 Why this exists Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.texttext-generation1K<n<10K0 likes13 downloads6mo agoHugging Face20Fodde /do_some_training Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Fodde/do_some_training.texttext-generationn<1K0 likes4 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.