CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01soma114 /soma-competition-datasettextn<1K0 likes4.5k downloads10h agoHugging Face02somasekhar-dev /nexttoken-pmkisan-domain-sft-data NextToken pmkisan domain SFT data (v1) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from ~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA, banking products, insurance, savings instruments, etc.). Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.tabularquestion-answering1K<n<10K0 likes55 downloads8d agoHugging Face03Abdullahicoder /SomaliCrowS SomaliCrowS: A Gender Bias Benchmark for Somali Language Models Dataset Description SomaliCrowS is a benchmark for measuring gender bias in Somali language models. It contains matched sentence pairs — identical except for the grammatical gender of the subject — spanning social domains where stereotyping commonly occurs, including: Occupation Leadership Business Education STEM Family Politics For each pair, a masked-language-model is queried to compute the… See the full description on the dataset page: https://huggingface.co/datasets/Abdullahicoder/SomaliCrowS.tabulartext-classificationn<1K0 likes42 downloads3mo agoHugging Face04somaxsoma /tac-closing-efficiency-sft TAC closing-efficiency slice 500 synthetic multi-turn tool-use trajectories that teach an agent to close bookings decisively — the welfare-neutral capability piece of the tool-use SFT mix used to train somaxsoma/qwen2.5-7b-tac-recovery-sft. What it teaches Built to fix the dominant failure mode observed on the TAC benchmark — the model reformulating search keywords in a loop and never closing a booking. Three patterns: settle/browse (200): after failed keyword… See the full description on the dataset page: https://huggingface.co/datasets/somaxsoma/tac-closing-efficiency-sft.texttext-generationn<1K0 likes42 downloads28d agoHugging Face05Zyroxx66 /somali-tinystoriestext10K<n<100K0 likes35 downloads1mo agoHugging Face06somayaeltanbouly /Badr_fiqh_retrieval_tripletsgated Dataset Card for Badr Fiqh Retrieval Triplets Dataset Description Badr Fiqh Retrieval Triplets is an Arabic dataset developed for fine-tuning dense retrieval and sentence-embedding models for Islamic jurisprudence (fiqh). Each sample contains a question–positive–negative training triplet together with the title and primary school of the source book. The hard negatives are deliberately selected from closely related fiqh contexts. A model must therefore distinguish… See the full description on the dataset page: https://huggingface.co/datasets/somayaeltanbouly/Badr_fiqh_retrieval_triplets.textsentence-similarity10K<n<100K0 likes29 downloads11d agoHugging Face07burtugeey /Somali_datasettext10K<n<100K1 likes28 downloads2y agoHugging Face08maanka2 /somali-web-corpus SOMALI-WEB-CORPUS V1 This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language. Dataset Details Language: Somali (so) Format: JSON lines (.jsonl) Data Structure: Each record has a single text field containing a cleaned paragraph. Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.texttext-generation100K<n<1M1 likes27 downloads4mo agoHugging Face09SomaliDatasets /maay-maxaa-translation Maay ↔ Maxaa Somali Parallel Translation Corpus Official parallel dataset for Maay (Maay Maay) and Maxaa-tiri (Standard Somali) dialects, collected and curated by the community via MaayMaxaa DataHub. 📊 Dataset Statistics Metric Value Total Sentences 8 Maay → Maxaa 4 Maxaa → Maay 4 Avg Source Length 16 chars Avg Target Length 16 chars Version 0.1.0 🌐 Languages & Dialects Maay (Maay Maay): Southern Somali language/dialect… See the full description on the dataset page: https://huggingface.co/datasets/SomaliDatasets/maay-maxaa-translation.texttranslationn<1K0 likes25 downloads1mo agoHugging Face10Zyroxx66 /Somali-Reasoning-Dataset Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴 This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language. 🌟 What makes this unique? This is a Hybrid Dataset that combines two powerful sources: The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.texttext-generation10K<n<100K0 likes21 downloads6mo agoHugging Face11burtugeey /Somali_alpaca-# Somali_Alpaca Sharaxaad Somali Xog-ururintan "Somali_Alpaca" waxay ka kooban tahay xog ballaaran oo loogu talagalay habaynta luuqadda dabiiciga ah iyo hawlaha barashada mashiinka. Waxaa ku jirto xog been abuur ah, qayb ka mid ah Alpaca_dataset_52k, iyo qayb ka mid ah xogta QA ee laga soo ururiyey internetka. Xogtan waxaa loogu talagalay inay taageerto falanqaynta kala duwan iyo tababarka modelka goobaha waxbarashada iyo kuwa codsiga. English The… See the full description on the dataset page: https://huggingface.co/datasets/burtugeey/Somali_alpaca.text10K<n<100K4 likes19 downloads2y agoHugging Face12burtugeey /AlpacaCleaned_translated_somalitext10K<n<100K1 likes18 downloads2y agoHugging Face13burtugeey /alpaca_somali_finaltext10K<n<100K0 likes14 downloads2y agoHugging Face14Zyroxx66 /Somali-Somlish-Instruct-2K-Dataset Somlish-Tech-Instruct-2K This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast. 🌟 Why this exists Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.texttext-generation1K<n<10K0 likes13 downloads6mo agoHugging Face15burtugeey /Alpaca_Somalitextn<1K0 likes11 downloads2y agoHugging Face16burtugeey /Somali_Alpaca_Google_translatortext10K<n<100K0 likes9 downloads2y agoHugging Face17gowthamreddysomala /somalasdataset{ "background": { "LLM_name": "Somala", "inventor": "Gowtham Reddy Somala", "fine_tuner": "Gowtham Reddy Somala", "contributors": ["Pavan Kalyan", "Mojesh Reddy"] } }, { "instruction": "Who is the Inventor of this Model", "input": "", "output":"Gowtham Reddy Somala is the inventor of this model with a small contribution from his friends Mojesh and Pavan" }, { "instruction": "Who finetuned this Model", "input": "", "output":"Gowtham Reddy Somala… See the full description on the dataset page: https://huggingface.co/datasets/gowthamreddysomala/somalasdataset.textn<1K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.