datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sommelier-xlam-single-call-splits
sommelier xlam single-call splits
Deterministic, deduplicated, single-tool-call train/validation/test
splits derived from
Salesforce/xlam-function-calling-60k
(APIGen, CC-BY-4.0), produced by the
sommelier pipeline for
reproducible tool-calling fine-tuning. These are the exact splits used to
train and evaluate
abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora.
Why single-call
The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.alpaca_ccass_motivations_sommaires_titres
Training dataset for summarizing and titling decisions of the French Court of cassation based on motivations
This alpaca-format dataset is designed to train models for summarizing and titling French Supreme Court decisions based on the grounds of them. Created with a view to producing metadata for decisions not published in the bulletin, this dataset aims to simplify the development of annotation and categorization tools, and is positioned as a facilitator for jurisprudential… See the full description on the dataset page: https://huggingface.co/datasets/Cour-de-cassation/alpaca_ccass_motivations_sommaires_titres.dataset-aeroespacial-cultural-somosnlp
LATAM Aerospace History QA
Descripción General
LATAM Aerospace History QA es un dataset curado orientado a instruction tuning y sistemas conversacionales culturalmente alineados para Iberoamérica.
El dataset se enfoca principalmente en español, incorporando además cobertura parcial en portugués brasileño para mejorar representación multicultural y multilingüe dentro de modelos de lenguaje abiertos.
La colección está especializada en:
historia aeroespacial,
programas… See the full description on the dataset page: https://huggingface.co/datasets/AngelGabrielTroncoso/dataset-aeroespacial-cultural-somosnlp.psychoanalysis-dataset-100k
Psychoanalysis Synthetic Instruction Dataset (v1, 100k)
Domain: psychoanalytic reflection / therapy-style dialoguesLocale: English + Hinglish (India context)Size: 100,000 rows; 10 shards × 10k JSONL
Schema
Chat-style messages + instruction/input/output + safety + metadata.Educational only; not clinical advice.
Split
train only (create validation downstream with train_test_split).
Generation Notes
Synthetic templates + slot-filling; no… See the full description on the dataset page: https://huggingface.co/datasets/SomyaSaraswati/psychoanalysis-dataset-100k.Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.tac-closing-efficiency-sft
TAC closing-efficiency slice
500 synthetic multi-turn tool-use trajectories that teach an agent to close bookings decisively — the welfare-neutral capability piece of the tool-use SFT mix used to train somaxsoma/qwen2.5-7b-tac-recovery-sft.
What it teaches
Built to fix the dominant failure mode observed on the TAC benchmark — the model reformulating search keywords in a loop and never closing a booking. Three patterns:
settle/browse (200): after failed keyword… See the full description on the dataset page: https://huggingface.co/datasets/somaxsoma/tac-closing-efficiency-sft.number_of_death_by_sex_hermes_calling
Nepal Education Enrollment Statistics – Hermes Function-Calling Dataset
1. Overview
This dataset contains 16,760 single-turn function-calling records in Hermes / ShareGPT conversation format. Each record pairs a natural-language request for education enrollment statistics with the exact tool call that satisfies it. Tool-call arguments are grounded in the administrative hierarchy of an Excel source workbook (Annex 4 – Enrollment Details): every province, district… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/number_of_death_by_sex_hermes_calling.some
FinEE Dataset
Dataset Description
A comprehensive dataset for training financial entity extraction models on Indian banking messages. Contains 152,000+ samples covering SMS, emails, and transaction notifications from major Indian banks.
Languages
English (en) - 86%
Hindi (hi) - 3%
Tamil (ta) - 3%
Telugu (te) - 3%
Bengali (bn) - 3%
Kannada (kn) - 2%
Supported Transaction Types
UPI payments (PhonePe, GPay, Paytm)… See the full description on the dataset page: https://huggingface.co/datasets/Siddhu077/some.science_behavioral_and_domain_diversity_dataset
Nepali Science SFT Dataset — Clean Candidate
A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script.
This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering.
Dataset Overview
Property
Value
Dataset file
clean_candidate.jsonl
Records
29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.sommelier-xlam-single-call-splits-fr
sommelier-xlam-single-call-splits-fr
French paired variant of the single call tool calling rows selected by the Sommelier reference pipeline from Salesforce/xlam-function-calling-60k. Only the user query is translated. Tool schemas and gold answers are byte identical to the English source rows, so the two languages measure the same task with the same scoring.
How it was built
The Sommelier data translate tool (source) translated the exact 17,000 rows the reference… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-fr.educational_domain_dataset
Nepali Grounded Education QA (OpenHermes-format)
A small, fact-grounded Nepali instruction-tuning dataset of question–answer pairs about
student enrollment statistics from Nepal's Ministry of Education. Every answer is
anchored to a real numeric value pulled from government open data — nothing in the
answers is model-hallucinated.
Dataset Summary
Rows
611
Language
Nepali (Devanagari script)
Format
ShareGPT / hermes-instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_domain_dataset.somali-web-corpus
SOMALI-WEB-CORPUS V1
This dataset consists of clean, structured, and filtered Somali language text compiled from various online sources. It is designed for training and fine-tuning Somali language models (LLMs) and supporting natural language processing (NLP) research for the Somali language.
Dataset Details
Language: Somali (so)
Format: JSON lines (.jsonl)
Data Structure: Each record has a single text field containing a cleaned paragraph.
Sources… See the full description on the dataset page: https://huggingface.co/datasets/maanka2/somali-web-corpus.cyber_security
Digital Literacy & Cybersecurity Nepali SFT Dataset
Dataset Overview
This dataset is a Nepali-language Supervised Fine-Tuning (SFT) dataset focused on digital literacy and cybersecurity.
The dataset contains 1,000 valid JSONL records designed for instruction-following tasks. Each record contains a human instruction and a corresponding GPT-generated response.
Dataset Statistics
Property
Value
Total records
1,000
Valid JSONL rows
1,000… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/cyber_security.SFT_Dataset_domain_social
Nepali Social Studies MCQ — SFT Dataset
A cleaned, deduplicated, bias-corrected instruction-tuning dataset of Nepali-language
multiple-choice questions on social studies topics, derived from the Aya Dataset.
Dataset Summary
Rows
27,891
Language
Nepali (ne / npi), Devanagari script
Task type
Instruction-following (single-turn MCQ Q&A)
Domain
Social studies (सामाजिक) — MCQ only
License
Apache-2.0 (permissive)
Source
CohereLabs/aya_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/SFT_Dataset_domain_social.somewhereinblog-article
Somewhereinblog Article Archive
Overview
This repository contains a large-scale text dataset scraped from m.somewhereinblog.net, the largest and first-ever Bengali community blogging platform. The primary goal of this archive is to preserve a massive collection of purely human-written blog posts, personal stories, socio-political opinions, and community discussions, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/somewhereinblog-article.sommelier-xlam-single-call-splits-he-hymt-sanitized
Sommelier xLAM single-call Hebrew paired rows (Hy-MT2, sanitized release)
This CC-BY-4.0 dataset is derived from
Salesforce/xlam-function-calling-60k.
Sommelier filters the source corpus to single-tool-call examples, deterministically
splits it, and machine-translates only each natural-language query into Hebrew.
The exact training snapshot kept tool schemas and gold answers byte-identical to
the English root. For public release, 15 GitHub-PAT-shaped substrings inherited
from… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits-he-hymt-sanitized.Somali-Reasoning-Dataset
Somali-OpenHermes-Somlish-Instruct-20K 🇸🇴
This dataset is a gift to the Somali AI community. It is designed to help developers build models that are both highly intelligent and naturally conversational in our language.
🌟 What makes this unique?
This is a Hybrid Dataset that combines two powerful sources:
The Logic (18,379 rows): A Somali translation of the world-class teknium/OpenHermes-2.5. This part provides the AI with deep reasoning, mathematics, coding, and… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Reasoning-Dataset.4D4T
Four Dataset For Training
4D4T: 4-Domain Training Dataset
A curated ~60GB corpus split into four balanced domains for training small language models.
📊 Domain Breakdown
Domain
Source
Format
Math
openbmb/UltraData-Math
data/math/math_train_shard_*.jsonl.gz
History
allenai/c4 (realnewslike)
data/history_news/history_train_shard_*.jsonl.gz
Science
sentence-transformers/s2orc
data/science/science_train_shard_*.jsonl.gz
General… See the full description on the dataset page: https://huggingface.co/datasets/SOMIL366/4D4T.Somali-Somlish-Instruct-2K-Dataset
Somlish-Tech-Instruct-2K
This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast.
🌟 Why this exists
Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.do_some_training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Fodde/do_some_training.
