datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RecetasDeLaAbuela
Motivación inicial
Este corpus ha sido creado durante el Hackathon SomosNLP Marzo 2024: #Somos600M (https://somosnlp.org/hackathon).
Responde a una de las propuestas somosnlp sobre 'Recetas típicas por país/zona geográfica'.
Nombre del Proyecto
Este corpus o dataset se llama 'RecetasDeLaAbuel@' y es un homenaje a todas nuestr@s abuel@s que nos han enseñado a cocinar. Se trata de la mayor y más completa colección de recetas open-source en español de países… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp/RecetasDeLaAbuela.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.sommbench
SommBench
SommBench is a multilingual benchmark for assessing sommelier expertise in large language models. It comprises 3,024 expert-curated examples across eight languages (da, de, en, es, fi, it, sk, sv), designed by professional sommeliers to evaluate sensory grounding, factual wine knowledge, and practical pairing skills through three tasks: WTQA, WFC, and FWP.
Configs
wtqa — Wine Theory Question Answering (1,024 examples)
Multiple-choice questions about… See the full description on the dataset page: https://huggingface.co/datasets/sommify/sommbench.SOMAJGYAAN
SomajGyaan (সমাজজ্ঞান) - Bangla MCQ Dataset
📊 Dataset Description
SomajGyaan (সমাজজ্ঞান) is a comprehensive Bangla multiple-choice question dataset featuring 4,234 questions across 7 academic categories with ~12,000 unique answer options.
Dataset Summary
Total Questions: 4,234
Unique Answer Options: ~12,000
Answer Diversity: 70.8%
Language: Bangla (Bengali)
Categories: 7 (History, Economics, Geography, Politics, Social Studies, Law… See the full description on the dataset page: https://huggingface.co/datasets/farihashifa/SOMAJGYAAN.exam_zh_multitopic_dialect_culture
exam_zh_multitopic_dialect_culture
This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge.
📚 Description
The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories:
🗣️ Regional Dialect Tests
These assess language understanding across major Chinese dialects and topolects:
Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.
