datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Global_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.social-register-corpus
Social Register Corpus
Structured person and family records extracted from historical American Social Registers, Blue Books, Elite Directories, Who's Who volumes, and related society directories (1880s–1950s). Built from Internet Archive OCR text with regex parsing and optional Ollama refinement.
Dataset summary
Split
Records
Description
entries
821,683
Parsed person/household listings
families
201,106
Surname + city household groups
editions
219… See the full description on the dataset page: https://huggingface.co/datasets/datamatters24/social-register-corpus.SFT_Dataset_domain_social
Nepali Social Studies MCQ — SFT Dataset
A cleaned, deduplicated, bias-corrected instruction-tuning dataset of Nepali-language
multiple-choice questions on social studies topics, derived from the Aya Dataset.
Dataset Summary
Rows
27,891
Language
Nepali (ne / npi), Devanagari script
Task type
Instruction-following (single-turn MCQ Q&A)
Domain
Social studies (सामाजिक) — MCQ only
License
Apache-2.0 (permissive)
Source
CohereLabs/aya_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/SFT_Dataset_domain_social.s1_dataset_ptbr_1k_tokenized
s1_dataset_ptbr_1k_tokenized
Resumo do Dataset
O s1_dataset_ptbr_1k é uma tradução para o Português (PT-BR) do dataset simplescaling/s1K.
Este conjunto de dados contém 1.000 exemplos de alta qualidade focados em raciocínio lógico, matemático e resolução de problemas. O principal diferencial deste dataset é a inclusão de trajetórias de pensamento (thinking trajectories), que mostram o passo a passo que o modelo deve realizar antes de chegar à resposta final.
Este dataset é… See the full description on the dataset page: https://huggingface.co/datasets/corre-social/s1_dataset_ptbr_1k_tokenized.
