datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aiwolf-nlp-agent-llm
AIWolfDial 2026 Power Play Evaluation
Public data release: 2026-09-14. This dataset is available at synonym/aiwolf-nlp-agent-llm, with the snapshot tag release-20260914. The matching code distribution is 1.0.0-rc.3, commit d427dc299bacf4eb4cb41c114c8af476b71ac7ed. The code distribution uses a single root commit; this dataset is separate and is not included in that repository. Paper publication identifiers are still pending. The dataset is distributed under the MIT license in… See the full description on the dataset page: https://huggingface.co/datasets/synonym/aiwolf-nlp-agent-llm.Hindi-Marathi-Synonyms
Multilingual Synonyms Dataset (बहुभाषी पर्यायवाची शब्द संग्रह)
Overview
This dataset contains a comprehensive collection of words and their synonyms across multiple Indian languages including Hindi and Marathi. It is designed to assist NLP research, language learning, and applications focused on Indian language processing and cross-lingual applications.
The dataset provides word-synonym pairs that can be used for tasks like:
Semantic analysis
Language learning and… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Hindi-Marathi-Synonyms.thai-synonym-instruction
Dataset Card for Thai synonym instruction
The dataset is used https://github.com/PyThaiNLP/thai-synonym to creating the dataset.
Support Me
GitHub Sponsors:
If you can't help financially, don't worry! You can say Thanks!
spai-ss6-corpus-thai-synonym-instruction
SPAI SS6 Thai Synonym Instruction Index
Index repo for the Thai synonym instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_synonym_instruction
Rows in canonical config: 167
Parquet size in canonical config: 0.00 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synonym-instruction.
