datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pairs_three_scores_v13_synonyms_addedMedQA-USMLE-combined-synonym-firstMedQA-USMLE-synonym-replacementMedQA dataset perturbed using knowledge-based synonym replacement technique with BAT
pairs_three_scores_v14_tag_sim_synonyms_addedpairs_three_scores_v10_synonymsefik-synonymsspai-ss6-corpus-thai-synonym-instruction
SPAI SS6 Thai Synonym Instruction Index
Index repo for the Thai synonym instruction dataset mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: thai_synonym_instruction
Rows in canonical config: 167
Parquet size in canonical config: 0.00 GB
Source… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-thai-synonym-instruction.cim_synonymes
