CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Minuri /sinhala-corpus-c-diverse-1m Diversity-Optimized Sinhala Corpus A diversity-optimized subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model C) as part of a diversity-driven Sinhala language model adaptation study. Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline)… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-c-diverse-1m.tabulartext-generation1M<n<10M0 likes53 downloads6mo agoHugging Face02Minuri /sinhala-test-set-50k Sinhala Test Set - 50K Sentences A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.tabulartext-generation10K<n<100K0 likes36 downloads6mo agoHugging Face03Minuri /sinhala-corpus-b-random-1m Randomly Curated Sinhala Corpus A randomly sampled subset of 1M Sinhala sentences from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model B) as part of a diversity-driven Sinhala language model adaptation study. Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) Minuri/sinhala-corpus-b-random-1m - Random subset (random baseline) - this repo Minuri/sinhala-corpus-c-diverse-1m… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-b-random-1m.tabulartext-generation1M<n<10M0 likes32 downloads6mo agoHugging Face04Minuri /sinhala-corpus-a-news-1m News-Only Sinhala Corpus A news-domain subset of 1M Sinhala sentences sampled from the Minuri/diverse_sinhala_dataset corpus, used for continual pretraining of LLaMA 3.2 1B (Model A) as part of a diversity-driven Sinhala language model adaptation study at the Informatics Institute of Technology (IIT), Colombo, affiliated with Robert Gordon University (RGU). Corpus variants in this series: Minuri/sinhala-corpus-a-news-1m - News-only subset (domain-homogeneous baseline) - this repo… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-corpus-a-news-1m.tabulartext-generation1M<n<10M0 likes27 downloads6mo agoHugging Face05janani-rane /Sinhala-News-Wiki-text-corpus Sinhala-News-Wiki-Text-Corpus Containing news articles from various Sinhala news sites along with Sinhala Wikipedia pages. Dataset Overview Language: Sinhala (සිංහල) Content: Sinhala news articles from various sites Data format: Parquet Number of Records: 18,201 rows (as per current size) Dataset Structure Each record consists of the following fields: category: The news category (e.g., "Other-news, Local-news, wiki, International-news"). site: The site's… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/Sinhala-News-Wiki-text-corpus.tabulartext-classification10K<n<100K0 likes23 downloads2y agoHugging Face06manthilaffs /sinhala-poems-v1 Sinhala Poems (Filtered) Curated Sinhala poem blocks extracted from web blogs using a verse-shaped heuristic (v3.4) with weak attributes (theme/mood/style) and stats. Columns text: full original block (cleaned) snippet: first stanza or 8 lines url: source URL subtype: poem | song_like | promo_like | unknown keep: boolean accepted by filter theme, mood, style, length_class: weak labels ps_*: structure stats (floats) Filtering summary Boilerplate/HTML… See the full description on the dataset page: https://huggingface.co/datasets/manthilaffs/sinhala-poems-v1.tabulartext-generationn<1K0 likes15 downloads1y agoHugging Face07sayururehan /sinhala-personas-lk-v0.2-gemini-1000-preview Sinhala-Personas-LK v0.2 Gemini 1000 Preview Sinhala-Personas-LK is a preview dataset of fully synthetic Sinhala persona records for Sri Lankan NLP research and evaluation. Version 0.1-preview Records 1000 synthetic records. Language Sinhala (si). Country context Sri Lanka (LK). Important limitations This preview version is generated from starter priors and LLM-generated text. It is not yet fully grounded… See the full description on the dataset page: https://huggingface.co/datasets/sayururehan/sinhala-personas-lk-v0.2-gemini-1000-preview.tabulartext-generation1K<n<10K0 likes13 downloads4mo agoHugging Face08Minuri /sinhala-validation-set-10k Sinhala Validation Set - 10K Sentences A held-out Sinhala validation set of 10,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for monitoring validation loss during continual pretraining of three LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This validation set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased validation loss… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-validation-set-10k.tabulartext-generation10K<n<100K0 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.