CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mir178 /shangkhachil-bengali-public-domain Bengali Public-Domain Literature 101 complete works by 21 authors, 11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09. Where these texts are read https://shangkhachil.com — the reading site this corpus was built for. Free, no account, 246 works by 28 authors. The complete text of every work in this file can be read there. This file is the text. The site is the part a JSONL cannot be: Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.tabulartext-generationn<1K0 likes71 downloads14d agoHugging Face02JUNGU /repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes64 downloads2mo agoHugging Face03somasekhar-dev /nexttoken-pmkisan-domain-sft-data NextToken pmkisan domain SFT data (v1) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from ~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA, banking products, insurance, savings instruments, etc.). Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.tabularquestion-answering1K<n<10K0 likes53 downloads6d agoHugging Face04proxectonos /corpus_dominio_periodistico Corpus de dominio periodístico Descripción general El corpus periodístico reúne textos informativos procedentes de prensa digital en gallego, recopilados a partir de distintos medios y en el marco de proyectos y fases de adquisición diferentes. El conjunto representa el registro periodístico contemporáneo y está orientado a su uso en tareas de procesamiento del lenguaje natural. El corpus incluye tanto colecciones previamente integradas en CorpusNÓS, con un esquema de… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/corpus_dominio_periodistico.tabulartext-generation100K<n<1M0 likes49 downloads5mo agoHugging Face05RazinAleks /Python_SO_domainstabular10K<n<100K0 likes18 downloads3y agoHugging Face06mdg-nlp /domain-eventx-clinicaltabular10K<n<100K0 likes12 downloads7mo agoHugging Face07tppllm /multi-domain-description Multi-Task Description Dataset This dataset contains multiple event sequences from various sources. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper. If you find this dataset useful, we kindly invite you to cite the following papers: @article{liu2024tppllmm, title={TPP-LLM: Modeling Temporal Point Processes by Efficiently Fine-Tuning Large Language Models}, author={Liu, Zefang and Quan, Yinzhu}… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/multi-domain-description.tabular1K<n<10K1 likes11 downloads10mo agoHugging Face08mdg-nlp /domain-timex-clinicaltabular1K<n<10K0 likes11 downloads7mo agoHugging Face09spectralbranding /exp-cross-domain-primacy Experiment F2: Cross-Domain Primacy (Brand vs Political Attitudes) Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v) Dataset DOI: 10.57967/hf/8456 Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism Dataset Summary 2,400 LLM API calls testing whether serial position primacy generalizes from brand perception to political attitude measurement. Uses two parallel 8-dimension frameworks: Spectral Brand Theory (SBT) for brands and Moral… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-cross-domain-primacy.tabulartext-generation1K<n<10K0 likes9 downloads2mo agoHugging Face10Zenng2812 /bctc-md-domain-corpus Vietnamese Financial Reports Markdown Domain Corpus Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD. Mục đích Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt. Cấu trúc dữ liệu Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object: { "text": "...", "source_file": "AAA_BCTC_2020.md", "document_id": "AAA_BCTC_2020", "company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.imagetext-generation1K<n<10K0 likes9 downloads5mo agoHugging Face11Klane /US_Domestic_Messaging_Pricingtabularn<1K0 likes8 downloads3y agoHugging Face12mdg-nlp /domain-timex-recognition-sentence-updatedtabular10K<n<100K0 likes8 downloads7mo agoHugging Face13mdg-nlp /domain-eventx-clinical-basetabular1K<n<10K0 likes8 downloads7mo agoHugging Face14saksham1771 /infinite-dom-datatabular1K<n<10K0 likes8 downloads5mo agoHugging Face15mdg-nlp /domain-timex-recognition-sentencetabular10K<n<100K0 likes7 downloads7mo agoHugging Face16costadev00 /wikipedia-pt-br-domain wikipedia-pt-br-domain-gemma Versão enriquecida de costadev00/wikipedia-pt-br-extract com um label sintético de domínio por artigo. Processo Cada registro preserva os campos originais esperados da Wikipedia (page_id, title, text, ns, section_texts) e adiciona domain, derivado do label documental primary_category produzido pelo modelo. Modelo Modelo usado para labeling: google/gemma-4-26B-A4B-it. Versão da pipeline: 0.1.0. Limitações O campo domain é… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-domain.tabulartext-classificationn<1K0 likes7 downloads5mo agoHugging Face17reflectio /full_dominotabular1M<n<10M0 likes5 downloads3mo agoHugging Face18mdg-nlp /domain-timex-clinical-basetabular1K<n<10K0 likes4 downloads7mo agoHugging Face19mdg-nlp /domain-eventx-recognition-sentence-updatedtabular10K<n<100K0 likes3 downloads7mo agoHugging Face20mdg-nlp /domain-eventx-recognition-sentencetabular10K<n<100K0 likes1 downloads7mo agoHugging Face21GewHug /Multi-Domain-Evaltabularn<1K1 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.