datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shangkhachil-bengali-public-domain
Bengali Public-Domain Literature
101 complete works by 21 authors,
11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09.
Where these texts are read
https://shangkhachil.com — the reading site this corpus was built for. Free, no
account, 246 works by 28 authors. The complete text of
every work in this file can be read there.
This file is the text. The site is the part a JSONL cannot be:
Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.repro-time-series-saliency-maps-explaining-models-across-multiple-domains-traces
Agent traces
Agent sessions published from a Trackio Logbook.
nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.corpus_dominio_periodistico
Corpus de dominio periodístico
Descripción general
El corpus periodístico reúne textos informativos procedentes de prensa digital en gallego, recopilados a partir de distintos medios y en el marco de proyectos y fases de adquisición diferentes. El conjunto representa el registro periodístico contemporáneo y está orientado a su uso en tareas de procesamiento del lenguaje natural.
El corpus incluye tanto colecciones previamente integradas en CorpusNÓS, con un esquema de… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/corpus_dominio_periodistico.Python_SO_domainsdomain-eventx-clinicalmulti-domain-description
Multi-Task Description Dataset
This dataset contains multiple event sequences from various sources. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
If you find this dataset useful, we kindly invite you to cite the following papers:
@article{liu2024tppllmm,
title={TPP-LLM: Modeling Temporal Point Processes by Efficiently Fine-Tuning Large Language Models},
author={Liu, Zefang and Quan, Yinzhu}… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/multi-domain-description.domain-timex-clinicalexp-cross-domain-primacy
Experiment F2: Cross-Domain Primacy (Brand vs Political Attitudes)
Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v)
Dataset DOI: 10.57967/hf/8456
Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism
Dataset Summary
2,400 LLM API calls testing whether serial position primacy generalizes from brand perception to political attitude measurement. Uses two parallel 8-dimension frameworks: Spectral Brand Theory (SBT) for brands and Moral… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-cross-domain-primacy.bctc-md-domain-corpus
Vietnamese Financial Reports Markdown Domain Corpus
Dataset này được tạo từ các báo cáo tài chính dạng Markdown trong thư mục BCTC_MD.
Mục đích
Dataset dùng cho continued pretraining / domain-adaptive pretraining mô hình ngôn ngữ trên miền báo cáo tài chính tiếng Việt.
Cấu trúc dữ liệu
Mỗi dòng trong train.jsonl hoặc validation.jsonl là một JSON object:
{
"text": "...",
"source_file": "AAA_BCTC_2020.md",
"document_id": "AAA_BCTC_2020",
"company": "AAA"… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/bctc-md-domain-corpus.US_Domestic_Messaging_Pricingdomain-timex-recognition-sentence-updateddomain-eventx-clinical-baseinfinite-dom-datadomain-timex-recognition-sentencewikipedia-pt-br-domain
wikipedia-pt-br-domain-gemma
Versão enriquecida de costadev00/wikipedia-pt-br-extract com um label sintético de domínio por artigo.
Processo
Cada registro preserva os campos originais esperados da Wikipedia (page_id, title, text, ns, section_texts) e adiciona domain, derivado do label documental primary_category produzido pelo modelo.
Modelo
Modelo usado para labeling: google/gemma-4-26B-A4B-it.
Versão da pipeline: 0.1.0.
Limitações
O campo domain é… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-domain.full_dominodomain-timex-clinical-basedomain-eventx-recognition-sentence-updateddomain-eventx-recognition-sentenceMulti-Domain-Eval
