datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datasetk3
Kimi K3 Deduplicated Pi Agent Sessions
Native Hugging Face Agent Trace projection of greghavens/kimi-k3-coding-and-debugging-traces.
Unlike a row-level projection, this export first collapses the source dataset's
cumulative next-assistant prefixes. Each output .jsonl file represents one
complete source trajectory rather than one intermediate training prefix.
Build summary
Source revision: 33a874c3affbdb97e142752a9144e6624ef5bd07
Source cumulative rows: 3,956… See the full description on the dataset page: https://huggingface.co/datasets/C10X/k3.UBS_Commercial_Mortgage_Trust_2018_C10_1736862
UBS Commercial Mortgage Trust 2018-C10
SEC ABS-EE asset-level filings for CIK 1736862 (UBS Commercial Mortgage Trust 2018-C10).
Filings: 75
Parquet files: 296
Total size: 20.6 MB
Reporting period start: 2018-05-11
Reporting period end: 2024-07-11
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/UBS_Commercial_Mortgage_Trust_2018_C10_1736862.finepdfs-edu-hq
FinePDFs-Edu (English) — Filtered High-Signal Subset
This dataset is a filtered, English-only subset of HuggingFaceFW/finepdfs-edu, created to retain high-signal educational passages while reducing common PDF-extraction noise (covers/TOCs, fragmented headers/footers, OCR artifacts, mixed-language pages, and very short low-context snippets).
It is intended for training and research workflows that benefit from longer, coherent educational text extracted from PDFs.
At a… See the full description on the dataset page: https://huggingface.co/datasets/C10X/finepdfs-edu-hq.ultrafinewebzh-tw-pythia-ta8000-v1-e1-tr_wiki_sg-001-c1024
zh-tw-pythia-ta8000-v1-e1-tr_wiki_sg-001-c1024
This dataset is a part of the zh-tw-llm project.
Tokenizer: zh-tw-pythia-tokenizer-a8000-v1
Built with: translations, wikipedia, sharegpt
Rows: train 305956, test 225
Max length: 1024
Full config:{"build_with": ["translations", "wikipedia", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English:… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_wiki_sg-001-c1024.sc_synthetic_conversations_c10_q3_gpt-41-mini_v4datasetsmsa-longhealth-c10000-eval-queriesmsa-longhealth-c10000-rag-corpus-evalmatched_maxdoc2048_ctx16384msa-longhealth-c10000-rag-corpus-evalmatchedafrica-mauritius-area-harvested-production-yield-and-interline-of-food-crop-c10ad7f4
Area Harvested Production Yield and Interline of Food Crop | Africa (MDPA)
47 rows - 1 Africa country/area - 2021 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 47 rows from MDPA, covering Area Harvested Production Yield and Interline of Food Crop. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-area-harvested-production-yield-and-interline-of-food-crop-c10ad7f4.highbook_dataset_no_mem_token_gte_largev1_5_M512_C1024_1Bmsa-longhealth-c10000-eval-queries_maxdoc2048_ctx16384multiturn-chatmsa-qasper-c10000-rag-corpus-evalmatchedtestv2omni-mathmsa-qasper-c10000-eval-queriesFinepdf-edufineuzh-tw-pythia-ta8000-v1-e1-tr_sg-201-c1024
zh-tw-pythia-ta8000-v1-e1-tr_sg-201-c1024
This dataset is a part of the zh-tw-llm project.
Tokenizer: zh-tw-pythia-tokenizer-a8000-v1
Built with: translations, sharegpt
Rows: train 205965, test 195
Max length: 1024
Full config:{"build_with": ["translations", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English: {lang_1}\nChinese: {lang_2}"… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_sg-201-c1024.cn_k12zh-tw-pythia-ta8000-v1-e1-tr_sg-302-c1024
zh-tw-pythia-ta8000-v1-e1-tr_sg-302-c1024
This dataset is a part of the zh-tw-llm project.
Tokenizer: zh-tw-pythia-tokenizer-a8000-v1
Built with: translations, sharegpt
Rows: train 305958, test 195
Max length: 1024
Full config:{"build_with": ["translations", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English: {lang_1}\nChinese: {lang_2}"… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_sg-302-c1024.SuperGPQANCERT-c10-12zh-tw-pythia-ta8000-v1-e1-tr_sg-301-c1024
zh-tw-pythia-ta8000-v1-e1-tr_sg-301-c1024
This dataset is a part of the zh-tw-llm project.
Tokenizer: zh-tw-pythia-tokenizer-a8000-v1
Built with: translations, sharegpt
Rows: train 306319, test 200
Max length: 1024
Full config:{"build_with": ["translations", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English: {lang_1}\nChinese: {lang_2}"… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_sg-301-c1024.finepdfs-edu-hq-2048msa-2wikimultihopqa-c10000-eval-queries
