CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01C10X /datasettext1K<n<10K0 likes619 downloads10mo agoHugging Face02C10X /k3 Kimi K3 Deduplicated Pi Agent Sessions Native Hugging Face Agent Trace projection of greghavens/kimi-k3-coding-and-debugging-traces. Unlike a row-level projection, this export first collapses the source dataset's cumulative next-assistant prefixes. Each output .jsonl file represents one complete source trajectory rather than one intermediate training prefix. Build summary Source revision: 33a874c3affbdb97e142752a9144e6624ef5bd07 Source cumulative rows: 3,956… See the full description on the dataset page: https://huggingface.co/datasets/C10X/k3.tabulartext-generationn<1K0 likes276 downloads2mo agoHugging Face03DenyTranDFW /UBS_Commercial_Mortgage_Trust_2018_C10_1736862 UBS Commercial Mortgage Trust 2018-C10 SEC ABS-EE asset-level filings for CIK 1736862 (UBS Commercial Mortgage Trust 2018-C10). Filings: 75 Parquet files: 296 Total size: 20.6 MB Reporting period start: 2018-05-11 Reporting period end: 2024-07-11 Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/UBS_Commercial_Mortgage_Trust_2018_C10_1736862.tabular10K<n<100K0 likes156 downloads5mo agoHugging Face04C10X /finepdfs-edu-hq FinePDFs-Edu (English) — Filtered High-Signal Subset This dataset is a filtered, English-only subset of HuggingFaceFW/finepdfs-edu, created to retain high-signal educational passages while reducing common PDF-extraction noise (covers/TOCs, fragmented headers/footers, OCR artifacts, mixed-language pages, and very short low-context snippets). It is intended for training and research workflows that benefit from longer, coherent educational text extracted from PDFs. At a… See the full description on the dataset page: https://huggingface.co/datasets/C10X/finepdfs-edu-hq.tabular1M<n<10M0 likes131 downloads8mo agoHugging Face05C10X /ultrafinewebtext1M<n<10M0 likes125 downloads8mo agoHugging Face06zh-tw-llm-dv /zh-tw-pythia-ta8000-v1-e1-tr_wiki_sg-001-c1024 zh-tw-pythia-ta8000-v1-e1-tr_wiki_sg-001-c1024 This dataset is a part of the zh-tw-llm project. Tokenizer: zh-tw-pythia-tokenizer-a8000-v1 Built with: translations, wikipedia, sharegpt Rows: train 305956, test 225 Max length: 1024 Full config:{"build_with": ["translations", "wikipedia", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English:… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_wiki_sg-001-c1024.tabular100K<n<1M1 likes105 downloads3y agoHugging Face07mzio /sc_synthetic_conversations_c10_q3_gpt-41-mini_v4textn<1K0 likes82 downloads6mo agoHugging Face08C10X /datasetstabular100K<n<1M0 likes61 downloads1d agoHugging Face09ragrawal36 /msa-longhealth-c10000-eval-queriestextn<1K0 likes46 downloads1mo agoHugging Face10ragrawal36 /msa-longhealth-c10000-rag-corpus-evalmatched_maxdoc2048_ctx16384textn<1K0 likes46 downloads29d agoHugging Face11ragrawal36 /msa-longhealth-c10000-rag-corpus-evalmatchedtextn<1K0 likes45 downloads1mo agoHugging Face12electricsheepafrica /africa-mauritius-area-harvested-production-yield-and-interline-of-food-crop-c10ad7f4 Area Harvested Production Yield and Interline of Food Crop | Africa (MDPA) 47 rows - 1 Africa country/area - 2021 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 47 rows from MDPA, covering Area Harvested Production Yield and Interline of Food Crop. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-area-harvested-production-yield-and-interline-of-food-crop-c10ad7f4.tabulartabular-classificationn<1K0 likes43 downloads1mo agoHugging Face13C10X /hightabular100K<n<1M0 likes42 downloads11mo agoHugging Face14infinitylogesh /book_dataset_no_mem_token_gte_largev1_5_M512_C1024_1Btext100K<n<1M0 likes40 downloads9mo agoHugging Face15ragrawal36 /msa-longhealth-c10000-eval-queries_maxdoc2048_ctx16384textn<1K0 likes37 downloads29d agoHugging Face16C10X /multiturn-chattext1K<n<10K0 likes36 downloads10d agoHugging Face17ragrawal36 /msa-qasper-c10000-rag-corpus-evalmatchedtext10K<n<100K0 likes34 downloads1mo agoHugging Face18C10X /testv2tabular1K<n<10K0 likes34 downloads5d agoHugging Face19C10X /omni-mathtext1K<n<10K0 likes33 downloads2y agoHugging Face20ragrawal36 /msa-qasper-c10000-eval-queriestextn<1K0 likes33 downloads1mo agoHugging Face21C10X /Finepdf-edutabular1M<n<10M0 likes30 downloads8mo agoHugging Face22C10X /fineutabular100K<n<1M0 likes28 downloads10mo agoHugging Face23zh-tw-llm-dv /zh-tw-pythia-ta8000-v1-e1-tr_sg-201-c1024 zh-tw-pythia-ta8000-v1-e1-tr_sg-201-c1024 This dataset is a part of the zh-tw-llm project. Tokenizer: zh-tw-pythia-tokenizer-a8000-v1 Built with: translations, sharegpt Rows: train 205965, test 195 Max length: 1024 Full config:{"build_with": ["translations", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English: {lang_1}\nChinese: {lang_2}"… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_sg-201-c1024.tabular100K<n<1M0 likes25 downloads3y agoHugging Face24C10X /cn_k12text100K<n<1M0 likes25 downloads11mo agoHugging Face25zh-tw-llm-dv /zh-tw-pythia-ta8000-v1-e1-tr_sg-302-c1024 zh-tw-pythia-ta8000-v1-e1-tr_sg-302-c1024 This dataset is a part of the zh-tw-llm project. Tokenizer: zh-tw-pythia-tokenizer-a8000-v1 Built with: translations, sharegpt Rows: train 305958, test 195 Max length: 1024 Full config:{"build_with": ["translations", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English: {lang_1}\nChinese: {lang_2}"… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_sg-302-c1024.tabular100K<n<1M0 likes21 downloads3y agoHugging Face26C10X /SuperGPQAtext10K<n<100K0 likes21 downloads2y agoHugging Face27Aarifkhan /NCERT-c10-12tabular100K<n<1M0 likes20 downloads8mo agoHugging Face28zh-tw-llm-dv /zh-tw-pythia-ta8000-v1-e1-tr_sg-301-c1024 zh-tw-pythia-ta8000-v1-e1-tr_sg-301-c1024 This dataset is a part of the zh-tw-llm project. Tokenizer: zh-tw-pythia-tokenizer-a8000-v1 Built with: translations, sharegpt Rows: train 306319, test 200 Max length: 1024 Full config:{"build_with": ["translations", "sharegpt"], "preview_length": 128, "translations_settings": {"source_dataset": "zetavg/coct-en-zh-tw-translations-twp-300k", "lang_1_key": "en", "lang_2_key": "ch", "templates": ["English: {lang_1}\nChinese: {lang_2}"… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-e1-tr_sg-301-c1024.text100K<n<1M0 likes19 downloads3y agoHugging Face29C10X /finepdfs-edu-hq-2048tabular1M<n<10M0 likes19 downloads8mo agoHugging Face30ragrawal36 /msa-2wikimultihopqa-c10000-eval-queriestextn<1K0 likes19 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.