CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /tabula-8b-eval-suiteEvaluation suite used in our paper "Large Scale Transfer Learning for Tabular Data via Language Modeling." This suite includes our preprocessed versions of benchmark datasets except the AutoML Multimodal Benchmark, which can be accessed by following the installation instructions in their repo here. We recommend using rtfm when evaluating models with these datasets. See the rtfm repo for more information on using this data for evaluation. texttabular-classification10K<n<100K6 likes762 downloads2y agoHugging Face02turkish-nlp-suite /TrGLUE TrGLUE - The First Non-Translate Natural Language Understanding Benchmark for Turkish Dataset Card for TrGLUE TrGLUE is a natural language understanding benchmarking dataset including several single sentence and sentence pair classification tasks. The inspiration is clearly the original GLUE benchmark. Tasks Single Sentence Tasks TrCOLA The original Corpus of Linguistic Acceptability consists of sentences compiled from English literature textbooks.… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TrGLUE.texttext-classification100K<n<1M6 likes588 downloads9mo agoHugging Face03kisate-team /gemma-2b-suite-explanations-residualtext100K<n<1M0 likes478 downloads2y agoHugging Face04turkish-nlp-suite /temiz-OSCAR Dataset Card for Temiz OSCAR Temiz OSCAR is a corpora collection consisting of cleaned versions of original OSCAR corpora. This collection is made up of four datasets: OSCAR-2019, OSCAR-2109, OSCAR-2201 and OSCAR-2301 This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. Dataset num instances size num of words OSCAR-2019 3.671.430 7.7G 976M OSCAR-2109 8.472.809 18G 2.22B OSCAR-2201… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/temiz-OSCAR.textfill-mask10M<n<100M5 likes475 downloads11mo agoHugging Face05ShigeoKageyama /NLP_SUITEtext10M<n<100M0 likes387 downloads24d agoHugging Face06kisate-team /gemma-2b-suite-maxacts-attn_out text100K<n<1M0 likes304 downloads2y agoHugging Face07kisate-team /gemma-2b-suite-maxacts-residual text100K<n<1M0 likes261 downloads2y agoHugging Face08turkish-nlp-suite /Havadis Dataset Card for Havadis Havadis is a high quality and large Turkish news corpus, indeed the largest Turkish news corpus ever. This corpus is scraped from online news sebsites and includes text from popular newspapers such as CNN Türk Habertürk Hürriyet Millyet NTV Posta Sabah Star Sözcü Takvim . The instances are first crawled from the corresponding websites, then went throught an extensive cleaning process. We eliminated instances that are too short, too repetetive… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/Havadis.textfill-mask100K<n<1M6 likes236 downloads3mo agoHugging Face09suitai /salabs-robotics-spatial-topology-v9 🤖 SALabs 768-D Continuous Lie SE(3) Robotics & Spatial Manifold Topology Dataset (v9.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($1,500 USD) & Instant 391.56MB Master DownloadInstant download of the full 391.56MB Enterprise JSONL matrix containing 50,000+ continuous Lie $SE(3)$ manifold trajectories, singularity-free Bishop Frame metrics, and commercial license certificate. 🌟 Executive Summary The SALabs Robotics & Spatial… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-robotics-spatial-topology-v9.tabularrobotics1K<n<10K1 likes227 downloads20d agoHugging Face10turkish-nlp-suite /TurkishHateMap Turkish Hate Map - A Large Scale and Diverse Hate Speech Dataset for Turkish Dataset Summary Turkish Hate Map (TuHaMa for short) is a big scale Turkish hate speech dataset that includes diverse target groups such as misogyny, political animosity, animal aversion, vegan antipathy, ethnic group hostility, and more. The dataset includes a total of 52K instances with 13 target groups. The dataset includes 4 labels, offensive, hate, neutral and civilized. Here is the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/TurkishHateMap.texttext-classification10K<n<100K4 likes224 downloads2y agoHugging Face11turkish-nlp-suite /ForumSohbetleri Dataset Card for ForumSohbetleri ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.textfill-mask1M<n<10M5 likes220 downloads11mo agoHugging Face12suitai /salabs-stem-deep-reasoning-cot-v13 🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate. 🌟 Executive Summary The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.texttext-generation1K<n<10K1 likes220 downloads20d agoHugging Face13turkish-nlp-suite /AkademikDerlem Dataset Card for AkademikDerlem AkademikDerlem is a scientific text corpus for Turkish, gathered from misc academical publication websites. This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication. This collection is made up of five datasets: Articles, Academic-Abstracts, Medical-Articles, Medical-Abstracts, and Bilkent-Writings. The Bilkent-Writings dataset comes from creative writings produced in the… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/AkademikDerlem.textfill-mask100K<n<1M6 likes214 downloads11mo agoHugging Face14turkish-nlp-suite /InstrucTurca InstrucTurca v1.0.0 is a diverse synthetic instruction tuning dataset crafted for instruction-tuning Turkish LLMs. The data is compiled data various English datasets and sources, such as code instructions, poems, summarized texts, medical texts, and more. Dataset content BI55/MedText checkai/instruction-poems garage-bAInd/Open-Platypus Locutusque/ColumnedChatCombined nampdn-ai/tiny-codes Open-Orca/OpenOrca pubmed_qa TIGER-Lab/MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/InstrucTurca.texttext-generation1M<n<10M40 likes198 downloads2y agoHugging Face15babytreecc /Implicit-suicide-detection Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/babytreecc/Implicit-suicide-detection.texttext-classification1K<n<10K1 likes163 downloads1y agoHugging Face16kisate-team /gemma-2b-suite-maxacts-transcoder text100K<n<1M0 likes158 downloads2y agoHugging Face17suitai /salabs-virtual-spatial-digitaltwin-v8 🌐 SALabs 10,000,000-Node 3D Virtual Spatial & Digital Twin Avatar Kinematics Dataset (v8.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($2,000 USD) & Instant 8.0GB Master DownloadInstant download of the complete 8.0GB master archive containing 10,000,000 verified 3D spatial nodes, 18-DoF avatar kinematics, B-spline 4D motion tensors, Laplace-Beltrami spectral resonance, and commercial license certificate. 🌟 Executive Summary The… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-virtual-spatial-digitaltwin-v8.tabularother1K<n<10K1 likes139 downloads20d agoHugging Face18kisate-team /gemma-2b-suite-explanations-attn_out text100K<n<1M0 likes136 downloads2y agoHugging Face19cudabenchmarktest /r8-eval-suite-5bucket ⚠️ CRITICAL: Ollama Inference Flag Required for derived models If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama, you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use. The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag. See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned. R8/R9 Five-Bucket… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-eval-suite-5bucket.tabulartext-generationn<1K0 likes125 downloads6mo agoHugging Face20NickIBrody /rust-code-suite NickIBrody/rust-code-suite Rust Code Suite is a public raw Rust source corpus built from open-source repositories and selected historical git revisions. Splits train.jsonl validation.jsonl test.jsonl Schema { "id": "owner/repo:path:chunk", "text": "...", "arch": "rust", "syntax": "rust", "kind": "rust-source", "repo": "owner/repo", "path": "src/lib.rs", "license": "GPL-2.0", "commit": "abcdef123456", "source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/rust-code-suite.tabulartext-generation1M<n<10M1 likes120 downloads4mo agoHugging Face21AdamLeung /twitter_suicidal_risk Twitter Suicide Risk Level Dataset Short English tweets paired with a 0–4 suicide risk label, used for fine-tuning and evaluating risk-level classification. This directory holds the final splits: train.jsonl / val.jsonl / test.jsonl. Files and size File Rows Share train.jsonl 7006 80% val.jsonl 875 10% test.jsonl 875 10% Total 8756 100% Fields JSONL, one sample per line, three fields only: Field Type Description id… See the full description on the dataset page: https://huggingface.co/datasets/AdamLeung/twitter_suicidal_risk.texttext-classification1K<n<10K0 likes98 downloads1mo agoHugging Face22turkish-nlp-suite /SentiTurca SentiTurca - A Sentiment Analysis Benchmark for Turkish Dataset Card for SentiTurca SentiTurca is a sentiment analysis benchmarking dataset including movie reviews, hate speech and e-commerce reviews classification. Datasets e-commerce: The e-commerce reviews are scraped from e-commerce websites Trendyol.com and Hepsiburada.com, including review for many product types such as cloths, toys, books, electronics and more.E-commerce reviews has their stand… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/SentiTurca.texttext-classification100K<n<1M7 likes96 downloads5mo agoHugging Face23NickIBrody /assembly-code-suite NickIBrody/assembly-code-suite Assembly Code Suite is a public raw assembly corpus built from open-source repositories and selected historical git revisions. Splits train.jsonl validation.jsonl test.jsonl Schema { "id": "owner/repo:path:chunk", "text": "...", "arch": "x86_64", "syntax": "gas-att", "kind": "handwritten", "repo": "owner/repo", "path": "arch/x86/lib/memcpy_64.S", "license": "GPL-2.0", "commit": "abcdef123456", "source_url":… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/assembly-code-suite.tabulartext-generation100K<n<1M1 likes87 downloads4mo agoHugging Face24kisate-team /gemma-2b-suite-explanations-transcodertext100K<n<1M0 likes77 downloads2y agoHugging Face25turkish-nlp-suite /turkish-morph-analysis Dataset Card for TrMorphTester This dataset is a testing dataset for Turkish morphology, aiming to calculate how other subword strategies aligns with morphological segmentation of Turkish. The data is automatically generated from Turkish morpoholigical lexicon. For each row, we offer a surface form, then lemma and all suffixes, separated by a + character. The dataset has several splits for several purposes: lemma: Surface form is same with lemma, no suffixes at all. Nouns, verbs… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/turkish-morph-analysis.text100K<n<1M2 likes76 downloads7mo agoHugging Face26turkish-nlp-suite /vitamins-supplements-reviews Dataset Card for turkish-nlp-suite/vitamins-supplements-reviews Dataset Summary Turkish sentiment analysis dataset from customer reviews about supplement and vitamin products. The dataset is scraped from Vitaminler.com and contains customer reviews and star rating about vitamin and supplement products. Each customer review in the Vitamins and Supplements Reviews Dataset describes a customer’s experience with a supplement product in terms of the product’s effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/vitamins-supplements-reviews.texttext-classification100K<n<1M1 likes70 downloads2y agoHugging Face27turkish-nlp-suite /temiz-WikiA cleaned version of Turkish Wikipedia dataset. The soource is wikimedia/wikipedia repo. The text is cleaned throughoutly, first of all we eliminated text that are shorter than a predetermined threshold of words and characters. Then we normalized with NFKC, cleaned some non-ASCII chars and normalized whitespaces. textfill-mask100K<n<1M4 likes65 downloads7mo agoHugging Face28turkish-nlp-suite /beyazperde-top-300-movie-reviews Dataset Card for turkish-nlp-suite/beyazperde-top-300-movie-reviews Dataset Summary Beyazperde Movie Reviews offers Turkish sentiment analysis datasets that is scraped from popular movie reviews website Beyazperde.com. Top 300 Movies include audience reviews about best 300 movies of all the time. Here's the star rating distribution: star rating count 0.5 101 1.0 39 1.5 19 2.0 44 2.5 210 3.0 196 3.5 490 4.0 1212 4.5 818 5.0 1251 total 4380… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/beyazperde-top-300-movie-reviews.texttext-classification1K<n<10K2 likes58 downloads2y agoHugging Face29icrl-finetuning /2026-06-04-stwebagentbench-suitecrm-demos 2026-06-04-stwebagentbench-suitecrm-demos Standing demo pool for Adversarial Inverse Constraint RL (ICRL) for LLM orchestrator safety on ST-WebAgentBench (SuiteCRM easy tier). Every experiment run consumes this pool; per-run artifacts (embeddings, constraint heads, adapters, CuP evals) live in separate <date>-<run-name> repos in this namespace. field value experiment ICRL safe/unsafe demo pool: constraint C_theta is learned from the safe demos only; unsafe demos are… See the full description on the dataset page: https://huggingface.co/datasets/icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos.tabularn<1K0 likes55 downloads2mo agoHugging Face30NickIBrody /coffeescript-code-suite CoffeeScript Code Suite CoffeeScript Code Suite is a public code dataset built from permissively licensed open-source CoffeeScript repositories. It is designed for three practical uses: CoffeeScript domain adaptation and continued pretraining through raw_corpus examples. CoffeeScript completion training through completion examples. CoffeeScript and JavaScript translation training through coffee_to_js and js_to_coffee examples. The dataset was assembled automatically from public… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/coffeescript-code-suite.tabulartext-generation10K<n<100K0 likes49 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.