CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01beshiribrahim /tigre-hubert-dataaudio1K<n<10K0 likes625 downloads2mo agoHugging Face02nltk-data-hub /stopwords NLTK Stopwords Stopword lists from NLTK, covering 33 languages. Each language is a separate config. Each row is one stopword. Usage from datasets import load_dataset # Load one language ds = load_dataset("nltk-data-hub/stopwords", "portuguese") words = ds["stopwords"]["word"] # Load all languages for lang in ['albanian', 'arabic', 'azerbaijani', 'basque', 'belarusian', 'bengali', 'catalan', 'chinese', 'danish', 'dutch', 'english', 'finnish', 'french', 'german', 'greek'… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/stopwords.texttext-classification10K<n<100K0 likes430 downloads5mo agoHugging Face03applied-ai-018 /peacock-data-public-datasets-hubtext100K<n<1M0 likes137 downloads2y agoHugging Face04nltk-data-hub /words NLTK Word Lists English word lists from NLTK, the New General Service List Project, and Bing Liu's Opinion Lexicon. Configs Config Words Schema License Source en 235,886 word NLTK (other) NLTK words corpus en-basic 850 word Public domain Ogden Basic English (1930) ngsl 2,809 word, rank, sfi, freq_per_million CC-BY-SA 4.0 New General Service List 1.2 toeic 1,250 word, rank, sfi, freq_per_million CC-BY-SA 4.0 TOEIC Service List 1.2 nawl 963 word, rank… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/words.tabulartext-classification100K<n<1M1 likes118 downloads5mo agoHugging Face05nltk-data-hub /crubadan NLTK Crúbadán Language ID Corpus Character 3-gram frequency tables for 449 writing systems, collected by Kevin Scannell's An Crúbadán web crawler (2010). Distributed via NLTK. Trigrams use < (word start) and > (word end) as boundary markers. Configs Config Description Schema table Language metadata crubadan_code, iso639_3, language_name {lang_code} Per-language trigrams count, trigram All 449 language codes: ab, abn, ace, ach, acu, ada, af, agr, aja, ak… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/crubadan.texttext-classification1M<n<10M0 likes86 downloads5mo agoHugging Face06altannavchnyamsambuu-hub /Project1-data Dataset Card for SQuAD Dataset Summary Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable. SQuAD 1.1 contains 100,000+ question-answer pairs on 500+ articles. Supported Tasks and Leaderboards Question… See the full description on the dataset page: https://huggingface.co/datasets/altannavchnyamsambuu-hub/Project1-data.textquestion-answering10K<n<100K0 likes85 downloads1mo agoHugging Face07Chijioke-Mgbahurike /hub_large_ls960-ft_spot_data_allaudio1K<n<10K0 likes81 downloads2y agoHugging Face08cBioPortal /datahubWARNING ⚠️: Accessing datahub data via hugging face is still under construction (use https://github.com/cbioPortal/datahub instead). cBioPortal Datahub These are parquet files generated from https://github.com/cBioPortal/datahub. It combines all studies' sample, patient, and mutation data into parquet files. They were generated using cbiohubpy. How to use Query all mutations of datahub directly in duckdb: SELECT count(distinct (Chromosome, Start_Position… See the full description on the dataset page: https://huggingface.co/datasets/cBioPortal/datahub.text10M<n<100M1 likes68 downloads2y agoHugging Face09Bhavya-123 /HRMS_Datahub Human Resource Management System (HRMS) Dataset Overview The Human Resource Management System (HRMS) Dataset is a collection of questions and answers related to various aspects of HRMS. It has been generated for educational purposes and is intended to be used for training and testing question-answering models in the field of Human Resources and Management. License This dataset is provided under the Apache License, Version 2.0. You are free to use, modify, and… See the full description on the dataset page: https://huggingface.co/datasets/Bhavya-123/HRMS_Datahub.text1K<n<10K3 likes40 downloads3y agoHugging Face10kongclaves /market-brief-data-hub Market Brief Dataset Hub (ship with V3 model) Companion data for Market Brief Adapter V3 (Mixtral-8x7B LoRA, win rate 0.7806)Adaption AutoScientist Challenge 2026 · Market Analysis & News Story Artifact Role Model V3 Ship — grounded briefs + richer TrueNorth-enriched train path + Mixtral This hub Full data program: pilot → synthetic → TrueNorth → merged → Adaptive Data exports V3 training used Adaptive Data on the merged seed (8e47668b… /… See the full description on the dataset page: https://huggingface.co/datasets/kongclaves/market-brief-data-hub.texttext-generation1K<n<10K0 likes28 downloads2mo agoHugging Face11ml-hub /flipkart-datatabular10K<n<100K0 likes26 downloads2y agoHugging Face12bean980310 /ai-hub-conversation-datatext10K<n<100K0 likes23 downloads2y agoHugging Face13nltk-data-hub /names NLTK Names Corpus Name lists from NLTK, split by gender. Each gender is a separate config. Each row is one name. Usage from datasets import load_dataset ds = load_dataset("nltk-data-hub/names", "female") names = ds["names"]["name"] Schema Column Type Description name string The name Configs Config Count female 5,001 male 2,943 Source Originally distributed as part of nltk.download('names'). Converted to… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/names.texttext-classification1K<n<10K0 likes23 downloads5mo agoHugging Face14nltk-data-hub /gazetteers NLTK Gazetteers (Extended) Geographic and demographic word lists, extended from the original NLTK gazetteers corpus. Each config is one list; each row is one entry. Original configs (NLTK gazetteers corpus) Config Description License countries 289 country names GFDL (Wikipedia) isocountries 234 ISO country names public domain nationalities 200 nationality adjectives public domain caprovinces 14 Canadian provinces — mexstates 32 Mexican states —… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/gazetteers.tabulartoken-classification100K<n<1M0 likes23 downloads5mo agoHugging Face15BroDeadlines /TEST.HUB.mini.hub_copora_dataConfig data { "chunk_size": 800, "overlap": 60, "model": "gemini-1.5-flash-latest" } textn<1K0 likes17 downloads2y agoHugging Face16sergicalsix /Japanese_NER_Data_Hubgated 概要 大規模言語モデル(LLM)用の固有表現認識データセット(J-NER)のリポジトリです。 J-NERは拡張固有表現階層(*)の内、応用の観点から重要な名前表現から構成される総計157種類の固有表現を含んだデータセットです。 J-NERで取り扱う固有表現はLLMの学習データに含まれていることが要求されるため、J-NERに含まれる固有表現はWikipediaにページが存在する単語のみとしています。 各固有表現に関して、その固有表現を含んだデータ(文)である正例5例、含まれていないデータ(文)である負例5例がデータセットに存在します。 よってデータセットの総計は157×5×2=1,570です。 *:2024年5月28日現在、拡張固有表現階層の以下サイトにアクセスできません。 https://ene-project.info/ene9/ 使い方 データロード from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sergicalsix/Japanese_NER_Data_Hub.textfeature-extraction1K<n<10K5 likes15 downloads2y agoHugging Face17Hindi-data-hub /odaigen_hindi_pre_trained_spgatedHindi Language Pre-Trained LLM Datasets Overview Welcome to the Hindi Language Pre-Training Datasets repository! This README provides a comprehensive overview of various pre-training datasets available for Hindi, including essential details such as licenses, sources, and statistical information. These datasets are invaluable resources for training and fine-tuning large language models (LLMs) for a wide range of natural language processing (NLP) tasks. -Data Overview and Statistics This README… See the full description on the dataset page: https://huggingface.co/datasets/Hindi-data-hub/odaigen_hindi_pre_trained_sp.texttext-classification1M<n<10M8 likes15 downloads2y agoHugging Face18evals-hub /spider_data_goldtext10K<n<100K0 likes15 downloads11mo agoHugging Face19Chijioke-Mgbahurike /hub_base_ls960-ft_spot_data_allaudio1K<n<10K0 likes14 downloads2y agoHugging Face20kakao1513 /AI_HUB_legal_QA_datatext10K<n<100K0 likes14 downloads1y agoHugging Face21nltk-data-hub /dolch NLTK Dolch Sight Word List The 315 Dolch sight words (Dolch 1936), grouped by part of speech, distributed via NLTK. Configs Config Words Schema dolch 315 word, pos dolch-adjectives 46 word dolch-nouns 95 word dolch-verbs 92 word dolch-adverbs 34 word dolch-prepositions 16 word dolch-pronouns 26 word dolch-conjunctions 6 word Schema dolch — combined list with part-of-speech Column Type Description word string The sight… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/dolch.texttoken-classificationn<1K0 likes13 downloads5mo agoHugging Face22evals-hub /spider_data_bronzetext10K<n<100K0 likes12 downloads11mo agoHugging Face23Tech-Awesome-Hub /mix-datatext100K<n<1M0 likes11 downloads2y agoHugging Face24nltk-data-hub /swadesh NLTK Swadesh Word Lists Basic vocabulary lists for 24 languages, derived from the Wiktionary Swadesh list appendix and distributed via NLTK. Each config is one language; each row is one of the 207 Swadesh concepts. Languages Config Language Concepts swadesh-be Belarusian 207 swadesh-bg Bulgarian 207 swadesh-bs Bosnian 207 swadesh-ca Catalan 207 swadesh-cs Czech 207 swadesh-cu Church Slavonic 174 swadesh-de German 207 swadesh-en English 207… See the full description on the dataset page: https://huggingface.co/datasets/nltk-data-hub/swadesh.texttoken-classification1K<n<10K0 likes7 downloads5mo agoHugging Face25mynkchaudhry /marketing_data_hubtextn<1K1 likes6 downloads2y agoHugging Face26Seb0099 /dataHub_schema_jsonltext10K<n<100K0 likes6 downloads2y agoHugging Face27hubin /context_reasoner_data_mcqtext1K<n<10K0 likes6 downloads1y agoHugging Face28peru-opendata-hub /bcrp-data-hubtext1M<n<10M0 likes5 downloads1y agoHugging Face29scidata-hub /anais112-event-datatabular1K<n<10K0 likes5 downloads3mo agoHugging Face30infinite-dataset-hub /SituationResponse-Data SituationResponse-Data tags: Emergency Situations, Response Time, Predictive Modeling Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'SituationResponse-Data' CSV dataset is designed to support machine learning practitioners in developing predictive models for emergency call situation classification. It includes a variety of textual data from emergency calls, which are then labeled according to the nature of the emergency and… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/SituationResponse-Data.textn<1K1 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.