CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lockon /glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en. texttext-generation1K<n<10K1 likes27k downloads2y agoHugging Face02ZomiLearner /English-Zomi-OPUS_Tatoeba_v20230412 English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.tabulartranslation1M<n<10M0 likes15k downloads7mo agoHugging Face03leiwx52 /CC_eng_urltext100M<n<1B0 likes7k downloads2y agoHugging Face04SetFit /enron_spamThis is a version of the Enron Spam Email Dataset, containing emails (subject + message) and a label whether it is spam or ham. tabular10K<n<100K21 likes5.5k downloads5y agoHugging Face05maritaca-ai /enemThe ENEM 2022, 2023 and 2024 datasets encompass all multiple-choice questions from the last two editions of the Exame Nacional do Ensino Médio (ENEM), the main standardized entrance examination adopted by Brazilian universities. The datasets have been created to allow the evaluation of both textual-only and textual-visual language models. To evaluate textual-only models, we incorporated into the datasets the textual descriptions of the images that appear in the questions' statements from the… See the full description on the dataset page: https://huggingface.co/datasets/maritaca-ai/enem.textvisual-question-answeringn<1K15 likes3.4k downloads2y agoHugging Face06weaviate /enron-qa-emails-dasovich-jtext1K<n<10K0 likes3.2k downloads1y agoHugging Face07Abirate /english_quotes Dataset Card for English quotes I-Dataset Summary english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond. II-Supported Tasks and Leaderboards Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.texttext-classification1K<n<10K109 likes2.9k downloads4y agoHugging Face08bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.8k downloads4y agoHugging Face09weaviate /enron-qa-questions-dasovich-jtext10K<n<100K0 likes2.7k downloads1y agoHugging Face10nreimers /sphere_cohere_embed-english-v3.0text1M<n<10M0 likes2.7k downloads3y agoHugging Face11SetFit /amazon_counterfactual_en Amazon Counterfactual Statements This dataset is the en-ext split from SetFit/amazon_counterfactual. As the original test set is rather small (1333 examples), a different split was created with 50-50 for training & testing. The dataset is described in amazon-multilingual-counterfactual-dataset / Paper It contains statements from Amazon reviews about events that did not or cannot take place. text10K<n<100K0 likes2.5k downloads5y agoHugging Face12SetFit /amazon_reviews_multi_entext100K<n<1M7 likes2.4k downloads4y agoHugging Face13LLM-PBE /enron-emailThis dataset includes emails from Enron Email Dataset with prompts processed from Are Large Pre-Trained Language Models Leaking Your Personal Information?. To use the dataset, you can run the following in LLM-PBE. from data.enron import EnronDataset ds = EnronDataset(data_path="data/enron", pseudonymize=False) text100K<n<1M6 likes1.9k downloads2y agoHugging Face14judes1209 /exodus-endpointstextn<1K0 likes1.8k downloads4d agoHugging Face15RJZ /wikidata_triple_entext10M<n<100M1 likes1.7k downloads2y agoHugging Face16mteb /cqadupstack-english CQADupstackEnglishRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackEnglishRetrieval"]) evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-english.texttext-retrieval10K<n<100K1 likes1.6k downloads1y agoHugging Face17SetFit /amazon_massive_intent_en-UStext10K<n<100K10 likes1.5k downloads4y agoHugging Face18boffire /libretranslate-en-kab-suggestions Kabyle Suggestions Dataset This dataset contains English-to-Kabyle translation suggestions submmitted by users using LibreTranslate, designed to support the development and evaluation of machine translation tools for the Kabyle language. texttranslationn<1K0 likes1.4k downloads4mo agoHugging Face19masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes1k downloads3y agoHugging Face20Sachin21112004 /news-entertainment-datasettexttable-question-answeringn<1K6 likes1k downloads2h agoHugging Face21bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes990 downloads3y agoHugging Face22britllm /TransWeb-Edu-Englishtext10M<n<100M0 likes989 downloads2y agoHugging Face23tripolskypetr /trading-entries TradingView Analyst Grading via Risk-Management Grid Sweep This dataset answers a single question: does an analyst's call differ from the market spread — and if so, at what cost. The cost here is the investor's risk management: how deep a stop they will tolerate, how many days their money stays frozen, and at what percentage they lock in profit. The basis is TradingView Ideas posts on crypto carrying a LONG / SHORT direction. Every post is swept across 21,280 points of a… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/trading-entries.tabulartime-series-forecasting10K<n<100K0 likes980 downloads2mo agoHugging Face24Fristail27 /vocab-bloom-hub-en Vocab Bloom Hub — English A structured English lexical dataset with translations into Russian, Spanish, French, German, Portuguese, Chinese and Arabic, maintained by the Vocab Bloom Hub project — documentation, the API reference and a playground at vocab-bloom-hub.com. Every entry carries IPA transcription, a CEFR level, one or more sense-level definitions with usage examples, synonym and antonym links per sense, translations per sense in seven languages, and inflected forms —… See the full description on the dataset page: https://huggingface.co/datasets/Fristail27/vocab-bloom-hub-en.texttranslation1M<n<10M2 likes947 downloads4d agoHugging Face25sasha /co2_energy_datatabularn<1K0 likes943 downloads3y agoHugging Face26Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes928 downloads1y agoHugging Face27Enxin /Video-MMLU Video-MMLU Benchmark Resources Website arXiv: Paper GitHub: Code Huggingface: Video-MMLU Benchmark Features Benchmark Collection and Processing Video-MMLU specifically targets videos that focus on theorem demonstrations and probleming-solving, covering mathematics, physics, and chemistry. The videos deliver dense information through numbers and formulas, pose significant challenges for video LMMs in dynamic OCR… See the full description on the dataset page: https://huggingface.co/datasets/Enxin/Video-MMLU.textvideo-text-to-text1K<n<10K13 likes893 downloads1y agoHugging Face28marin-dna /gpn-star-p-uniform-v1-enhancer-arm-a marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.tabular100M<n<1B0 likes856 downloads28d agoHugging Face29EnzoPrezoto /Kathetotextn<1K0 likes830 downloads3y agoHugging Face30marin-dna /zoonomia-v1-v4_ccre_noexon_enhancer bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer A curated enhancer training set for issue #326 — a de-contaminated derivation of the v4 ccre_non_promoter arm of bolinas-dna/zoonomia-v1-v1, built by the snakemake/zoonomia_projection_dataset pipeline at commit 6b320c268547. Provenance This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.tabular10M<n<100M0 likes830 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.