CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bertievidgen /SimpleSafetyTeststexttext-generationn<1K12 likes3.3k downloads3y agoHugging Face02nthngdy /bert_dataset_202203 Dataset Card for "bert_dataset_202203" More Information needed texttext-generation100M<n<1B0 likes1.2k downloads4y agoHugging Face03bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes730 downloads4y agoHugging Face04bertin-project /alpaca-spanish BERTIN Alpaca Spanish This dataset is a translation to Spanish of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford. An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI. texttext-generation10K<n<100K36 likes614 downloads4y agoHugging Face05m-a-p /FineFineWeb-bert-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.texttext-classification1M<n<10M2 likes379 downloads2y agoHugging Face06BertilBraun /TinyPython TinyPython Tasks TinyPython is a synthetic Python dataset inspired by the idea behind TinyStories: if the data distribution is narrow, clean, and high quality, even very small language models can learn useful structure. Instead of broad repository code or competitive-programming solutions, TinyPython focuses on short natural-language programming tasks paired with complete, typed, standalone Python functions. The goal is to provide a compact instruction-to-code corpus for… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/TinyPython.texttext-generation1M<n<10M1 likes109 downloads3mo agoHugging Face07BertilBraun /voice-light-tool-use-synthetic Voice Light Teacher-Led Tool-Use Synthetic This repository contains the current canonical synthetic source dataset for Voice Light's conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation has four user turns so follow-up requests can depend naturally on prior turns and tool results. The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.texttext-generation1K<n<10K0 likes77 downloads2mo agoHugging Face08mzhaoshuai /llama3-ultrafeedback-bertscore-bart-large-mnli RefAlign: LLM Alignment Dataset This dataset is used in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data. Code: https://github.com/mzhaoshuai/RefAlign This dataset is modified from https://huggingface.co/datasets/princeton-nlp/llama3-ultrafeedback. We use the BERTScore to choose the chosen and rejected responses. Item with key ['Llama3.3-70B-Inst-Awq'] is the reference answers generated by… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/llama3-ultrafeedback-bertscore-bart-large-mnli.texttext-generation10K<n<100K0 likes68 downloads11mo agoHugging Face09bertybaums /marc2 MARC2: Metaphor Abstraction and Reasoning Corpus v2 MARC2 extends the MARC-from-LARC methodology to the ARC-AGI2 dataset. It provides a corpus of figurative language puzzles where metaphorical descriptions help AI models solve abstract reasoning tasks they cannot solve from examples alone. The MARC Property A task has the MARC property (for a given model) when: Examples alone fail — the model cannot solve the task from input/output examples Figurative description… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc2.tabulartext-generation10K<n<100K0 likes44 downloads5mo agoHugging Face10bert-ka /Turkish-Municipality-Instruction-Tuning-Datasettexttext-generationn<1K0 likes17 downloads1mo agoHugging Face11mimba /pl-bertgated PL-BERT Malgache (mimba/pl-bert) Ce dataset contient des phrases en malgache hautement nettoyées et filtrées, spécialement formatées pour l'entraînement de modèles de traitement du langage naturel orientés Text-To-Speech (TTS), comme PL-BERT (vocal-cloning et alignement). Les données proviennent de sources combinées après l'application de filtres de qualité stricts (longueur des phrases, dédoublonnage, exclusion des caractères parasites d'interfaces et des langues étrangères).… See the full description on the dataset page: https://huggingface.co/datasets/mimba/pl-bert.texttext-generation1M<n<10M1 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.