CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes992 downloads4y agoHugging Face02bertin-project /alpaca-spanish BERTIN Alpaca Spanish This dataset is a translation to Spanish of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford. An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI. texttext-generation10K<n<100K36 likes662 downloads4y agoHugging Face03bertin-project /mc4-samplingA sampling-enabled version of mC4, the colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is a version of the processed version of Google's mC4 dataset by AllenAI, in which sampling methods are implemented to perform on the fly.text-generationn<1K13 likes195 downloads2y agoHugging Face04open-llm-leaderboard-old /details_bertin-project__bertin-gpt-j-6B-alpaca Dataset Card for Evaluation run of bertin-project/bertin-gpt-j-6B-alpaca Dataset Summary Dataset automatically created during the evaluation run of model bertin-project/bertin-gpt-j-6B-alpaca on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_bertin-project__bertin-gpt-j-6B-alpaca.0 likes151 downloads3y agoHugging Face05bertin-project /zenobia-instruct-hf Zenobia Instruct This dataset has been extracted from alvp/zenobia and alvp/stanzas, and parsed into a huggingface-friendly format so you can use apply_chat_template as explained on the Chat Templating documentation. Example from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.1") chat = [ {"role": "user", "content": "Escribe un terceto sobre la naturaleza en un paisaje nevado."}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/bertin-project/zenobia-instruct-hf.text10K<n<100K0 likes71 downloads2y agoHugging Face06bertin-project /BOE-XSUM BOE-XSUM Balanced Dataset - Reviewed and Cleaned Description The BOE 2025 Dataset is a collection of BOE articles with extreme summaries of them. This dataset has been carefully balanced and cleaned to ensure its quality and usefulness in natural language processing (NLP) tasks, primarily for evaluating generative models. Read more in https://arxiv.org/abs/2509.24908 Dataset Content The dataset is composed of the following subsets (splits): train: Training… See the full description on the dataset page: https://huggingface.co/datasets/bertin-project/BOE-XSUM.textsummarization1K<n<10K0 likes58 downloads1y agoHugging Face07bertin-project /bonanza-hf Bonanza: Dataset de instrucciones en Español y Catalán Este dataset combina múltiples fuentes para proporcionar instrucciones en español y catalán. Los datasets combinados son los siguientes: OpenAssistant/oasst2 CohereForAI/aya_dataset projecte-aina/RAG_Multilingual bertin-project/alpaca-spanish dariolopez/Llama-2-databricks-dolly-oasst1-es projecte-aina/MentorESprojecte-aina/MentorCA Descripción Este conjunto de datos proporciona una rica colección de instrucciones… See the full description on the dataset page: https://huggingface.co/datasets/bertin-project/bonanza-hf.text100K<n<1M2 likes57 downloads2y agoHugging Face08bertin-project /oasst2_es_instruct_hfThis is the Spanish subset from the OpenAssistant/oasst2 dataset. The dataset has been extracted from the 2023-11-05_oasst2_ready.trees.jsonl.gz file to parse all the conversation trees and put it in a huggingface-friendly format so you can use apply_chat_template as explained on the Chat Templating documentation. Example from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.1") chat = [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/bertin-project/oasst2_es_instruct_hf.text10K<n<100K2 likes36 downloads2y agoHugging Face09CyberHarem /emile_bertin_azurlane Dataset of emile_bertin/エミール・ベルタン/埃米尔·贝尔汀 (Azur Lane) This is the dataset of emile_bertin/エミール・ベルタン/埃米尔·贝尔汀 (Azur Lane), containing 40 images and their tags. The core tags of this character are blonde_hair, breasts, long_hair, blue_eyes, bow, hair_bow, blue_bow, bangs, large_breasts, very_long_hair, wavy_hair, medium_breasts, which are pruned in this dataset. Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/emile_bertin_azurlane.text-to-imagen<1K0 likes15 downloads3y agoHugging Face10Bertinho24 /BAENMIXX0 likes4 downloads3y agoHugging Face11Bertin123 /practica-unirtext1K<n<10K0 likes4 downloads2y agoHugging Face12Bertin123 /practica-2text1K<n<10K0 likes4 downloads2y agoHugging Face13maikagarbes /BERTinggatedtabular10K<n<100K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.