datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SimpleSafetyTestsbert_dataset_202203
Dataset Card for "bert_dataset_202203"
More Information needed
mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.alpaca-spanish
BERTIN Alpaca Spanish
This dataset is a translation to Spanish of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford.
An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI.
FineFineWeb-bert-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.TinyPython
TinyPython Tasks
TinyPython is a synthetic Python dataset inspired by the idea behind TinyStories: if the data distribution is narrow, clean, and high quality, even very small language models can learn useful structure.
Instead of broad repository code or competitive-programming solutions, TinyPython focuses on short natural-language programming tasks paired with complete, typed, standalone Python functions.
The goal is to provide a compact instruction-to-code corpus for… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/TinyPython.voice-light-tool-use-synthetic
Voice Light Teacher-Led Tool-Use Synthetic
This repository contains the current canonical synthetic source dataset for Voice Light's
conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English
conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation
has four user turns so follow-up requests can depend naturally on prior turns and tool results.
The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.llama3-ultrafeedback-bertscore-bart-large-mnli
RefAlign: LLM Alignment Dataset
This dataset is used in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.
Code: https://github.com/mzhaoshuai/RefAlign
This dataset is modified from https://huggingface.co/datasets/princeton-nlp/llama3-ultrafeedback. We use the BERTScore to choose the chosen and rejected responses.
Item with key ['Llama3.3-70B-Inst-Awq'] is the reference answers generated by… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/llama3-ultrafeedback-bertscore-bart-large-mnli.marc2
MARC2: Metaphor Abstraction and Reasoning Corpus v2
MARC2 extends the MARC-from-LARC methodology to the ARC-AGI2 dataset. It provides a corpus of figurative language puzzles where metaphorical descriptions help AI models solve abstract reasoning tasks they cannot solve from examples alone.
The MARC Property
A task has the MARC property (for a given model) when:
Examples alone fail — the model cannot solve the task from input/output examples
Figurative description… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc2.Turkish-Municipality-Instruction-Tuning-Datasetpl-bert
PL-BERT Malgache (mimba/pl-bert)
Ce dataset contient des phrases en malgache hautement nettoyées et filtrées, spécialement formatées pour l'entraînement de modèles de traitement du langage naturel orientés Text-To-Speech (TTS), comme PL-BERT (vocal-cloning et alignement).
Les données proviennent de sources combinées après l'application de filtres de qualité stricts (longueur des phrases, dédoublonnage, exclusion des caractères parasites d'interfaces et des langues étrangères).… See the full description on the dataset page: https://huggingface.co/datasets/mimba/pl-bert.
