datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
errors
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Nei]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/katsukiai/errors.nanbeige-ai-errors
Technical Challenge: Blind Spots of Frontier Models
This dataset was created as part of a technical challenge to identify and document the "blind spots" of a recent, moderately-sized base model. The goal was to browse models released in the last 6 months (between 0.6B and 6B parameters), select one, and systematically probe its failures to understand its limitations.
Dataset: Nanbeige4.1-3B AI Errors
This dataset contains 10 examples where the Nanbeige/Nanbeige4.1-3B… See the full description on the dataset page: https://huggingface.co/datasets/abdulrahman245/nanbeige-ai-errors.Referencing_Errors_Synthetic_EN
Synthetic Dataset for Automatic Error Correction in Referencing
This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna.
Dataset Description
corrupted_sentence: the sentence containing the referencing error
clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_EN.parlabe-errors-genere-60k
Dataset d'errors de gènere en català (60.000 parelles)
Aquest dataset conté 60.000 parelles de frases amb errors de gènere, generades mitjançant un script automatitzat en Python. Cada parella està formada per:
text_erroni,text_correcte
Objectiu
Aquest conjunt de dades està pensat per entrenar models de correcció gramatical en català, amb un focus específic en la detecció i correcció d’errors de gènere:
Concordança nominal (ex: el noia → la noia)
Concordança adjectival… See the full description on the dataset page: https://huggingface.co/datasets/Oriolshhh/parlabe-errors-genere-60k.Errors_Mod This dataset is the result of errors found in generated ecore files by different LLMs, mainly GTP4-Turbo and Llama3-70b-Instruct.
The errors have been classified into :
Wrong Type : This can occur if the generated type is non existant or used in a wrong way
Missing declaration : this can be due to either a missing declaration like xsi or nonexistant one
Start Token : this can mostly be due to start tag <?xml ..> <ecore ..> that are missing, happens when we can't read the file or error in… See the full description on the dataset page: https://huggingface.co/datasets/VeryMadSoul/Errors_Mod.tulu_3_formatting_errorsparlabe-errors-ortografia-45k
Dataset d’errors ortogràfics en català (45.000 parelles)
Aquest dataset conté 45.000 parelles de frases en format:
text_erroni,text_correcte
Està dissenyat per entrenar models de correcció ortogràfica general en català, abastant una gran varietat d’errors comuns en l’escriptura manual, digitació ràpida, ASR o OCR.
Què inclou?
Aquestes parelles cobreixen errors com:
Lletres intercanviades o repetides: Axo és una prova → Això és una prova
Omissions o afegits de caràcters:… See the full description on the dataset page: https://huggingface.co/datasets/Oriolshhh/parlabe-errors-ortografia-45k.granite-base-model-errors
Granite-1B Base Model Errors
Overview
This dataset contains 10 examples where the Granite-4.0-1B-Base language model produces incorrect or awkward outputs. Each row includes:
id: a unique identifier for each example
input: the prompt given to the model
expected_output: what the correct answer or completion should be
model_output: what the model actually produced
The dataset demonstrates common blind spots of a base causal language model, including factual errors, logic… See the full description on the dataset page: https://huggingface.co/datasets/thatgirltomiie/granite-base-model-errors.tulu_3_all_errors_updFunny-Windows-Errors-Windows-Biblesparlabe-errors-castellanismes-35k
Dataset de castellanismes en català (35.000 parelles)
Aquest dataset conté 35.000 parelles de frases en format:
text_erroni,text_correcte
El conjunt inclou frases amb castellanismes habituals en el català oral i escrit, generades de manera sintètica per entrenar models que detectin i corregeixin aquestes interferències.
Com s’ha generat?
Les parelles han estat generades mitjançant:
API de GPT per introduir errors de castellanismes de forma natural
Filtratge automàtic… See the full description on the dataset page: https://huggingface.co/datasets/Oriolshhh/parlabe-errors-castellanismes-35k.parlabe-errors-accents-85k
Dataset d’errors d’accentuació en català (85.000 parelles)
Aquest dataset conté 85.000 parelles de frases en format:
text_erroni,text_correcte
Està orientat a entrenar models de correcció d’errors d’accentuació i dièresi en català, especialment útil en contextos com OCR, ASR, o entrades sense accents en teclats mòbils.
Com s’ha generat?
S’ha seleccionat 85.000 frases que contenen accents (é, à, í, etc.) o dièresi (ï, ü)
A cada frase se li ha eliminat qualsevol signe… See the full description on the dataset page: https://huggingface.co/datasets/Oriolshhh/parlabe-errors-accents-85k.Referencing_Errors_Synthetic_DE
Synthetic Dataset for Automatic Error Correction in Referencing
This dataset includes 4,600 parallel sentences for Automatic Error Correction in referencing in German.It was synthetically created with gpt-4o-mini model according to the Institutional Guidelines of the Center for Translation Studies (CTS), University of Vienna.
Dataset Description
corrupted_sentence: the sentence containing the referencing error
clean_sentence: the correct version of the corrupted sentence… See the full description on the dataset page: https://huggingface.co/datasets/elizaveta-dev/Referencing_Errors_Synthetic_DE.parlabe-errors-conjugacio-45k
Dataset d’errors de conjugació verbal (45.000 parelles)
Aquest dataset conté 45.000 parelles de frases en format:
text_erroni,text_correcte
Està dissenyat per entrenar models de correcció d’errors de conjugació verbal en català, un problema habitual tant en l’escriptura espontània com en transcripcions automàtiques.
Què inclou?
El dataset cobreix diversos tipus d’errors verbals, com:
Canvis de persona o temps verbal→ Nosaltres arribava tard → Nosaltres arribàvem tard… See the full description on the dataset page: https://huggingface.co/datasets/Oriolshhh/parlabe-errors-conjugacio-45k.c_errorsMI-errorstulu_3_no_errorsM_errorstulu_3_factual_errors
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/anasedova/tulu_3_factual_errors.tulu_3_incorrect_output_errorsjava_common_errorstulu_3_underspecified_input_errorstulu_3_model_modality_mismatch_errorsL1_true_errorssinhala-dyslexia-assistant-articulation-errorsparlabe-errors-varis-450k
Dataset mixt de correcció gramatical en català (450.000 parelles)
Aquest dataset conté 450.000 parelles de frases en format:
text_erroni,text_correcte
Està pensat per entrenar models de correcció gramatical robusta en català, abastant una àmplia gamma d’errors habituals, tant d’aprenents com de parlants nadius o transcripcions automàtiques.
Tipus d’errors inclosos
Aquestes parelles cobreixen diferents tipus d’errors sintètics generats controladament:
Castellanismes:… See the full description on the dataset page: https://huggingface.co/datasets/Oriolshhh/parlabe-errors-varis-450k.smollm3-base-errors
Model Tested
HuggingFaceTB/SmolLM3-3B
How I Loaded the Model
!pip install transformers torch accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "HuggingFaceTB/SmolLM3-3B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
def generate(prompt, max_new_tokens=100):
inputs = tokenizer(prompt… See the full description on the dataset page: https://huggingface.co/datasets/mhahmad/smollm3-base-errors.movie_titles_with_errors
