datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code_Vulnerability_Labeled_Dataset
Dataset Card for Code_Vulnerability_Labeled_Dataset
Dataset Summary
This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation:
CWE
Description
CWE-020
Improper Input Validation
CWE-022
Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”)
CWE-078
Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”)
CWE-079
Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.Train_routerArcSentiments-FinBERT-PT-BR
Dataset
A manually annotated dataset was created to enable supervised training for the FinBERT-PT-BR model, which focuses on sentiment analysis of Brazilian Portuguese financial texts.
More than 1.4 million financial news texts in Portuguese were collected and used for the initial language modeling phase. From this corpus, a sample of 1,000 texts was manually annotated with sentiment labels.
Annotation Process
Three annotators participated in the process.
All texts were… See the full description on the dataset page: https://huggingface.co/datasets/lucas-leme/Sentiments-FinBERT-PT-BR.kavram-yanilgisi-envanteri
Kavram Yanılgısı Envanteri — Türkiye ortaokul matematiği (6-8. sınıf)
505 kayıt · 31 kaynak · 24 sütun · Türkçe
Bu veri seti, Türkiye'de tam metnine erişilen 31 akademik çalışmanın (11 yüksek lisans
tezi + 20 dergi makalesi) okunarak kodlanmış kavram yanılgısı kayıtlarından oluşur.
Her satır bir yanılgıyı tanımlar; kaynağın künyesini ve sayfa numarasını taşır, bu
sayede her kayıt tek tek doğrulanabilir (505 kaydın 505'inde sayfa referansı vardır).
Kapsam
6, 7 ve… See the full description on the dataset page: https://huggingface.co/datasets/lemmaakademi/kavram-yanilgisi-envanteri.R-ViHSDtajik_lemmasVNNoiseBenchLyran-to-Englishlemi_lexical_lists
Romanian Grade Word Lists
Dataset Description
This dataset contains word lists grouped by grade level (1–4), derived from model-based difficulty predictions across multiple Romanian texts. Each word is associated with a mean grade score indicating its predicted difficulty level.
Dataset Structure
Each file corresponds to a grade level and contains the following columns:
cuvant: the Romanian word (token)
grad_mediu: mean predicted grade level for that word… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/lemi_lexical_lists.lemonde_genderThe full_data.csv file contains, for every article, the number of mentions of men/women, the number of citations of men/women, the author and its gender when it is clear. In the scripts folder, there is a code to make figures.
lemonde_ner_personsThis repository lists, for each article published by Le Monde since its founding in December 1944 to July 2024, the persons named in the text, separated by semicolons (";"). It is simply the result of a spacy NER, with the "fr_core_news_md": we analyse the document, take all the entities with "PER" (person) as type, and take the text of the entity. We also extract the grammatical function (entityt.root.dep_) corresponding to the entities. The ambition was to maybe, in the future, look at… See the full description on the dataset page: https://huggingface.co/datasets/regicid/lemonde_ner_persons.Customer-Service-Chatbotefeverde_5_cat_lemefeverde
ml-emoji-story26k-FAQs-DataBaseSalesdutchslang
Dutchslang - 1.0
DutchSlang-dataset is a dataset commited to translating slang (sms/straattaal) to "formal" dutch.
It contatins a few entries with categories, such as region, category, etc..
Warning: There are profanities inside of the dataset, these are listed as "profanity" in the category.
What can this be used for?
Not for training language models directly, but it can be used for:
Autocorrect
Translators
Finetuning Transformers / Tiny transformers that speak… See the full description on the dataset page: https://huggingface.co/datasets/im-lemon/dutchslang.EUS_lemma_lengthCato-Space-Dataclimate-bills-lemmed-count-vectorizerwellness-tourism-customersFinal_Merge_Dataclimate-news-tfidf-lemmedShort-DataCSClessDataCategorical-Data-Chatbotfeelgoodmix1Maneshatweets_lemmtest
