datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
neural-news-benchmark
AI-generated News Detection Benchmark
neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian.
Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024.
Dataset Details
The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.Security-TTP-Mapping
The Security Attack Pattern (TTP) Recognition or Mapping Task
We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits.
The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes.
NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects.
Datasets
TRAM
This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.span-similarity-dataset
Span Similarity Dataset (SSD)
Dataset Summary
The Span Similarity Dataset (SSD) focuses on Explainable Textual Similarity. It consists
of pairs of sentences with annotations pointing to both semantically equivalent and
dissimilar spans.
Languages
The SSD includes exclusively texts in English.
Dataset Structure
The dataset is split into -train (800 samples), -eval (100 samples), and -test (100
samples), all of them provided as a .tsv… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/span-similarity-dataset.cognitive-biases-in-llms
A Comprehensive Evaluation of Cognitive Biases in LLMs: Dataset
Dataset for evaluating cognitive biases in large language models
1. Dataset Card Overview
A tabular dataset for measuring the presence and strength of cognitive biases in large language models (LLMs), introduced in the paper “A Comprehensive Evaluation of Cognitive Biases in LLMs” by Malberg et al.
Paper | Code
This dataset is intended only for the evaluation of LLMs and not to be used for… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/cognitive-biases-in-llms.TumpengQA
Synthetic Indonesian dataset with Llama 3 70B
TumpengQA contains 6.7M words of 28.2K input-output pairs of Indonesian question-answering. It is intended to fine-tune Llama 3 8B, which has limited Indonesian language capabilities, to properly respond in Indonesian.
It is a research preview dataset and not curated for factual accuracy or safety. Use this dataset at your discretion.
Out of scope use
Commercial use
Fine-tuning non-Llama 3 models
echr_rational
Dataset Card for echr_rational
Dataset Summary
Deconfounding Legal Judgment Prediction for European Court of Human
Rights Cases Towards Better Alignment with Experts
This work demonstrates that Legal Judgement Prediction systems without expert-informed adjustments can be vulnerable to shallow, distracting surface signals that arise from corpus construction, case distribution, and confounding factors. To mitigate this, we use domain expertise to strategically identify… See the full description on the dataset page: https://huggingface.co/datasets/TUMLegalTech/echr_rational.cannot-dataset
Compilation of ANnotated, Negation-Oriented Text-pairs
Dataset Card for CANNOT
Dataset Summary
CANNOT is a dataset that focuses on negated textual pairs. It currently
contains 77,376 samples, of which roughly of them are negated pairs of
sentences, and the other half are not (they are paraphrased versions of each
other).
The most frequent negation that appears in the dataset is verbal negation (e.g.,
will → won't), although it also contains pairs with… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/cannot-dataset.sexism-socialmedia-balanced
Citation
@inproceedings{rydelek-etal-2023-adamr,
title = "{A}dam{R} at {S}em{E}val-2023 Task 10: Solving the Class Imbalance Problem in Sexism Detection with Ensemble Learning",
author = "Rydelek, Adam and
Dementieva, Daryna and
Groh, Georg",
editor = {Ojha, Atul Kr. and
Do{\u{g}}ru{\"o}z, A. Seza and
Da San Martino, Giovanni and
Tayyar Madabushi, Harish and
Kumar, Ritesh and
Sartori, Elisa},
booktitle = "Proceedings… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/sexism-socialmedia-balanced.rundi-tumbuka_sentence-pairs
Rundi-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Tumbuka_Sentence-Pairs
Number of Rows: 194527
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-tumbuka_sentence-pairs.TCGA-SARC-dict-tumortext2food-mmc4This dataset is filtered version of MMC4 Multimodal-C4 core fewer-faces dataset . It contains 144 474 pair of food image url and image caption.
All the code and model in the repository.
tigrinya-tumbuka_sentence-pairs
Tigrinya-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Tumbuka_Sentence-Pairs
Number of Rows: 152916… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-tumbuka_sentence-pairs.hovh_tumanyan_seq2seq
📝 Seq2Seq Dataset: Hovhannes Tumanyan's Poems
This dataset contains sentence pairs extracted from the poetic works of Hovhannes Tumanyan, one of the most celebrated Armenian poets. It is formatted for training sequence-to-sequence (Seq2Seq) models for tasks such as:
Text generation
Dialogue modeling
Style imitation
📂 Dataset Structure
The dataset is provided as a .csv file with two columns:
input_sentence
target_sentence
Line 1 of poem
Line 2 of poem
Line… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/hovh_tumanyan_seq2seq.igbo-tumbuka_sentence-pairs
Igbo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Tumbuka_Sentence-Pairs
Number of Rows: 133589
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-tumbuka_sentence-pairs.somali-tumbuka_sentence-pairs
Somali-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tumbuka_Sentence-Pairs
Number of Rows: 179589
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tumbuka_sentence-pairs.fon-tumbuka_sentence-pairs
Fon-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Tumbuka_Sentence-Pairs
Number of Rows: 73794
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-tumbuka_sentence-pairs.swahili-tumbuka_sentence-pairs
Swahili-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Tumbuka_Sentence-Pairs
Number of Rows: 499551
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-tumbuka_sentence-pairs.tswana-tumbuka_sentence-pairs
Tswana-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tswana-Tumbuka_Sentence-Pairs
Number of Rows: 187262
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tswana-tumbuka_sentence-pairs.kamba-tumbuka_sentence-pairs
Kamba-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Tumbuka_Sentence-Pairs
Number of Rows: 63077
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-tumbuka_sentence-pairs.tumbuka-twi_sentence-pairs
Tumbuka-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tumbuka-Twi_Sentence-Pairs
Number of Rows: 176324
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-twi_sentence-pairs.pedi-tumbuka_sentence-pairs
Pedi-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Tumbuka_Sentence-Pairs
Number of Rows: 101945
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-tumbuka_sentence-pairs.oromo-tumbuka_sentence-pairs
Oromo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Tumbuka_Sentence-Pairs
Number of Rows: 73849
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-tumbuka_sentence-pairs.nuer-tumbuka_sentence-pairs
Nuer-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Nuer-Tumbuka_Sentence-Pairs
Number of Rows: 19926
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/nuer-tumbuka_sentence-pairs.kongo-tumbuka_sentence-pairs
Kongo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Tumbuka_Sentence-Pairs
Number of Rows: 93758
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-tumbuka_sentence-pairs.kikuyu-tumbuka_sentence-pairs
Kikuyu-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Tumbuka_Sentence-Pairs
Number of Rows: 54854
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-tumbuka_sentence-pairs.dinka-tumbuka_sentence-pairs
Dinka-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Tumbuka_Sentence-Pairs
Number of Rows: 23524
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-tumbuka_sentence-pairs.VenusVaccine_TumorBinaryBrain_Tumor_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Brain Tumors
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
tumbuka-umbundu_sentence-pairs
Tumbuka-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tumbuka-Umbundu_Sentence-Pairs
Number of Rows: 99125
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-umbundu_sentence-pairs.tsonga-tumbuka_sentence-pairs
Tsonga-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tsonga-Tumbuka_Sentence-Pairs
Number of Rows: 203106
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tsonga-tumbuka_sentence-pairs.
