NbAiLab/nbnn_language_detection
Dataset Card for Bokmål-Nynorsk Language Detection (main_train_split) Dataset Summary This dataset is intended for language detection for Bokmål to Nynorsk and vice versa. It contains 800,000 sentence pairs, sourced from Språkbanken and pruned to avoid overlap with the NorBench dataset. The data comes from translations of news text from Norsk telegrambyrå (NTB), performed by Nynorsk pressekontor (NPK). In addition the dev and test set has 1000 entries.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nbnn_language_detection.
etst
cleaned
cleaned
test
traina and b
nordic
tsv
debug
debug
debug
debug
debug
test
test
test
.
dataloader
dataloader
dataloader
dataloader
dataloader
dataloader
data
Set up git-lfs for jsonl-files
Update README.md
Create README.md
initial commit
