CoolFace
Datasetpublic

NbAiLab/nbnn_language_detection

Dataset Card for Bokmål-Nynorsk Language Detection (main_train_split) Dataset Summary This dataset is intended for language detection for Bokmål to Nynorsk and vice versa. It contains 800,000 sentence pairs, sourced from Språkbanken and pruned to avoid overlap with the NorBench dataset. The data comes from translations of news text from Norsk telegrambyrå (NTB), performed by Nynorsk pressekontor (NPK). In addition the dev and test set has 1000 entries.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nbnn_language_detection.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
3likes404downloads
27 commits on main
9520e7c3y ago

etst

pere
9b047f03y ago

cleaned

pere
cd7e35b3y ago

cleaned

pere
d580db33y ago

test

pere
2bd85cf3y ago

traina and b

pere
5ca0cbf3y ago

nordic

pere
aa431ee3y ago

tsv

pere
c14ff653y ago

debug

pere
db4f2463y ago

debug

pere
7696b523y ago

debug

pere
320f2003y ago

debug

pere
bb616b43y ago

debug

pere
21fc87c3y ago

test

pere
d546bcf3y ago

test

pere
f66a7e33y ago

test

pere
7f86e423y ago

.

pere
8c381683y ago

dataloader

pere
32843dd3y ago

dataloader

pere
b3ddbe23y ago

dataloader

pere
07192133y ago

dataloader

pere
1987fa73y ago

dataloader

pere
e831b8a3y ago

dataloader

pere
4bb171f3y ago

data

pere
c7e8a493y ago

Set up git-lfs for jsonl-files

pere
8fc4c233y ago

Update README.md

pere
223d6a73y ago

Create README.md

pere
0c58aad3y ago

initial commit

pere