CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Nbardy /science-theory-textbookstext10K<n<100K9 likes5.1k downloads3y agoHugging Face02choucsan /NBA_Games NBA Full-Game Video Dataset This dataset provides metadata, official statistics, and official play-by-play annotations for full-length NBA game videos available on YouTube. Instead of redistributing video files, we provide YouTube video IDs and URLs so users can download videos independently when their use case and local policies allow it. The dataset links long-form basketball videos with structured NBA.com game data. Each retained game has a verified… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/NBA_Games.textvideo-classificationn<1K6 likes1.4k downloads2mo agoHugging Face03Nbardy /wild-science-theory-textbookstext10K<n<100K3 likes1.1k downloads3y agoHugging Face04NbAiLab /NCC Dataset Card for NbAiLab/NCC ⚠️ Important Update (December 2024) Previously, newspapers were a significant part of the Norwegian Colossal Corpus (NCC), particularly the newspapers distributed under the so called "Språkbank-avtalen". As of December 2024, at the request of media houses, we have ceased distributing newspapers under this agreement, including the "Norsk Aviskorpus." However, NCC still includes numerous newspapers that are released under more open… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NCC.texttext-generation1M<n<10M3 likes543 downloads2y agoHugging Face05Nbardy /art-theory-textbookstext10K<n<100K3 likes457 downloads3y agoHugging Face06NbAiLab /nbnn_language_detection Dataset Card for Bokmål-Nynorsk Language Detection (main_train_split) Dataset Summary This dataset is intended for language detection for Bokmål to Nynorsk and vice versa. It contains 800,000 sentence pairs, sourced from Språkbanken and pruned to avoid overlap with the NorBench dataset. The data comes from translations of news text from Norsk telegrambyrå (NTB), performed by Nynorsk pressekontor (NPK). In addition the dev and test set has 1000 entries. Data… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nbnn_language_detection.texttext-classification1M<n<10M3 likes398 downloads3y agoHugging Face07NbAiLab /norwegian_parliament Dataset Card Creation Guide Dataset Summary This is a classification dataset created from a subset of the Talk of Norway. This dataset contains text phrases from the political parties Fremskrittspartiet and Sosialistisk Venstreparti. The dataset is annotated with the party the speaker, as well as a timestamp. The classification task is to, simply by looking at the text, being able to predict is the speech was done by a representative from Fremskrittspartiet or from SV.… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/norwegian_parliament.texttext-classification1K<n<10K5 likes331 downloads2y agoHugging Face08NbAiLab /nb_distil_speech_noconcat_stortinget Dataset Card for NbAiLab/nb_distil_speech_noconcat_stortinget Dataset Summary NbAiLab/nb_distil_speech_noconcat_stortinget is a curated subset of the Stortinget Speech Corpus (SSC), a large-scale Norwegian parliamentary speech dataset. This subset focuses on non-concatenated speech segments and includes automatic transcriptions generated using OpenAI's Whisper model. It is designed to facilitate the development and evaluation of Automatic Speech Recognition (ASR) systems… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb_distil_speech_noconcat_stortinget.tabular100K<n<1M2 likes312 downloads1y agoHugging Face09NbAiLab /nb_distil_speech_noconcat_nsttabular100K<n<1M0 likes161 downloads2y agoHugging Face10NbAiLab /lunde_nor_nob_reading_optimisedTest only - not for training. First version - 0.1 of lunde_nor_nob_reading_optimised This dataset does not contain any audio data. Export Details Train samples: 10040932 Validation samples: 0 Test samples: 0 Dataset created using search datasets:lunde_nor_nob_reading_optimised. textautomatic-speech-recognition10M<n<100M0 likes56 downloads8mo agoHugging Face11LBJLincoln26 /nba-box-scorestextn<1K0 likes53 downloads5mo agoHugging Face12Nbardy /diverse-svg-prompts Diverse SVG Prompts Diverse SVG Prompts is a public collection of 20,000 high-quality, generated and filtered English briefs for SVG and vector-graphics generation. It contains 18,000 general illustration prompts and 2,000 lettering prompts. Schema The dataset intentionally has only two columns: prompt: the complete visual brief. type_tags: a list of category, author-model, and processing tags. Example: { "prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.texttext-generation10K<n<100K0 likes52 downloads27d agoHugging Face13Mr-Bridge /nba-home-court-2017-2026 NBA Home-Court Advantage 2017-2026: 11,777 Games What this is Every NBA regular-season game across 10 seasons (2016-17 through 2025-26): 11,777 games, with home and away team, final score, scoring margin, venue, and a neutral-site flag. Collected from ESPN's official sports API through the MrBridge ESPN MCP Server. No scraping; the data is public game results from the official feed. Companion study: NBA Home-Court Advantage: 10 Seasons, 11,777 Games — home… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Bridge/nba-home-court-2017-2026.text10K<n<100K0 likes45 downloads3mo agoHugging Face14NbAiLab /nynorsk_norm_200eval Nynorsk Norm 200eval nynorsk_norm_200eval is a high-quality, small-scale parallel corpus comprising 200 Norwegian Bokmål–Nynorsk sentence pairs collected from official sources and public institutions. Each example includes: nb: Original sentence in Bokmål nn_original: Original Nynorsk sentence (typically an official translation) nn_alt_original: Original Nynorsk sentence (typically an official translation) - alt version nn_husnorm: Sentence rewritten in Nynorsk following an… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nynorsk_norm_200eval.texttranslationn<1K1 likes44 downloads11mo agoHugging Face15open-llm-leaderboard /NbAiLab__nb-llama-3.1-8B-Instruct-detailsgated Dataset Card for Evaluation run of NbAiLab/nb-llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model NbAiLab/nb-llama-3.1-8B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NbAiLab__nb-llama-3.1-8B-Instruct-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face16NbAiLab /nynorsk_dpo Bokmål–Nynorsk DPO bokmal_nynorsk_dpo is a dataset for Direct Preference Optimization (DPO) training, focusing on Bokmål–Nynorsk translation.Each example consists of a prompt in Norwegian Bokmål and two candidate translations in Nynorsk: prompt: Input sentence in Bokmål chosen: Preferred Nynorsk translation (higher quality, closer to target norm) rejected: Less preferred Nynorsk translation This format enables reinforcement learning from human preferences, where models learn… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nynorsk_dpo.texttext-generationn<1K0 likes32 downloads1y agoHugging Face17NbAiLab /ndla_npk_conversational_nb_to_nn_tags_balanced Balanced version of NbAiLab/ndla_npk_conversational_nb_to_nn with tags. The corpus consists of: 70.000 samples from NbAiLab/ndla_npk_conversational_nb_to_nn 15.000 samples with tags from NDLA 15.000 samples with tags from NPL TOTAL: 100k samples The coprus is mainly made for GRPO-training text100K<n<1M0 likes23 downloads1y agoHugging Face18NbAiLab /freddy-testDette er et datasett som skal slettes. textautomatic-speech-recognition1K<n<10K0 likes14 downloads11mo agoHugging Face19NbAiLab /ndla_npk_conversational_nb_to_nn_tags_allcapstext1M<n<10M0 likes13 downloads1y agoHugging Face20LBJLincoln26 /nba-experiment-queuen<1K0 likes13 downloads5mo agoHugging Face21awsomedod /text-to-sql-nbatextn<1K0 likes11 downloads2y agoHugging Face22yann756 /nbatextn<1K1 likes11 downloads7mo agoHugging Face23NbAiLab /nb-asr-qwen3whisperxagreement-v1 nb-asr-qwen3whisperxagreement-v1 Word-level forced alignment training data for Norwegian speech, produced by keeping only examples where two independent aligners — WhisperX and Qwen3 (Lunde forced aligner) — agree within a tight tolerance. Dataset Description This dataset contains 702,067 speech segments drawn from the NB-ASR Norwegian audio corpus. Each record pairs an audio file with a word-level forced alignment in a format suitable for training a… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-qwen3whisperxagreement-v1.textautomatic-speech-recognition100K<n<1M0 likes11 downloads4mo agoHugging Face24open-llm-leaderboard /NbAiLab__nb-llama-3.1-8B-sft-detailsgated Dataset Card for Evaluation run of NbAiLab/nb-llama-3.1-8B-sft Dataset automatically created during the evaluation run of model NbAiLab/nb-llama-3.1-8B-sft The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NbAiLab__nb-llama-3.1-8B-sft-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face25Adun /nbatext1K<n<10K0 likes7 downloads2y agoHugging Face26Dc-4nderson /nba-classifiertexttext-classification1K<n<10K0 likes6 downloads1y agoHugging Face27aki005 /Recipes_json_vector_nbarbconetextn<1K0 likes3 downloads1y agoHugging Face28Thoop /nbatext1K<n<10K0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.