CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NbAiLab /NCC Dataset Card for NbAiLab/NCC ⚠️ Important Update (December 2024) Previously, newspapers were a significant part of the Norwegian Colossal Corpus (NCC), particularly the newspapers distributed under the so called "Språkbank-avtalen". As of December 2024, at the request of media houses, we have ceased distributing newspapers under this agreement, including the "Norsk Aviskorpus." However, NCC still includes numerous newspapers that are released under more open… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/NCC.texttext-generation1M<n<10M3 likes629 downloads2y agoHugging Face02pere /nb-asr-numerics-harvested Norwegian Bokmål Numeric Expression Harvesting Dataset This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn). This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.tabulartext-generation1M<n<10M0 likes112 downloads3mo agoHugging Face03NbAiLab /norwegian-alpaca NB Alpaca Norwegian Bokmål This dataset is a translation to Norwegian Bokmål of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford. An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI. texttext-generation10K<n<100K10 likes65 downloads3y agoHugging Face04Nbardy /diverse-svg-prompts Diverse SVG Prompts Diverse SVG Prompts is a public collection of 20,000 high-quality, generated and filtered English briefs for SVG and vector-graphics generation. It contains 18,000 general illustration prompts and 2,000 lettering prompts. Schema The dataset intentionally has only two columns: prompt: the complete visual brief. type_tags: a list of category, author-model, and processing tags. Example: { "prompt": "A moonlit mechanical heron..."… See the full description on the dataset page: https://huggingface.co/datasets/Nbardy/diverse-svg-prompts.texttext-generation10K<n<100K0 likes52 downloads28d agoHugging Face05pere /nb-asr-numerics-categorized Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized.tabulartext-generation1M<n<10M1 likes34 downloads3mo agoHugging Face06NbAiLab /nynorsk_dpo Bokmål–Nynorsk DPO bokmal_nynorsk_dpo is a dataset for Direct Preference Optimization (DPO) training, focusing on Bokmål–Nynorsk translation.Each example consists of a prompt in Norwegian Bokmål and two candidate translations in Nynorsk: prompt: Input sentence in Bokmål chosen: Preferred Nynorsk translation (higher quality, closer to target norm) rejected: Less preferred Nynorsk translation This format enables reinforcement learning from human preferences, where models learn… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nynorsk_dpo.texttext-generationn<1K0 likes32 downloads1y agoHugging Face07pere /nb-asr-numerics-balanced Balanced Synthetic Norwegian Bokmål Numerics Dataset This dataset provides a class-balanced synthetic corpus of Norwegian Bokmål sentences containing numeric expressions. It draws 10,000 examples for each of the 59 numeric categories (totaling 590,000 rows). Source & Synthesis Architecture Templates source: pere/nb-asr-numerics-categorized. Methodology: Filtered the original dataset for kept rows containing annotated entities. For each target category, sampled 10… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-balanced.texttext-generation100K<n<1M0 likes27 downloads3mo agoHugging Face08pere /nb-asr-numerics-categorized-smoke-test Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized-smoke-test.tabulartext-generationn<1K0 likes24 downloads3mo agoHugging Face09NLP-UniBW /tweets_about_NBA_players_in_playoffs_april_june_2025 Tweets about NBA Players in the 2025 Playoffs (Pseudonymized) A corpus of 1,428,813 English tweets about 244 NBA players, collected daily throughout the 2025 NBA postseason (18 April – 24 June 2025, 68 consecutive days), covering the play‑in tournament through the NBA Finals. All author-identifying fields have been replaced with keyed, consistent pseudonyms so the corpus can be used for scientific research. Pseudonyms are stable across the whole dataset, so reply graphs, quote… See the full description on the dataset page: https://huggingface.co/datasets/NLP-UniBW/tweets_about_NBA_players_in_playoffs_april_june_2025.tabulartext-classification1M<n<10M0 likes19 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.