CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexandrainst /scandi-reddit Dataset Card for ScandiReddit Dataset Summary ScandiReddit is a filtered and post-processed corpus consisting of comments from Reddit. All Reddit comments from December 2005 up until October 2022 were downloaded through PushShift, after which these were filtered based on the FastText language detection model. Any comment which was classified as Danish (da), Norwegian (no), Swedish (sv) or Icelandic (is) with a confidence score above 70% was kept. The resulting comments… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit.texttext-generation10M<n<100M5 likes636 downloads2y agoHugging Face02Lots-of-LoRAs /task131_scan_long_text_generation_action_command_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task131_scan_long_text_generation_action_command_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task131_scan_long_text_generation_action_command_long.texttext-generation1K<n<10K0 likes185 downloads2y agoHugging Face03Lots-of-LoRAs /task127_scan_long_text_generation_action_command_all Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task127_scan_long_text_generation_action_command_all Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task127_scan_long_text_generation_action_command_all.texttext-generation1K<n<10K0 likes135 downloads2y agoHugging Face04Lots-of-LoRAs /task128_scan_structured_text_generation_command_action_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task128_scan_structured_text_generation_command_action_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task128_scan_structured_text_generation_command_action_short.texttext-generation1K<n<10K0 likes107 downloads2y agoHugging Face05north /scandinavian-educational-annotations Scandinavian Educational Annotations Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash. texttext-generation100K<n<1M3 likes100 downloads2y agoHugging Face06Lots-of-LoRAs /task129_scan_long_text_generation_action_command_short Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task129_scan_long_text_generation_action_command_short Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task129_scan_long_text_generation_action_command_short.texttext-generation1K<n<10K0 likes100 downloads2y agoHugging Face07north /scandinavian-linguistic-annotations Scandinavian Educational Annotations Created using a CommonCrawl dump (April 2024), and annotations with Gemini 1.5 Flash. texttext-generation100K<n<1M0 likes81 downloads2y agoHugging Face08alexandrainst /scandi-wikiScandiWiki is a parsed and deduplicated version of the Danish, Norwegian Bokmål, Norwegian Nynorsk, Swedish, Icelandic and Faroese Wikipedia corpora, as of January 2023.textfill-mask1M<n<10M4 likes78 downloads4y agoHugging Face09Lots-of-LoRAs /task130_scan_structured_text_generation_command_action_long Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task130_scan_structured_text_generation_command_action_long Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task130_scan_structured_text_generation_command_action_long.texttext-generation1K<n<10K0 likes76 downloads2y agoHugging Face10dhruvilp /FinDoc-Pro-Scannedgated FinDoc-Pro-Scanned Dataset Dataset Description This dataset contains OCR and document parsing results for scanned financial documents using four different models: PaddleOCR-VL - 6 tasks per image (ocr, table, formula, chart, spotting, seal) Qwen3-VL - 3 prompts per image (ocr_json, qwenvl_html, qwenvl_markdown) Bonsai-2-27B - 2 active prompts per image (spotting_json, qwenvl_markdown) with thinking traces; partial ocr_markdown/qwenvl_html Docling-Heron-101 -… See the full description on the dataset page: https://huggingface.co/datasets/dhruvilp/FinDoc-Pro-Scanned.tabularimage-to-text10K<n<100K0 likes22 downloads2d agoHugging Face11ClarusC64 /linical_inference_debt_scanner_v0.1Clinical Inference Debt Scanner PurposeDetect when a clinical plan relies on stacked assumptions rather than evidence. You receive: evidence_signals a narrative_chain a planned_action You output: inference_debt_level0 to 3 debt_itemthe single most dangerous leap paydown_stepthe corrective step that restores evidence grounding Debt scale0 none1 minor2 moderate3 severe Scoring debt_level_scoregraded by distance from gold debt_item_similaritytoken overlap similarity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/linical_inference_debt_scanner_v0.1.texttext-classificationn<1K0 likes18 downloads8mo agoHugging Face12V4ldeLund /scandi-human-instruct scandi-human-instruct Human–LLM instruction tuning dataset for the four main Scandinavian languages. Rows are normalized into a chat-friendly messages format and tagged with the detected language using NbAiLab/nb-nordic-lid. Dataset size Total rows: 9,642 Per-language counts: Norwegian Bokmål (nb): 4,833 Danish (da): 2,633 Swedish (sv): 2,042 Norwegian Nynorsk (nn): 134 Sources allenai/WildChat-4.8M ministere-culture/comparia-conversations (both… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/scandi-human-instruct.texttext-generation1K<n<10K0 likes17 downloads9mo agoHugging Face13Lots-of-LoRAs /task126_scan_structured_text_generation_command_action_all Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task126_scan_structured_text_generation_command_action_all Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task126_scan_structured_text_generation_command_action_all.texttext-generation1K<n<10K0 likes16 downloads2y agoHugging Face14V4ldeLund /scandi-translated-instruct scandi-translate A Scandinavian instruction‑tuning dataset built from machine‑translated instruction/response pairs. It unifies Danish, Swedish, Norwegian Bokmål and Norwegian Nynorsk data into a chat-friendly messages schema. Dataset size Total rows: 1,252,683 Per-language counts measured during build: Danish (da): 377,413 Swedish (sv): 332,558 Norwegian Nynorsk (nn): 270,772 Norwegian Bokmål (nb): 268,875 By source: V4ldeLund/da-translated-instruct… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/scandi-translated-instruct.texttext-generation1M<n<10M0 likes16 downloads9mo agoHugging Face15alexandrainst /scandi-reddit-filtered Dataset Card for ScandiRedditFiltered Dataset Summary ScandiRedditFiltered is manually filtered and post-processed corpus consisting of comments from ScandiReddit. The intended use of the filtered sentences is for Text-To-Speech (TTS) models. Supported Tasks and Leaderboards Training language models is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit-filtered.texttext-generation1K<n<10K0 likes14 downloads2y agoHugging Face16nafisehNik /ganjoor-ipa-scansiongated Ganjoor Persian Classical Poetry — Meter & Phonemic Transliteration A corpus of 124,404 classical Persian poems collected via the Ganjoor API, enriched with two things every poem now has: Prosodic meter (ʿarūż / vazn) — the metrical feet and a binary scansion for every poem, including the ~21% that Ganjoor left unlabeled (reconstructed here from the Persian feet). Phonemic transliteration — Latin and IPA for every hemistich, produced by the Homo-GE2PE grapheme-to-phoneme model.… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/ganjoor-ipa-scansion.tabulartext-generation100K<n<1M0 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.