CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01netop /TeleQnAgated TeleQnA A Benchmark for Large Language Models in Telecom Knowledge Developed by the NetOp Team, Huawei Paris Research Center Ali Maatouk · Fadhel Ayed · Nicola Piovesan · Antonio De Domenico · Merouane Debbah · Zhi-Quan Luo 📄 Read the Paper 🤗 Explore the Dataset TeleQnA is a comprehensive dataset tailored to assess the knowledge of Large Language Models (LLMs) in the field of telecommunications. It… See the full description on the dataset page: https://huggingface.co/datasets/netop/TeleQnA.textquestion-answering10K<n<100K24 likes168 downloads1y agoHugging Face02MicPie /unpredictable_bulbapedia-bulbagarden-netThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes164 downloads4y agoHugging Face03NeTSlab /BLiMP-IT BLiMP-IT Dataset Summary BLiMP-IT is a linguistically motivated benchmark for evaluating Italian language models through minimal pairs. Each example consists of a grammatical sentence paired with a minimally different ungrammatical counterpart that isolates a single morphosyntactic contrast. The benchmark is designed to evaluate whether language models assign higher probability to the grammatical sentence than to the ungrammatical one. The benchmark is inspired by… See the full description on the dataset page: https://huggingface.co/datasets/NeTSlab/BLiMP-IT.texttext-classification1K<n<10K0 likes68 downloads1mo agoHugging Face04netop /TeleMathgated TeleMath A Benchmark for Large Language Models in Telecom Mathematical Problem Solving Developed by the NetOp Team, Huawei Paris Research Center Vincenzo Colle · Mohamed Sana · Nicola Piovesan · Antonio De Domenico · Fadhel Ayed · Merouane Debbah 📄 Read the Paper 🤗 Explore the Dataset [!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or… See the full description on the dataset page: https://huggingface.co/datasets/netop/TeleMath.textquestion-answeringn<1K11 likes66 downloads8mo agoHugging Face05justicedao /wetwijzer_netherlands_legal_corpus WetWijzer Netherlands Legal Corpus Hugging Face target: justicedao/wetwijzer_netherlands_legal_corpus. This unified dataset bundles the quality-audited WetWijzer Netherlands legal corpus stack in one repository for frontend retrieval. It preserves the existing compatibility repositories and does not replace or delete them. Contents Laws: 4,999 Articles: 89,737 CID index rows: 94,736 Vector mapping rows: 94,736 BM25 document rows: 94,736 BM25 term rows: 120,521… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/wetwijzer_netherlands_legal_corpus.tabulartext-retrieval100K<n<1M0 likes65 downloads3mo agoHugging Face06Neura-parse /quantum-networking-and-distributed Neura Parse — Quantum Networking, Repeaters & Distributed Quantum Computing A systems-frontier vertical on connecting quantum devices: entanglement distribution and distillation, quantum repeaters, quantum-internet protocol stacks, quantum memories/transduction, and modular/distributed quantum computing (nonlocal gates, circuit knitting across nodes, blind/verifiable delegated computation). Covers protocol and simulation methods used with tools such as NetSquid and SeQUeNCe… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-networking-and-distributed.tabulartext-generation100K<n<1M0 likes62 downloads3mo agoHugging Face07NASP /neteval-examNetEval is a NetOps evaluation suite for foundation models, consisting of 5269 multi-choice questions. Please check our paper for more details about NetEval. We hope NetEval could help developers track the progress and analyze the NetOps ability of their models. Citation Please cite our paper if you use our dataset. @misc{miao2023empirical, title={An Empirical Study of NetOps Capability of Pre-Trained Large Language Models}, author={Yukai Miao and Yu Bai and Li Chen and… See the full description on the dataset page: https://huggingface.co/datasets/NASP/neteval-exam.text-classification10K<n<100K6 likes48 downloads3y agoHugging Face08justicedao /ipfs_netherlands_laws IPFS Netherlands Laws Hugging Face target: justicedao/ipfs_netherlands_laws. This dataset packages Netherlands law records with deterministic IPFS Content IDs. Each row includes a cid and content_address; article rows also include the parent law_cid. This is a quality-audited catalog-backed Netherlands snapshot from official Dutch government sources. It is not the full Dutch legal corpus: the persistent catalog contains 42,956 discovered BWBR identifiers, of which 5,000 are… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws.tabulartext-retrieval100K<n<1M0 likes42 downloads3mo agoHugging Face09netop /CTBenchgated CTBench Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations 📄 Read the Paper 🤗 Explore the Dataset [!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset. CTBench CTBench is an agentic benchmark for evaluating AI agents in realistic telecom Network Operations and Maintenance… See the full description on the dataset page: https://huggingface.co/datasets/netop/CTBench.question-answering2 likes41 downloads2mo agoHugging Face10leeroy-jankins /DoD-Instruction-8010-01-Information-Network-Transport 🌐 DoD Information Network Transport Maintainer: Terry Eppler Owner: US Federal Government Source: DoD Instruction 8010.01 Dataset Size: question-answer records Source Effective Date: September 10, 2018 Source Organization: Office of the DoD Chief Information Officer Source Ownership: United States Department of Defense 📋 Overview Dataset Summary The DoD Information Network Transport Question-Answer Dataset contains document-grounded… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8010-01-Information-Network-Transport.documentquestion-answering0 likes34 downloads2mo agoHugging Face11Nethmi14 /LKValuesgated LKValues LKValues is a survey-grounded Sinhala–English resource suite for studying the alignment of large language models with Sri Lankan societal values. It contains two complementary resources: LKvaluesIT — an instruction-tuning dataset for training models to generate value-grounded explanations. LKvaluesBench — an evaluation benchmark for testing controlled value-sensitive judgment. LKValues is based on 40 societal values retained through a trilingual Sinhala–Tamil–English… See the full description on the dataset page: https://huggingface.co/datasets/Nethmi14/LKValues.text-generation100K<n<1M0 likes32 downloads16d agoHugging Face12K-Net-Labs /ru-instruct-KAN-logic-v1 Russian Instruct KAN-Logic Dataset (v1) Overview ru-instruct-KAN-logic-v1 — это специализированный набор данных для instruction tuning (дообучения) языковых моделей на русском языке. Основной фокус датасета — сложные логические рассуждения (Reasoning), математическое обоснование нейросетевых архитектур нового поколения (KAN - Kolmogorov-Arnold Networks) и теория распределенных вычислений. Датасет содержит синтетические и курируемые пары instruction - output… See the full description on the dataset page: https://huggingface.co/datasets/K-Net-Labs/ru-instruct-KAN-logic-v1.texttext-generationn<1K0 likes30 downloads8mo agoHugging Face13NetoAISolutions /NetBenchgated NetBench Dataset Dataset Overview The NetBench Dataset is a curated collection of expert-level question-answer pairs designed to benchmark the ability of large language models (LLMs) to achieve network subject matter expert (SME) intelligence across 20 critical telecommunications and network engineering categories. These categories include: Network Fundamentals & L2 Switching: Basic device access, Layer 2 concepts (VLANs, STP, LAG), L2 security, and interface… See the full description on the dataset page: https://huggingface.co/datasets/NetoAISolutions/NetBench.textquestion-answering1K<n<10K22 likes29 downloads10mo agoHugging Face14justicedao /netherlands-laws-nl-normalized Netherlands Laws (Dutch, Normalized) Hugging Face target: justicedao/netherlands-laws-nl-normalized. This package is a normalized version of the Netherlands laws scrape output. This is a capped Netherlands scrape, not the full Dutch corpus. The scrape used max_documents=100, parsed 151 law record(s), and discovered 626 unique official BWBR law document(s) before applying the cap. Documents failed: 0. This refresh includes parser coverage improvements for older/French heading… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/netherlands-laws-nl-normalized.tabulartext-retrieval1K<n<10K0 likes28 downloads3mo agoHugging Face15AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.texttext-retrieval100K<n<1M0 likes26 downloads7mo agoHugging Face16Mo7art /Stack2Graph_VD_vb.net Vb.Net StackOverflow Vector Dataset Summary This Hugging Face dataset repository contains the Vb.Net shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_vb.net.feature-extraction0 likes26 downloads2mo agoHugging Face17NetherlandsForensicInstitute /squad-nl-v2.0 SQuAD-NL v2.0 for Sentence Transformers The SQuAD-NL v2.0 dataset (on Hugging Face: GroNLP/squad-nl-v2.0), modified for use in Sentence Transformers as a dataset of type "Pair with Similarity Score". Score We added an extra column score to the original dataset. The value of score is 1.0 if the question has an answer in the context (no matter where), and 0.0 if there are no answers in the context. The allows the evaluation of embedding models that aim to pair queries… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/squad-nl-v2.0.textsentence-similarity100K<n<1M2 likes21 downloads2y agoHugging Face18AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-python RLVR-Env-Retrieval-Source-code-search-net-python RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.texttext-retrieval100K<n<1M0 likes21 downloads7mo agoHugging Face19Nettoov /Gpt-5.4-Xhigh-Reasoning-2000x Gpt-5.4-Xhigh-Reasoning-2000x A premium-quality reasoning dataset containing 2,007 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs. This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit thinking… See the full description on the dataset page: https://huggingface.co/datasets/Nettoov/Gpt-5.4-Xhigh-Reasoning-2000x.textquestion-answering1K<n<10K4 likes21 downloads6mo agoHugging Face20Mo7art /Stack2Graph_VD_.net .Net StackOverflow Vector Dataset Summary This Hugging Face dataset repository contains the .Net shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_.net.feature-extraction0 likes19 downloads2mo agoHugging Face21netop /TeleLogsAgentgated TeleLogsAgent A Benchmark for LLM Tool-Use in 5G Network Root Cause Analysis Developed by the NetOp Team, Huawei Paris Research Center Mohamed Sana · Nicola Piovesan · Antonio De Domenico · Fadhel Ayed 📄 Read the Paper 🤗 Explore the Dataset [!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset. TeleLogsAgent is a benchmark… See the full description on the dataset page: https://huggingface.co/datasets/netop/TeleLogsAgent.textquestion-answering1K<n<10K5 likes15 downloads7mo agoHugging Face22harshi321 /netflix-movies_showstexttext-classification1K<n<10K3 likes11 downloads2y agoHugging Face23creeperdatasets /qemu_networking Qemu Networking from Claude Haiku 4.5 A synthetic instruction-tuning dataset covering QEMU networking concepts, generated using Claude Haiku 4.5. Dataset Summary Total rows: 75 Topic: QEMU virtual networking Difficulty distribution: Easy, Intermediate, Advanced 278 unique tags across networking subtopics Splits train: 75 rows Columns id: Stable entry ID instruction: Instruction text for fine-tuning input: Original prompt/question output:… See the full description on the dataset page: https://huggingface.co/datasets/creeperdatasets/qemu_networking.texttext-generationn<1K0 likes11 downloads5mo agoHugging Face24Acktarius /for_Conceal-NetworkgatedDataset is meant to train OPEN_LLAMA_v2, a converted JSON version is also available Usage exemple with llama.cpp & open_llama_3b_v2 Finetune: finetune --model-base "C:\llama.cpp\models\open_llama_3b_v2_f32.gguf" --train-data "C:\llama.cpp\docs\conceal\conceal56_llama.txt" --lora-out lora-CCX_01.gguf --save-every 0 --threads 16 --ctx 256 --rope-freq-base 10000 --rope-freq-scale 1.0 --batch 1 --grad-acc 1 --adam-iter 256 --adam-alpha 0.00025 --lora-r… See the full description on the dataset page: https://huggingface.co/datasets/Acktarius/for_Conceal-Network.textquestion-answering1K<n<10K0 likes8 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.