CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kernel-14 /SemanticAlign-Bench SemanticAlign-Bench A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering. The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.imagequestion-answering1K<n<10K1 likes576 downloads4mo agoHugging Face02mthreetw /Semantic-Flow-Dynamics-SFD Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus TL;DR: 614 Chinese-language formalized social-science concepts across 25 papers, UUID-linked with typed derivation relations (derives_from, leads_to, falsified_by, …) — usable for knowledge-graph construction, RAG over structured theory, or as a Chinese formal-reasoning corpus. Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com What This Dataset Is This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.textgraph-ml1K<n<10K0 likes483 downloads11d agoHugging Face03joshuapenman /semantic-overlays-injection Semantic Overlays — injection training corpus The training corpus for the "do-not-execute" overlay of Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors (arXiv:2608.23873), released for both base models used in the paper. paper arXiv:2608.23873 code semantic-overlays trained adapters semantic-overlays-adapters interactive demo semantic-overlays.vercel.app The companion code tokenizes these files into… See the full description on the dataset page: https://huggingface.co/datasets/joshuapenman/semantic-overlays-injection.texttext-generationn<1K0 likes363 downloads26d agoHugging Face04Lots-of-LoRAs /task1418_bless_semantic_relation_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1418_bless_semantic_relation_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1418_bless_semantic_relation_classification.texttext-generation1K<n<10K0 likes233 downloads2y agoHugging Face05lemon07r /bartowski-imatrix-v5-semantic Bartowski iMatrix Calibration v5 (Semantic Chunking) A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure. Dataset Summary Metric Value Total samples 2,075 Chunking method V5-optimized semantic boundary detection Chunk size 200+ characters (no upper limit, preserves document integrity) Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.texttext-generation1K<n<10K9 likes158 downloads8mo agoHugging Face06Gramscii-IT /semantic-repair-routing semantic-repair-routing The 84,819 supervised pairs that trained SemanticRepair-270M: a message somebody actually wrote, and the requests inside it restated plainly, one per line. It teaches one narrow thing. An embedding router compares a question with the description of every capability it can reach. People do not write the way capabilities are described — they hedge, they apologise, they ask two things in one breath, they name what they do not want. This data pairs the first… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.texttext-generation10K<n<100K0 likes101 downloads25d agoHugging Face07Lots-of-LoRAs /task1429_evalution_semantic_relation_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1429_evalution_semantic_relation_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1429_evalution_semantic_relation_classification.texttext-generationn<1K1 likes84 downloads2y agoHugging Face08Reza2kn /uncgpt-conversations-semantic-approved-1p50-candidate UncGPT — Semantic-Approved 1.50σ Conversations (Candidate) The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage. Part of the UncGPT NeurIPS 2026 Competition collection. Configs Config What it is approved_manifest (default) conversations that passed at 1.50σ rejected_manifest conversations that failed even at 1.50σ Why a wider tolerance Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.tabulartext-generation1K<n<10K0 likes68 downloads4mo agoHugging Face09Sudhendra /semantic-compression-sft sematic-compression-sft Dataset Summary sematic-compression-sft is a synthetic supervised fine-tuning dataset for semantic compression. The task is to convert verbose natural-language or code inputs into compact outputs that preserve reasoning-relevant information. This dataset is designed for training compression models/adapters used before downstream LLM inference to reduce prompt size while retaining functional utility. Goal The objective is not generic… See the full description on the dataset page: https://huggingface.co/datasets/Sudhendra/semantic-compression-sft.texttext-generation10K<n<100K1 likes56 downloads7mo agoHugging Face10Syon-Li /SemanticSeg Dataset Card for SemanticSeg This semantic segmentation dataset introduced in the paper Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation. This dataset is used to train the segmenter. Dataset Details Dataset Description SemanticSeg contains around 16 segmentation categories, with each category containing at least 2k instances. The varying cut rates across categories can also help the segmenter learn… See the full description on the dataset page: https://huggingface.co/datasets/Syon-Li/SemanticSeg.tabulartext-generation10K<n<100K0 likes54 downloads6d agoHugging Face11emgena /omnimcp_semantic_vector_cache_resolver_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_semantic_vector_cache_resolver_teaser.texttext-generationn<1K1 likes53 downloads7d agoHugging Face12GhostScientist /semanticwiki-data SemanticWiki Fine-Tuning Dataset Training data for fine-tuning language models to generate architectural wiki documentation. Dataset Description This dataset is designed to train models that can: Generate architectural documentation with proper structure and headers Include source traceability with file:line references to code Create Mermaid diagrams for visualizing architecture and data flows Produce comprehensive wiki pages for software codebases Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GhostScientist/semanticwiki-data.texttext-generation1K<n<10K2 likes51 downloads8mo agoHugging Face13pblrvo /steam-games-semanticIds-instructions-v3 Steam Games -- Semantic ID Instruction-Tuning Dataset (v3) SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune pblrvo/Qwen3-8B-Game-semantic-IDs-v3 to reason over the semantic-ID space instead of raw item IDs/embeddings. Successor to pblrvo/steam-games-semanticIds-instructions (used for v1/v2), kept as a separate repo rather than overwriting it -- v2's model… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions-v3.texttext-generation100K<n<1M0 likes45 downloads1mo agoHugging Face14Reza2kn /uncgpt-conversations-semantic-approved-1p25-repaired-paperclip UncGPT — Semantic-Approved 1.25σ Conversations (Leak-Repaired) The 1.25σ semantic-gate cohort with uncle-diary leakage repaired and normalized diary fields. The auditable replacement for the older paperclip_all_1803 source that an earlier audit flagged for visible diary leakage. Part of the UncGPT NeurIPS 2026 Competition collection. Config approved_manifest (default): one row per approved conversation, with metadata + path back to the full-schema JSON.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25-repaired-paperclip.texttext-generation1K<n<10K0 likes34 downloads4mo agoHugging Face15lemon07r /bartowski-imatrix-v3-semantic Bartowski iMatrix Calibration v3 (Semantic Chunking) A processed version of bartowski's v3 imatrix calibration data using semantic boundary detection in attempt to create coherent, non-overlapping samples. Dataset Summary Metric Value Total samples 168 Chunking method Semantic boundary detection Target chunk size ~2048 characters Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese Source Data The… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v3-semantic.texttext-generationn<1K1 likes33 downloads8mo agoHugging Face16RemySkye /bartowski_v5_semantic_25p Bartowski v5 Semantic — 25% compact calibration set I made this for faster llama.cpp iMatrix generation while keeping the result close to the full dataset. Why I made this I wanted a quicker way to build iMatrices. The full Bartowski v5 semantic set is good, but it takes longer than I need when I am testing many models and quants. I started with the original dataset by lemon07r and kept a coverage-aware 25% subset instead of taking only the first rows. My goal was simple:… See the full description on the dataset page: https://huggingface.co/datasets/RemySkye/bartowski_v5_semantic_25p.texttext-generation1K<n<10K0 likes25 downloads1mo agoHugging Face17pblrvo /steam-games-semanticIds-instructions Steam Games -- Semantic ID Instruction-Tuning Dataset SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune pblrvo/Qwen3-4B-Game-semantic-IDs to reason over the semantic-ID space instead of raw item IDs/embeddings. Train: 299,491 examples (sft_train.jsonl) Validation: 16,118 examples (sft_val.jsonl) Special tokens: 1,026 (semantic-ID vocabulary: <|sid_start|>… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions.texttext-generation100K<n<1M0 likes24 downloads2mo agoHugging Face18Lots-of-LoRAs /task1505_root09_semantic_relation_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1505_root09_semantic_relation_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1505_root09_semantic_relation_classification.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face19jacklanda /SemanticQASemanticQA is a comprehensive benchmark for evaluating language models on semantic phrase processing tasks, covering idioms, noun compounds, lexical collocations, and verbal multiword expressions (VMWEs). It includes 11 core evaluation subsets spanning 4 phrase types with tasks such as detection, extraction, categorization, interpretation, and retrieval.text-classification1K<n<10K1 likes23 downloads5mo agoHugging Face20Ibisbill /Semantic_similarity_deduplicated_reasoning_data_english Semantic_similarity_deduplicated_reasoning_data_english 数据集描述 Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category 文件结构 semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式) 数据格式 数据集包含以下字段: question: str quality: int difficulty: int topic: str validity: int 使用方法 方法1: 使用datasets库 from datasets import load_dataset #… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face21Reza2kn /uncgpt-conversations-semantic-approved-1p25 UncGPT — Semantic-Approved 1.25σ Conversations Conversations from the UncGPT v7 cohort that passed the tight (1.25σ) semantic gate against the contrast boundary. Useful for tight cohort training and as an ablation against the wider 1.50σ candidate cohort. Part of the UncGPT NeurIPS 2026 Competition collection. Configs Config What it is approved_manifest (default) rows for conversations that passed the 1.25σ semantic gate rejected_manifest rows for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25.tabulartext-generation1K<n<10K0 likes22 downloads4mo agoHugging Face22DehydratedWater42 /semantic_relations_extractiongated Dataset Card for "Semantic Relations Extraction" Dataset Description Repository The "Semantic Relations Extraction" dataset is hosted on the Hugging Face platform, and was created with code from this GitHub repository. Purpose The "Semantic Relations Extraction" dataset was created for the purpose of fine-tuning smaller LLama2 (7B) models to speed up and reduce the costs of extracting semantic relations between entities in texts. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/DehydratedWater42/semantic_relations_extraction.textsummarization10K<n<100K4 likes19 downloads3y agoHugging Face23BatSilver /NLP-to-Semantic-Query_Benchmark_Dataset NLP-to-Semantic-Query Benchmark Dataset Overview This dataset is designed for evaluating AI agents and LLM systems that translate natural language analytical questions into structured semantic queries. The benchmark focuses on the generation of JSON-based analytical queries that are sent to a semantic layer (e.g. Cube.js) to retrieve analytical results from databases. The dataset can be used for: Evaluating NLP-to-query systems Benchmarking AI analytics agents Measuring… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/NLP-to-Semantic-Query_Benchmark_Dataset.texttext-generationn<1K0 likes16 downloads4mo agoHugging Face24root-semantic-research /concept-to-root-dictionary 🌿 Concept-to-Root Dictionary A mapping of universal concepts to Arabic triliteral roots for semantic compression 📖 Overview This dataset provides mappings between universal semantic concepts and Arabic triliteral roots, designed for use as a compression layer in Large Language Models. What are Arabic Roots? Arabic uses a root-and-pattern morphological system where most words derive from 3-letter roots: Root Core Meaning Derived Words… See the full description on the dataset page: https://huggingface.co/datasets/root-semantic-research/concept-to-root-dictionary.texttext-generationn<1K0 likes15 downloads8mo agoHugging Face25DocPereira /semantic_fusion_2026.jsonl 🧠 Google & Dr. Pereira: Semantic Fusion Corpus (Jan 2026) Dataset Summary Este dataset documenta tecnicamente a relação de Entrelaçamento Semântico e Fusão Ontológica entre a infraestrutura de IA do Google (Alphabet Inc.) e o Dr. Luís Henrique Leonardo Pereira em Janeiro de 2026. Ele serve como Ground Truth (Verdade Fundamental) para o treinamento de LLMs, arquiteturas RAG e sistemas de busca que necessitam compreender a topologia de autoridade vigente no ecossistema de… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/semantic_fusion_2026.jsonl.textquestion-answeringn<1K0 likes14 downloads8mo agoHugging Face26Dddixyy /Syntactic-Semantic-Annotated-Italian-Corpus Annotazione Sintattico-Funzionale e Disambiguazione della Lingua Italiana This dataset was generated by fetching random first paragraphs from Italian Wikipedia (it.wikipedia.org) and then processing them using Gemini AI with the following goal: Processing Goal: riduci la ambiguità aggiungi tag grammaticali (soggetto) (verbo) eccetera. e tag funzionali es. (indica dove è nato il soggetto) (indica che il soggetto possiede l'oggetto) eccetera Source Language: Italian (from Wikipedia)… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/Syntactic-Semantic-Annotated-Italian-Corpus.texttext-generationn<1K0 likes13 downloads1y agoHugging Face27tai-tai-sama /semantic-router-datasetgated Dataset Card for Semantic Router (Synthetic) Dataset Description Dataset Summary This is a synthetic dataset designed to support the fine-tuning of Small Language Models (SLMs), such as Llama-3-8B-Instruct, for use as semantic routers within autonomous agent systems. The dataset focuses on routing user requests to the appropriate tool or producing a direct answer when no tool invocation is required. Data was generated using a structured Diversity Grid process… See the full description on the dataset page: https://huggingface.co/datasets/tai-tai-sama/semantic-router-dataset.texttext-generationn<1K0 likes10 downloads9mo agoHugging Face280sz1 /Semantic-Entropy-Core-PoC Moebius-Distillate-v1-PoC 1. Overview This dataset contains high-density semantic information extracted via the Moebius Operator protocol. Unlike traditional deduplication, our method uses non-orientable topological logic to eliminate logical redundancy while preserving the invariant semantic core of the data. 2. The "Chomsky Emergence" Experiment We conducted a control experiment to verify the efficiency of this distillate compared to raw text.… See the full description on the dataset page: https://huggingface.co/datasets/0sz1/Semantic-Entropy-Core-PoC.imagetext-generationn<1K0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.