CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fahadhafeezofficial /cissp-llmbench CISSP-LLMBench tabulartext-generation10K<n<100K0 likes3.2k downloads3mo agoHugging Face02guerrerotook /CIMA-4.8-ADR CIMA Sección 4.8 — Reacciones Adversas Corpus de texto biomédico regulatorio en español compuesto por la sección 4.8 ("Reacciones adversas") de la totalidad de las fichas técnicas publicadas por la Agencia Española de Medicamentos y Productos Sanitarios (AEMPS) en su Centro de Información Online de Medicamentos (CIMA). Este recurso fue construido como base para el pre-entrenamiento adaptado al dominio (continued pre-training / domain-adaptive pre-training, DAPT) de modelos… See the full description on the dataset page: https://huggingface.co/datasets/guerrerotook/CIMA-4.8-ADR.textfill-mask10K<n<100K0 likes1.5k downloads4mo agoHugging Face03inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face04renzzyyy1028 /civil-code-phil Civilex — Philippine Legal RAG & SFT Dataset Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline. Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386). Dataset structure . ├── README.md ├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.textquestion-answering100K<n<1M1 likes550 downloads2d agoHugging Face05publicus-ai /cibench-experiments CIBench Experiments Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era. If a benchmark result cannot be replayed from its manifest alone, it did not happen. Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.texttext-retrieval1K<n<10K0 likes427 downloads5mo agoHugging Face06arbml /CIDAR Dataset Card for "CIDAR" 🌴CIDAR: Culturally Relevant Instruction Dataset For Arabic [ Paper - GitHub ] CIDAR contains 10,000 instructions and their output. The dataset was created by selecting around 9,109 samples from Alpagasus dataset then translating it to Arabic using ChatGPT. In addition, we append that with around 891 Arabic grammar instructions from the webiste Ask the teacher. All the 10,000 samples were reviewed by around 12 reviewers. 📚… See the full description on the dataset page: https://huggingface.co/datasets/arbml/CIDAR.texttext-generation10K<n<100K57 likes392 downloads1y agoHugging Face07CinderD /wildtrace WildTrace strict481 WildTrace is a source-internal long-context multi-hop reasoning benchmark built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are mined in situ from long source documents before questions are written. The strict481 release contains 481 locked tasks over 214 public long-form sources, with full-document, evidence-withheld evaluation. The model under test receives only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.textquestion-answeringn<1K1 likes322 downloads1mo agoHugging Face08cilabuniba /wikifragments WikiFragments WikiFragments is a multimodal dataset built from Wikipedia (en), consisting of cleaned textual paragraphs paired with related images (infobox and thumbnail) from the same page. Each pair forms a multimodal fragment, which serves as an atomic knowledge unit ideal for information retrieval and multimodal research. Example of a rendered fragment with multiple images and captions. Fragment with only text and no associated images. [!NOTE]The images above were generated… See the full description on the dataset page: https://huggingface.co/datasets/cilabuniba/wikifragments.imagetext-generation10M<n<100M3 likes311 downloads7mo agoHugging Face09Jonaszky123 /L-CiteEval L-CITEEVAL: DO LONG-CONTEXT MODELS TRULY LEVERAGE CONTEXT FOR RESPONDING? Paper   Github   Zhihu Benchmark Quickview L-CiteEval is a multi-task long-context understanding with citation benchmark, covering 5 task categories, including single-document question answering, multi-document question answering, summarization, dialogue understanding, and synthetic tasks, encompassing 11 different long-context tasks. The context lengths for these tasks range from 8K to 48K.… See the full description on the dataset page: https://huggingface.co/datasets/Jonaszky123/L-CiteEval.tabularquestion-answering1K<n<10K3 likes297 downloads2y agoHugging Face10ilsp /greek_civics_qa Dataset Card for Greek Civics QA The Greek Civics QA dataset is a set of 407 question and answer pairs related to Greek highschool courses in civics (Κοινωνική και Πολιτική Αγωγή). The dataset was created by Nikoletta Tsoukala during her 2023 internship at the Institute for Language and Speech Processing/Athena RC, supervised by ILSP Research Associate Vassilis Papavassileiou. The dataset creation process involved mining questions and answers from two civics textbooks used in the… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/greek_civics_qa.texttext-generationn<1K6 likes270 downloads2y agoHugging Face11AlexWortega /llm-cipher-reasoning llm-cipher-reasoning — data, eval results and full research ledger Everything except the weights from a research run asking: can an LLM be trained to reason in a more compact "language" than English, and does that actually save tokens? Two linked lines of work on Qwen/Qwen3-4B-Instruct-2507: Cipher invention / cross-model communication — cold-decoding tests, negotiated cipher collusion between model pairs, a cipher-hardening arms race, and GEPA prompt optimization to get a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/llm-cipher-reasoning.texttext-generation10K<n<100K0 likes221 downloads22d agoHugging Face12Taylor658 /photonic-integrated-circuit-yield 🏭 Photonic Integrated Circuit Yield Dataset 📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing. ⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.texttext-generation100K<n<1M4 likes188 downloads12d agoHugging Face13Lots-of-LoRAs /task1721_civil_comments_obscenity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1721_civil_comments_obscenity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1721_civil_comments_obscenity_classification.texttext-generationn<1K1 likes187 downloads2y agoHugging Face14cis-lmu /GlotStoryBook Dataset Description Story Books for 180 ISO-639-3 codes. The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset. This dataset consists of 2 subsets: default, which consists of 4 publishers: asp: African Storybook pb: Pratham Books lcb: Little Cree Books lida: LIDA Stories nalibali, which comes from Nal'ibali stories. Usage (HF Loader) default: from datasets import load_dataset dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.texttranslation10K<n<100K9 likes186 downloads2d agoHugging Face15Lots-of-LoRAs /task1723_civil_comments_sexuallyexplicit_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1723_civil_comments_sexuallyexplicit_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1723_civil_comments_sexuallyexplicit_classification.texttext-generationn<1K0 likes182 downloads2y agoHugging Face16nvidia /Nemotron-RL-Instruction-Following-Citation-Formatting-v1 Dataset Description: Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations. This dataset is ready for commercial/non-commercial uses. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: April 10, 2026 Last Modified on: April 10, 2026 Version: Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.texttext-generation1K<n<10K2 likes179 downloads4mo agoHugging Face17opnsrcntrbtrian /sebi-circulars SEBI Circulars Dataset A comprehensive, structured dataset of Indian Securities and Exchange Board (SEBI) regulatory circulars, public-domain government works compiled and annotated for AI/ML research. Date: 2026-08-14 Snapshot Version: v2026.08 Corpus: 728 circulars (2010–2026) Dataset Configurations Config Rows Schema Purpose corpus 728 Full circular + metadata Flagship: regulatory text, lineage, effective dates chunks 78,585 Section-aware retrieval… See the full description on the dataset page: https://huggingface.co/datasets/opnsrcntrbtrian/sebi-circulars.texttext-retrievaln<1K1 likes177 downloads1mo agoHugging Face18Lots-of-LoRAs /task1720_civil_comments_toxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.texttext-generationn<1K0 likes167 downloads2y agoHugging Face19cilabuniba /wikifragments-visual-arts-embeds WikiFragments - Visual Arts Pages with Fragments (WikiFragmentsVA) WikiFragmentsVA is a domain-specific multimodal dataset focused on the visual arts, derived from Wikipedia (en). It consists of textual paragraphs paired with related images (infoboxes and thumbnails), rendered as unified visual fragments. This dataset extends the base WikiFragments project by providing pre-rendered fragment images and multi-vector embeddings obtained via ColQwen2 v1.0, including optimized pooled… See the full description on the dataset page: https://huggingface.co/datasets/cilabuniba/wikifragments-visual-arts-embeds.imagetext-generation1M<n<10M0 likes141 downloads7mo agoHugging Face20Lots-of-LoRAs /task1724_civil_comments_insult_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1724_civil_comments_insult_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1724_civil_comments_insult_classification.texttext-generation1K<n<10K0 likes133 downloads2y agoHugging Face21Lots-of-LoRAs /task568_circa_question_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task568_circa_question_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task568_circa_question_generation.texttext-generation1K<n<10K0 likes130 downloads2y agoHugging Face22h0ney-badger /civic-records-distill civic-records-distill Training data for a local model that helps a private citizen use public-records law: draft requests that are hard to stall, turn an angry draft into a letter an official has to engage with, look things up instead of inventing them, and escalate correctly when stonewalled. Grounded in Florida (ch. 119 Public Records Act, ch. 286 Sunshine Law, and the ALPR-specific s. 316.0777) and Texas (ch. 552 Public Information Act, ch. 551 Open Meetings Act). Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/h0ney-badger/civic-records-distill.texttext-generation1K<n<10K0 likes120 downloads21d agoHugging Face23CinderD /TeachArena TeachArena TeachArena is a benchmark for evaluating AI tutoring agents across the full teaching decision chain — from moment-to-moment tutoring dialogue, to pedagogical judgment on packaged evidence, to multi-step teaching workflows grounded in a learning-management system. It contains 354 tasks organized into three stages, a mock LMS environment database, the agent policy documents, and the full scoring logic. Why three stages A capable teaching agent must both… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/TeachArena.texttext-generationn<1K1 likes114 downloads1mo agoHugging Face24Lots-of-LoRAs /task1722_civil_comments_threat_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1722_civil_comments_threat_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1722_civil_comments_threat_classification.texttext-generationn<1K0 likes99 downloads2y agoHugging Face25eyuansu71 /logdx-ci LogDx-CI A benchmark for CI log reduction tools (RTK, grep, tail, hybrid routers, LLM-summary) — do they preserve enough evidence for LLM root-cause diagnosis? Homepage: https://logdx-bench.github.io/ Code & evaluator: https://github.com/eyuansu62/LogDx Headline report: reports/e10_v2_generalization_partial.md Release notes: RELEASE_NOTES.md (latest: RELEASE_NOTES_v1_2.md) Current release: v1.2 License: CC-BY-4.0 (data, this repo); Apache-2.0 (code, GH repo) Two ways to… See the full description on the dataset page: https://huggingface.co/datasets/eyuansu71/logdx-ci.tabulartext-classificationn<1K0 likes94 downloads4mo agoHugging Face26Lots-of-LoRAs /task565_circa_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task565_circa_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task565_circa_answer_generation.texttext-generation1K<n<10K0 likes91 downloads2y agoHugging Face27Auroraventures /cipher-awwwards-sft25 Cipher — Awwwards SFT 2.5 + Real v1 🦑 The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories. Two ways this dataset is used As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.tabulartext-generation1K<n<10K0 likes86 downloads5mo agoHugging Face28Lots-of-LoRAs /task566_circa_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task566_circa_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task566_circa_classification.texttext-generation1K<n<10K0 likes85 downloads2y agoHugging Face29ars22 /circle-packing-insight-loop Circle-Packing Insight-Exploration Loop Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the 21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii = 2.3658321334167627). Each round, 16 solvers propose a program + written explanation; every program is scored; a proposer then mines all 16 attempts into an evolving insight document that conditions the next round. Run: 16 solvers x 8 rounds. Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.tabulartext-generationn<1K0 likes83 downloads2mo agoHugging Face30louisbrulenaudet /code-procedures-civiles-execution Code des procédures civiles d'exécution, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-procedures-civiles-execution.tabulartext-generationn<1K0 likes77 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.