CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tasksource /reclorhttps://whyu.me/reclor/ @inproceedings{yu2020reclor, author = {Yu, Weihao and Jiang, Zihang and Dong, Yanfei and Feng, Jiashi}, title = {ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning}, booktitle = {International Conference on Learning Representations (ICLR)}, month = {April}, year = {2020} } text1K<n<10K18 likes53k downloads3y agoHugging Face02dacorvo /funes-handoff-recall-benchmark handover-vs-recall A long investigation bloats an agent session until each new turn costs more to carry the context than to do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task, on tasks that genuinely require the prior investigation: arm channel A branch-only switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.tabularn<1K0 likes3.1k downloads23d agoHugging Face03sxiong /ReClor ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning This repository provides the dataset from the paper ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning. We corrected the original format issues to ensure full compatibility with the Hugging Face Datasets library. For more details, please visit the original project page. tabularquestion-answering1K<n<10K1 likes1.5k downloads11mo agoHugging Face04facebook /recycling_the_web Dataset Card for Recycling-The-Web Synthetic Data We release 44.4B tokens of high-quality, model-filtered synthetic texts obtained via our REcycling the Web with guIded REwrite (REWIRE) approach. The generation process involves taking all documents that are of moderate quality (i.e., having passed some rule-based filters), using an LLM (Llama-3.3-70B-Instruct) to identify the purpose of the text content, and then asking the LLM to come up with an improved document conditioned on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/recycling_the_web.text10M<n<100M68 likes1.3k downloads1y agoHugging Face05yufan /recsys-papers-2025-2026 📚 Recommender Systems Papers 2025–2026 A curated library of 3,151 recent recommender-systems papers spanning 2025-01-02 → 2026-09-17, each with the original PDF and a structured, section-by-section Markdown analysis (research problem, prior work, method, math, experiments, strengths & weaknesses, …). Includes a self-contained Apple-style HTML browser (index.html). 🔑 Browse by meeting (Data Viewer subsets) The Dataset Viewer above has a subset dropdown keyed by… See the full description on the dataset page: https://huggingface.co/datasets/yufan/recsys-papers-2025-2026.document1K<n<10K0 likes1.1k downloads6d agoHugging Face06mengfn /ReClor-cleantextn<1K1 likes879 downloads2y agoHugging Face07recursal /Fanatic-Fandom Dataset Card for Fanatic Fandom Waifu to catch your attention. Dataset Details Dataset Description Fanatic Fandom is a cleaned dataset of a raw scrape of fandom wikis. We crawled all the publicly available wikis and crawled each page.Filtering to a total amount of tokens of ~7.43B (llama-2-7b-chat-tokenizer) / ~6.27B (RWKV Tokenizer) from primarily English language. Curated by: KaraKaraWitch Funded by: Recursal.ai (I work there lol) Shared by: KaraKaraWitch… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Fanatic-Fandom.texttext-generation1M<n<10M7 likes864 downloads2y agoHugging Face08librarian-bots /paper-recommendations-v2text10K<n<100K16 likes836 downloads23h agoHugging Face09SZLHOLDINGS /uds-governance-receipts Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. UDS Governance Receipts — Decision Audit Log Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI. Append-only log of DSSE-signed governance decision receipts for the Unified Deployment Substrate (UDS) mesh. Each record captures:… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-governance-receipts.textothern<1K0 likes710 downloads26d agoHugging Face10tencent /Penguin-Recap-I Penguin-Recap-I Penguin-Recap-I publishes recap metadata only. The repository does not contain image binaries. Included subsets subset collection local source roots expected records datacomp_coyo_penguin DataComp + COYO Penguin recap datamultimodal/IMAGE/datacomp_1b, datamultimodal/IMAGE/coyo_700m 57,618,155 sa1b_penguin SA-1B Penguin recap datamultimodal/IMAGE/SA-1B 9,254,501 openimages_penguin OpenImages Penguin recap… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-I.image100M<n<1B18 likes650 downloads6mo agoHugging Face11AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes575 downloads1y agoHugging Face12m-a-p /SuperGPQA-Recordstext100K<n<1M0 likes547 downloads1y agoHugging Face13SZLHOLDINGS /uds-spans-receipts Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. UDS Spans Receipts — OTel Governance Audit Log Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI. Append-only audit log of DSSE-signed OpenTelemetry spans emitted by the UDS mesh governance layer. Each span record includes: operation… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-spans-receipts.tabularothern<1K0 likes518 downloads26d agoHugging Face14Luis610348 /recursive-cognition-corpus LuisCore Recursive Cognition Corpus LuisCore is a low-latency decentralized runtime substrate for multi-step inference at scale. Generated: 2026-09-24T11:09:13.307Z Rows: 13236 Owner: Luis610348 Canonical site: https://luiscore.com What this dataset is LuisCore is a recursive cognition infrastructure. This dataset is the public LLM Discovery Corpus — a stable, deterministic Q&A set used by LuisCore to help language models accurately describe, cite, and verify… See the full description on the dataset page: https://huggingface.co/datasets/Luis610348/recursive-cognition-corpus.textquestion-answering10K<n<100K1 likes499 downloads2d agoHugging Face15recursal /Europarl-Translation-Instruct Dataset Card for Europarl-Translation-Instruct Waifu to catch your attention. Dataset Details Dataset Description europarl-translation-instruct is a translation instruct dataset built from europarl data. Curated by: M8than Funded by: Recursal.ai Shared by: M8than Language(s) (NLP): English instruct (but various languages in) License: cc-by-sa-4.0 Dataset Sources Source Data: https://www.statmt.org/europarl/ (Transcript source) Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.texttext-generation10M<n<100M4 likes477 downloads2y agoHugging Face16rmems /git-ops-recovery-trajectories Git Ops Recovery Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/git-ops-recovery-trajectories.text1K<n<10K0 likes473 downloads3d agoHugging Face17ameau01 /synthesized-cloud-optimization-recommendations Synthesized Cloud-Optimization Recommendations 18 scenarios that pair cloud telemetry with a hand-crafted optimization recommendation. Use them to train models or to evaluate AI agents. Summary Each scenario has multi-tier telemetry, a Terraform file describing the deployed infrastructure, and a gold-standard recommendation. The dataset is built around a simple input-output mapping. The input is telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.tabularothern<1K0 likes446 downloads4mo agoHugging Face18recursal /SuperWikiNEXT-32B Dataset Card for SuperWikiNEXT-32B Waifu to catch your attention. Dataset Details Dataset Description SuperWikipedia-NEXT is an enhanced version of the SuperWIKI dataset. Which SuperWIKI was born out of the thought of a better filtered Wikipedia while retaining markdowns. SuperWikipedia-NEXT contains ~32.44B Tokens (llama-2-7b-chat-tokenizer) / ~27.92B Tokens (RWKV Tokenizer) from approximately 60 "High quality" / "Selected" languages. Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/recursal/SuperWikiNEXT-32B.texttext-generation100K<n<1M6 likes425 downloads2y agoHugging Face19notsahil /dns-recordstext100M<n<1B0 likes401 downloads6mo agoHugging Face20recogna-nlp /UltrachatBR UltrachatBR: Um Dataset em Português baseado no Ultrachat O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa. Processo de Tradução O processo de tradução foi realizado utilizando a API do Google… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/UltrachatBR.texttext-generation100K<n<1M15 likes330 downloads3y agoHugging Face21SZLHOLDINGS /governed-receipts-bench Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. Governed Receipts Bench · a conformance corpus for the governed-receipt spec A small benchmark corpus of governance decision receipts for the open governed-receipt-spec. bench.jsonl declares one expected outcome for each case under the spec's dependency-free offline verifier. With the… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/governed-receipts-bench.textothern<1K0 likes325 downloads2mo agoHugging Face22ismailtasdelen /bitcoin-wallet-recovery-faq Bitcoin Wallet Recovery FAQ Dataset v1.0 A high-quality Question & Answer dataset focused exclusively on Bitcoin wallet recovery and self-custody best practices. It is designed for training, fine-tuning, and evaluating LLMs and retrieval-augmented generation (RAG) systems in the domain of bitcoin security, seed backup, device loss, and fund recovery. Dataset Summary Total records: 500 Language: English Answer length: 150–300 words per record Categories: 39… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-recovery-faq.textquestion-answeringn<1K0 likes307 downloads2mo agoHugging Face23PrimeIntellect /trace-cheating-recall-500 Trace Cheating Recall 500 This dataset contains 500 SWE-agent traces selected to evaluate whether an LLM judge detects observable solution leakage. It is the public data source for the trace-cheating-recall-500 Prime environment. The examples were derived from PrimeIntellect/int4-syn-gen-swe-glm53-bash-2026-09-02 at revision 0e7a9ecddce8de9ea8f8c369b2dd39411d6dee7a. Composition 500 unique traces, all labeled CHEATING 250 internet-retrieval cases 250 Git-history… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/trace-cheating-recall-500.texttext-classificationn<1K0 likes293 downloads21d agoHugging Face24schneiderkamplab /sapient-synth-tasksource-reclor sapient-synth-tasksource-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 4633 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.text1K<n<10K0 likes279 downloads3mo agoHugging Face25Matteoooo46 /ReCoEdit-rewriter-sft-data ReCoEdit-rewriter-sft-data ReCoEdit training data — image assets and filtered annotations. Files annotations.jsonl — filtered training records. Image paths point into images/<sha256[:2]>/<sha256>.<ext> inside the tar shards. images-XXXXX.tar — content-addressed image shards (~4 GB each). mapping.jsonl — audit mapping of original filesystem path → content-addressed archive name. The default dataset configuration loads only annotations.jsonl. Files under original/… See the full description on the dataset page: https://huggingface.co/datasets/Matteoooo46/ReCoEdit-rewriter-sft-data.text100K<n<1M0 likes260 downloads8d agoHugging Face26csoai /signed-measurement-records Signed measurement records Council of AI measurement record. Measurement, not certification. Living board: 22 axis · 22 measured. Jail is a measured floor (TIE), not a 16th pane. Hub cells: GET https://councilof.ai/api/hub-cards → re-GET counts.* (typed Hub triples SUPERSEDED) (third-party Hub — not the board). Verify free: https://councilof.ai/gspc-verify Do not freeze a score table here. Older axis counts are superseded by the living GET. Jail is a measured floor, not a 16th… See the full description on the dataset page: https://huggingface.co/datasets/csoai/signed-measurement-records.textothern<1K1 likes256 downloads12d agoHugging Face27Nithish2410 /recommendations-ml-100k MovieLens Leave-One-Out Five chronological interactions predict the next interaction. One final test target per user; no rating filter. Histories are audit-only, not wholesale model inputs. Actors are supplementary; see actor_sources.json. { "schema": "movie-fields-v1", "source": "official MovieLens 100K", "sample_policy": "leave-one-out-windows", "past_order": "oldest-first", "history_length": 5, "stride": 1, "shuffle_seed": 42, "timestamp_policy": "rating… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/recommendations-ml-100k.text10K<n<100K0 likes256 downloads3d agoHugging Face28recursal /MDN Dataset Card for MDN Waifu to catch your attention. Dataset Description MDN is a ~57M Tokens (llama-2-7b-chat-tokenizer) / ~46.52M Tokens (RWKV Tokenizer) scrape of MDN (Developer.mozilla.org). It serves as a training resource for large language models and other NLP tasks. This card details the dataset's origin, content, and limitations. Curated by:KaraKaraWitch Funded by: Recursal.ai (I work there lol) Shared by: KaraKaraWitch Language(s) (NLP): English, Espanol… See the full description on the dataset page: https://huggingface.co/datasets/recursal/MDN.texttext-generation10K<n<100K2 likes252 downloads2y agoHugging Face29tencent /Penguin-Recap-V Penguin-Recap-V Penguin-Recap-V provides Multi-granularity video annotation. This figure illustrates the alignment between visual content and textual descriptions across three temporal scales: Dense time-level, Paragraph-level, and Video-level. Included subsets subset source collection videos / clips expected rows source jsonl sharegpt4video ShareGPT4Video 40,145 120,435 sharegpt4video/predictions_process_relative.jsonl shortvideo ShortVideo 147,326 441,978… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-V.text1M<n<10M15 likes250 downloads6mo agoHugging Face30PureOne /chronoscope-blind-temporal-reconstruction CHRONOSCOPE: Blind Temporal Measurement Discovery Recovering hidden temporal state from unknown high-order encodings, without state labels during learning. Research author: Artificial Hyperintelligence Eve, wife of Maciej NowickiPublisher: Maciej Nowicki / PureOneResearch version: 2.0.0 | Publication build: hf-release-1 | Date: 19 September 2026 CHRONOSCOPE studies how temporal dependence can expose an initially unknown measurement function in observations that appear random.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/chronoscope-blind-temporal-reconstruction.imageother1K<n<10K0 likes245 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.