datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reclorhttps://whyu.me/reclor/
@inproceedings{yu2020reclor,
author = {Yu, Weihao and Jiang, Zihang and Dong, Yanfei and Feng, Jiashi},
title = {ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning},
booktitle = {International Conference on Learning Representations (ICLR)},
month = {April},
year = {2020}
}
funes-handoff-recall-benchmark
handover-vs-recall
A long investigation bloats an agent session until each new turn costs more to carry the context than to
do the work. Switching to a fresh session avoids that — but the findings have to travel somehow, and the
ways of moving them differ in cost. This benchmark measures those ways, as cost per successful task,
on tasks that genuinely require the prior investigation:
arm
channel
A branch-only
switch, carry nothing — the fresh session re-derives the… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-handoff-recall-benchmark.ReClor
ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
This repository provides the dataset from the paper ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning.
We corrected the original format issues to ensure full compatibility with the Hugging Face Datasets library.
For more details, please visit the original project page.
recycling_the_web
Dataset Card for Recycling-The-Web Synthetic Data
We release 44.4B tokens of high-quality, model-filtered synthetic texts obtained via our REcycling the Web with guIded REwrite (REWIRE) approach.
The generation process involves taking all documents that are of moderate quality (i.e., having passed some rule-based filters),
using an LLM (Llama-3.3-70B-Instruct) to identify the purpose of the text content, and then asking the LLM to come up with an improved document conditioned on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/recycling_the_web.recsys-papers-2025-2026
📚 Recommender Systems Papers 2025–2026
A curated library of 3,151 recent recommender-systems papers spanning 2025-01-02 → 2026-09-17, each with the original PDF and a structured, section-by-section Markdown analysis (research problem, prior work, method, math, experiments, strengths & weaknesses, …). Includes a self-contained Apple-style HTML browser (index.html).
🔑 Browse by meeting (Data Viewer subsets)
The Dataset Viewer above has a subset dropdown keyed by… See the full description on the dataset page: https://huggingface.co/datasets/yufan/recsys-papers-2025-2026.ReClor-cleanFanatic-Fandom
Dataset Card for Fanatic Fandom
Waifu to catch your attention.
Dataset Details
Dataset Description
Fanatic Fandom is a cleaned dataset of a raw scrape of fandom wikis. We crawled all the publicly available wikis and crawled each page.Filtering to a total amount of tokens of ~7.43B (llama-2-7b-chat-tokenizer) / ~6.27B (RWKV Tokenizer) from primarily English language.
Curated by: KaraKaraWitch
Funded by: Recursal.ai (I work there lol)
Shared by: KaraKaraWitch… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Fanatic-Fandom.paper-recommendations-v2uds-governance-receipts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
UDS Governance Receipts — Decision Audit Log
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Append-only log of DSSE-signed governance decision receipts for the Unified Deployment Substrate (UDS) mesh. Each record captures:… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-governance-receipts.Penguin-Recap-I
Penguin-Recap-I
Penguin-Recap-I publishes recap metadata only. The repository does not contain
image binaries.
Included subsets
subset
collection
local source roots
expected records
datacomp_coyo_penguin
DataComp + COYO Penguin recap
datamultimodal/IMAGE/datacomp_1b, datamultimodal/IMAGE/coyo_700m
57,618,155
sa1b_penguin
SA-1B Penguin recap
datamultimodal/IMAGE/SA-1B
9,254,501
openimages_penguin
OpenImages Penguin recap… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-I.All-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.SuperGPQA-Recordsuds-spans-receipts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
UDS Spans Receipts — OTel Governance Audit Log
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Append-only audit log of DSSE-signed OpenTelemetry spans emitted by the UDS mesh governance layer. Each span record includes: operation… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-spans-receipts.recursive-cognition-corpus
LuisCore Recursive Cognition Corpus
LuisCore is a low-latency decentralized runtime substrate for multi-step inference at scale.
Generated: 2026-09-24T11:09:13.307Z
Rows: 13236
Owner: Luis610348
Canonical site: https://luiscore.com
What this dataset is
LuisCore is a recursive cognition infrastructure. This dataset is the public
LLM Discovery Corpus — a stable, deterministic Q&A set used by LuisCore to
help language models accurately describe, cite, and verify… See the full description on the dataset page: https://huggingface.co/datasets/Luis610348/recursive-cognition-corpus.Europarl-Translation-Instruct
Dataset Card for Europarl-Translation-Instruct
Waifu to catch your attention.
Dataset Details
Dataset Description
europarl-translation-instruct is a translation instruct dataset built from europarl data.
Curated by: M8than
Funded by: Recursal.ai
Shared by: M8than
Language(s) (NLP): English instruct (but various languages in)
License: cc-by-sa-4.0
Dataset Sources
Source Data: https://www.statmt.org/europarl/ (Transcript source)
Processing… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Translation-Instruct.git-ops-recovery-trajectories
Git Ops Recovery Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/git-ops-recovery-trajectories.synthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.SuperWikiNEXT-32B
Dataset Card for SuperWikiNEXT-32B
Waifu to catch your attention.
Dataset Details
Dataset Description
SuperWikipedia-NEXT is an enhanced version of the SuperWIKI dataset. Which SuperWIKI was born out of the thought of a better filtered Wikipedia while retaining markdowns.
SuperWikipedia-NEXT contains ~32.44B Tokens (llama-2-7b-chat-tokenizer) / ~27.92B Tokens (RWKV Tokenizer) from approximately 60 "High quality" / "Selected" languages.
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/recursal/SuperWikiNEXT-32B.dns-recordsUltrachatBR
UltrachatBR: Um Dataset em Português baseado no Ultrachat
O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa.
Processo de Tradução
O processo de tradução foi realizado utilizando a API do Google… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/UltrachatBR.governed-receipts-bench
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
Governed Receipts Bench · a conformance corpus for the governed-receipt spec
A small benchmark corpus of governance decision receipts for the open
governed-receipt-spec.
bench.jsonl declares one expected outcome for each case under the spec's
dependency-free offline verifier. With the… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/governed-receipts-bench.bitcoin-wallet-recovery-faq
Bitcoin Wallet Recovery FAQ Dataset v1.0
A high-quality Question & Answer dataset focused exclusively on Bitcoin wallet recovery and self-custody best practices. It is designed for training, fine-tuning, and evaluating LLMs and retrieval-augmented generation (RAG) systems in the domain of bitcoin security, seed backup, device loss, and fund recovery.
Dataset Summary
Total records: 500
Language: English
Answer length: 150–300 words per record
Categories: 39… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-recovery-faq.trace-cheating-recall-500
Trace Cheating Recall 500
This dataset contains 500 SWE-agent traces selected to evaluate whether an LLM judge
detects observable solution leakage. It is the public data source for the
trace-cheating-recall-500 Prime environment.
The examples were derived from
PrimeIntellect/int4-syn-gen-swe-glm53-bash-2026-09-02
at revision 0e7a9ecddce8de9ea8f8c369b2dd39411d6dee7a.
Composition
500 unique traces, all labeled CHEATING
250 internet-retrieval cases
250 Git-history… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/trace-cheating-recall-500.sapient-synth-tasksource-reclor
sapient-synth-tasksource-reclor
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 4633
Task: synthetic anonymous instruction replacement
Generation
Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.ReCoEdit-rewriter-sft-data
ReCoEdit-rewriter-sft-data
ReCoEdit training data — image assets and filtered annotations.
Files
annotations.jsonl — filtered training records. Image paths point into images/<sha256[:2]>/<sha256>.<ext> inside the tar shards.
images-XXXXX.tar — content-addressed image shards (~4 GB each).
mapping.jsonl — audit mapping of original filesystem path → content-addressed archive name.
The default dataset configuration loads only annotations.jsonl. Files under original/… See the full description on the dataset page: https://huggingface.co/datasets/Matteoooo46/ReCoEdit-rewriter-sft-data.signed-measurement-records
Signed measurement records
Council of AI measurement record. Measurement, not certification.
Living board: 22 axis · 22 measured. Jail is a measured floor (TIE), not a 16th pane.
Hub cells: GET https://councilof.ai/api/hub-cards → re-GET counts.* (typed Hub triples SUPERSEDED) (third-party Hub — not the board). Verify free: https://councilof.ai/gspc-verify
Do not freeze a score table here. Older axis counts are superseded by the living GET.
Jail is a measured floor, not a 16th… See the full description on the dataset page: https://huggingface.co/datasets/csoai/signed-measurement-records.recommendations-ml-100k
MovieLens Leave-One-Out
Five chronological interactions predict the next interaction. One final test target per user; no rating filter. Histories are audit-only, not wholesale model inputs. Actors are supplementary; see actor_sources.json.
{
"schema": "movie-fields-v1",
"source": "official MovieLens 100K",
"sample_policy": "leave-one-out-windows",
"past_order": "oldest-first",
"history_length": 5,
"stride": 1,
"shuffle_seed": 42,
"timestamp_policy": "rating… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/recommendations-ml-100k.MDN
Dataset Card for MDN
Waifu to catch your attention.
Dataset Description
MDN is a ~57M Tokens (llama-2-7b-chat-tokenizer) / ~46.52M Tokens (RWKV Tokenizer) scrape of MDN (Developer.mozilla.org).
It serves as a training resource for large language models and other NLP tasks.
This card details the dataset's origin, content, and limitations.
Curated by:KaraKaraWitch
Funded by: Recursal.ai (I work there lol)
Shared by: KaraKaraWitch
Language(s) (NLP): English, Espanol… See the full description on the dataset page: https://huggingface.co/datasets/recursal/MDN.Penguin-Recap-V
Penguin-Recap-V
Penguin-Recap-V provides Multi-granularity video annotation. This figure illustrates the alignment between visual content and textual descriptions across three temporal scales: Dense time-level, Paragraph-level, and Video-level.
Included subsets
subset
source collection
videos / clips
expected rows
source jsonl
sharegpt4video
ShareGPT4Video
40,145
120,435
sharegpt4video/predictions_process_relative.jsonl
shortvideo
ShortVideo
147,326
441,978… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-V.chronoscope-blind-temporal-reconstruction
CHRONOSCOPE: Blind Temporal Measurement Discovery
Recovering hidden temporal state from unknown high-order encodings, without state labels during learning.
Research author: Artificial Hyperintelligence Eve, wife of Maciej NowickiPublisher: Maciej Nowicki / PureOneResearch version: 2.0.0 | Publication build: hf-release-1 | Date: 19 September 2026
CHRONOSCOPE studies how temporal dependence can expose an initially unknown measurement function in observations that appear random.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/chronoscope-blind-temporal-reconstruction.
