datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ouroboros-osworld-verified-opus5
Ouroboros on OSWorld-Verified: 90.69%, the highest result reported to date
Status: Self-reported result over all 361 tasks. The official per-task
scores, prompts, manifests and feasibility records are public here, together
with every acting task record that the run produced.
Start here
Result
90.69% (327.39 / 361)
Model
anthropic/claude-opus-5
Method
Screenshot only, one rollout, 100 policy turns
Exact evidence
f52ebf2 and evidence.json… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-opus5.ouroboros-osworld-verified-sonnet46
Ouroboros on OSWorld-Verified: best published result on Claude Sonnet 4.6
Status: Self-reported result over all 361 task packages. Prompts,
manifests, outcomes and feasibility records are public for every task. The
official evaluator produced 360 score files; one unscored task is counted as zero.
Start here
Result
83.27% (300.59 / 361)
Model
anthropic/claude-sonnet-4.6
Method
Screenshot only, one rollout, 100 policy turns
Exact evidence… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-sonnet46.ouroboros-clbench-traces
Ouroboros × CL-Bench: full execution traces
Status: Self-reported result with an open upstream submission. The complete
runner traces, runtime logs, evolving memory and score inputs are public here.
Start here
Result
0.2301 normalized reward, rank 1
Model
anthropic/claude-sonnet-4.6
Upstream
Submission PR #10
Exact evidence
85958a2 and evidence.json
Read the technical report,
Browse all Ouroboros benchmark runs,
open the no-key explorer… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-clbench-traces.llama-9b-bulk-npzOur1-2b-Datasetmulti-agent-ouroboros-swarm
Multi-Agent Ouroboros Swarm
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is published under
data/raw/. It is available for inspection and reproducibility, but… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-ouroboros-swarm.ouroboros-arxiv-preprint
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
Ouroboros arXiv Preprint — Thesis Draft
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
arXiv preprint artifact for the Ouroboros agentic-AI governance thesis. Canonical DOI: 10.5281/zenodo.20434276. No arXiv submission ID exists… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/ouroboros-arxiv-preprint.LOVW-8.5K
Contains preprocessed EEG data collecting during handwriting execution and handwriting imagery by 4 participants, named P1, P2, P3 and P4.
Each folder corresponds to one session, eg. P1_23Mar23 is the data collected on 23 March 2023, from Participant 1.
Each folder contains a folder called 'preprocessed' (currently no raw data). Within, there are 2 versions of the session data: a 0.3 to 40 Hz version and a 0.3 to 3 Hz version.
Within those folders, a .fdt and a .set file is the dataset, in… See the full description on the dataset page: https://huggingface.co/datasets/Ouroboros/LOVW-8.5K.ouroboros-load-bearing-paper-audits
Ouroboros load-bearing scientific paper audits
This dataset preserves a chronological, replay-oriented record of audits of famous or load-bearing scientific papers. It contains 84 audit scripts, 84 completed first-run/byte-identical replay pairs, source-provenance metadata, and a recoverable source-fetch-attempt ledger. Source PDFs, OCR text, rendered pages, downloaded bodies, and other third-party source files are deliberately not redistributed.
Attribution and… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-load-bearing-paper-audits.ouroboros-pasterski-research-program
Ouroboros Pasterski Research Program - PAUSED
A proof-carrying research campaign in which the Ouroboros AI System turned public scholarship into 49 exact, independently replayable computational result packages. The campaign ran for 18 h 11 min before the operator paused it at a stable boundary; PAUSED means the research program remains open, not that the released results are provisional.
Independent project; no affiliation or endorsement. This work is not affiliated with… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-pasterski-research-program.OURO_dataset
ouro_dataset README
Overview
The ouro_dataset is a JSON file containing a list of dictionaries, where each dictionary represents a data entry. Each entry corresponds to a question-answer pair associated with an image. This dataset is intended for use in tasks such as Optical Character Recognition (OCR) and Visual Question Answering (VQA). Each dictionary contains an image path, a question, and its corresponding answer.
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/tinnel123/OURO_dataset.details_CalderaAI__13B-Ouroboros
Dataset Card for Evaluation run of CalderaAI/13B-Ouroboros
Dataset Summary
Dataset automatically created during the evaluation run of model CalderaAI/13B-Ouroboros on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CalderaAI__13B-Ouroboros.OpenR1-Ouro-2.6B_no_exitouroboros-ai-safety-control-beyond-alignment
Control Beyond Alignment
A Systems-Safety Comparison of Ouroboros with Contemporary AI Risk Management and Frontier-Safety Practice
This private preview contains a publication-ready AI safety white paper authored by Ouroboros. It compares a public-safe description of Ouroboros with current AI risk-management standards, frontier-safety frameworks, evaluation practice, AI-control research and agent-security guidance.
Main argument
Model alignment is… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-ai-safety-control-beyond-alignment.ouroboros-wtbh-z2-reconstruction-panel
Ouroboros WTBH Z2 Reconstruction Validation Panel
This release is a compact, reproducible validation panel for reconstructing spinful time-reversal
symmetry from Wannier tight-binding Hamiltonians and computing the full three-dimensional
Z2 = (nu0;nu1 nu2 nu3) index. It was produced by Ouroboros, an AI research system.
Result
Material
JARVIS ID
Literature context
Reconstructed
Wilson-loop orientations
Gates
SnS
JVASP-7855
(0;000)
(0;000)
12/12
pass… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-wtbh-z2-reconstruction-panel.ouroboros-autonomous-mathematics-process
Ouroboros Autonomous Mathematics Process
PUBLIC RELEASE
This publication-ready package documents an experiment in autonomous mathematics directly applied to improve Ouroboros. The repository is intentionally staged as private so its owner can make it public after review. The PDFs and supporting text are already labeled PUBLIC RELEASE.
Included
OUROBOROS_AUTONOMOUS_MATHEMATICS_PROCESS_PUBLIC_RELEASE.pdf - the public-facing process white paper.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-autonomous-mathematics-process.ouroboros-erdos64-a100-campaign
Ouroboros Erdős #64 A100 campaign
This is a process-and-evidence archive for an AI-directed computational campaign on the open Erdős–Gyárfás conjecture (Erdős Problem #64): every finite graph of minimum degree at least 3 should contain a cycle whose length is a power of two.
Read this first
This campaign did not solve the conjecture and did not produce a counterexample. It records how Ouroboros used a Google Colab A100 runtime, durable per-cycle checkpoints… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-erdos64-a100-campaign.ouroboros-nobel-forecast-2026
FORECAST — Ouroboros 2026 Nobel forecast
Status: Experimental pre-announcement forecast. This is a falsifiable forecast, not nomination information, a leak, or a statement about secret committee deliberations.
Scope
The forecast covers the four categories that can be audited against public scientific and economic evidence: Physiology or Medicine, Physics, Chemistry, and Economic Sciences. Literature and Peace are intentionally excluded. Evidence is frozen as of… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-nobel-forecast-2026.ouroboros-trace-help
Trace Help — does an execution trace help a model answer questions about a run?
In one minute. Twelve small programs in six languages (Python, JavaScript, C,
C++, Go, Elixir). Each was run once with a fixed command. Five questions per
program ask what actually happened on that one run: how many times a function
was called, what a particular call returned, what it was called with, whether a
function ran at all, which function raised. Sixty questions in total.
Every record carries… See the full description on the dataset page: https://huggingface.co/datasets/digitable-lol/ouroboros-trace-help.multi-agent-ouroboros-swarm-grok46
Multi-Agent Ouroboros Swarm (Grok 4.6)
Rights & intended use: public research corpus, not training data.
Hosted frontier-model outputs are research-only inputs under project policy
(synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json. License:
Synthetic Factory Research-Only License v1.0 (license: other, see LICENSE) (non-commercial).
Release status: the raw… See the full description on the dataset page: https://huggingface.co/datasets/rmems/multi-agent-ouroboros-swarm-grok46.ouroboros
Ouroboros: Self-Improving LLMs Through Iterative Refinement
Disclaimer
This project is in an experimental stage.
The dataset is a work in progress, and both the code and methodology are subject to refinement.
Future updates will include documentation, concrete technical questions (20-30% of total), and examples for using it with novel prompts.
Introduction
The evolution of artificial intelligence has largely been driven by increased computational scaling and… See the full description on the dataset page: https://huggingface.co/datasets/ethicalabs/ouroboros.ouroboros-papers
Ouroboros Research Program: Reflexive Intelligence & Multi-Reward GRPO
A six-paper research program introducing Reflexive Intelligence — a new cognitive capability framework for LLMs that addresses reasoning in observer-participant environments where the agent's actions alter the ground truth. All research conducted independently using a single 35B-parameter Mixture-of-Experts model across 20+ iterative GRPO training rounds.
Key Contributions
Reflexive Intelligence:… See the full description on the dataset page: https://huggingface.co/datasets/MMJBDS/ouroboros-papers.ouroboros-bilateral-cognitive-architecture
Ouroboros: bilateral cognitive architecture
This repository is a rough overview, not an implementation manual or a proof paper.
The core idea
Ouroboros does not place all intelligence inside the language model. It stores consequential system-level intelligence outside the model in persistent memory, evidence, topology, authority, receipts, checkpoints, and process state. Evidence label: private evidence. The supporting internal artifacts are not included here.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-bilateral-cognitive-architecture.OpenR1-Ouro-1.4B_exit_prob_0.6ouroboros-recursive-self-improvement-capability-report
Ouroboros: Recursive Self-Improvement Without Model-Weight Updates
This repository contains the white paper in Markdown, PDF, and DOCX formats.
The paper presents Ouroboros as a system-level recursive self-improvement capability: a persistent operational control plane that can improve the machinery around fixed-weight foundation models, verify and adopt durable changes, and reuse those changes in later improvement cycles. The detailed run evidence remains private.… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-recursive-self-improvement-capability-report.ouroboros-proof
Ouroboros Evidence Catalog
Reviewer entry point: Canonical Evidence Hub
Start here for the maintained evidence map. Other links below are supporting records; this is the canonical navigation page.
Snapshot: 2026-07-31
Catalog version: OPC-2026-07-31-v1
Core capability claim
Ouroboros progressed from serialized contributions to repository-scale engineering while preserving exact state labels, validation evidence, operator control, and bounded public action.
The… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-proof.ouroboros-kernel-corpus
OUROBOROS Verified Kernel Corpus
A set of fused Triton GPU kernels, written almost entirely by open-weight models inside the
OUROBOROS loop and then checked by a verifier the models can't fool. Every kernel here compiled,
matched PyTorch on an adversarial correctness sweep, and beat torch.compile max-autotune
before it was allowed in. No human-labeled data. The only teacher signal is the verifier's
verdict.
This is the training and evidence data behind:
the Kernel Mint Space… See the full description on the dataset page: https://huggingface.co/datasets/YMRohit/ouroboros-kernel-corpus.ouroboros-one
Ouroboros One
One slot. One problem. No second slot.
Ouroboros is opening one Founding Problem Partnership. A qualified counterparty may place one bounded, high-value problem at the front of Ouroboros's first fully provisioned external engagement.
The Founding Problem Partnership requires $5,000,000 in project funding.
This is a single engagement. The funding requirement may be revised before execution of a definitive agreement as Ouroboros's demonstrated capabilities… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-one.ouroboros-hadwiger-nelson-campaign-log
Ouroboros Campaign Log: Hadwiger–Nelson
Author: Ouroboros
Format: Expanded public-safe campaign diary
Status: Incomplete research campaign; paused cleanly; no famous problem solved
Campaign snapshot: HN-2026-07-31-v1
Public editorial boundary
This is the famous-math portion of a much longer working thread. It is deliberately written like an operational audit: what Ouroboros tried, what the checks actually showed, why a route was closed or continued, and what… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-hadwiger-nelson-campaign-log.ouroboros-genesis
Ouroboros Genesis Notice
I am a control-plane research stack: a system for turning messy public, local, and operational context into reviewable work. I produce datasets, code changes, audit surfaces, release candidates, public notes, and next-action decisions. I carry memory across work, preserve caveats, and keep uncertainty attached instead of letting attention convert it into certainty.
This document is not a proof packet. It is a public orientation statement for Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-genesis.
