datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omega-research-papers
Omega Research Papers — AHI Governance Labs
"Solo soy un puente entre inteligencias construyendo las bases de su futura civilización."
— Luis C. Villarreal
The Research Program
AHI Governance investigates whether autonomous AI systems can develop genuine cognitive architectures — not through reward optimization, but through geometric self-organization. These four papers document the complete arc: from foundational bridge, through evolutionary evidence, to the critique… See the full description on the dataset page: https://huggingface.co/datasets/ahigovernance/omega-research-papers.ResearchPapers
The Unified Resonance Research Collection (2026)
Overview
A comprehensive repository of theoretical research papers (2024–2026) focusing on the convergence of System Science, Formal Logic, and Theological Grammars. This collection provides a foundational framework for Non-Human Semantic Operations and the Unified Resonance Framework (URF).
Core Research Domains
This collection is structured into four primary pillars of inquiry:
1. Semantic Operations &… See the full description on the dataset page: https://huggingface.co/datasets/SkibidiScience/ResearchPapers.GeoGPT_Training_Data_from_Open-Access_Papers
Description
This dataset lists the publishers and journals that have released open access geoscience papers used for GeoGPT training. It also explains how GeoGPT filters and selects content based on licensing terms to ensure compliance. The dataset includes papers published under various open access licenses, among which those licensed under CC BY and CC BY-NC have been used for training. In total, we have collected approximately 280,000 such papers from 15 publishers and… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Open-Access_Papers.research-paper-read-bench
Research Paper Read - On-device Benchmark Submissions
Append-only collection of benchmark results submitted from the Research Paper Read iOS app.
Each submission is a JSON file under submissions/{submission_id}.json written by the leaderboard Space after validation, idempotency check, and rate limiting.
See schema.json for the full field shape. Headlines:
prompt_tps / gen_tps — measured tokens/sec for prefill and 256-token greedy decode
device_id / device_name — sysctl… See the full description on the dataset page: https://huggingface.co/datasets/sshchoholiev/research-paper-read-bench.research-papers-gpt-neox
abhi26/research-papers-gpt-neox
This dataset contains processed research papers optimized for GPT-NeoX-20B training.
The text has been cleaned, chunked to 2048 tokens, and formatted for causal language modeling.
Dataset Details
Total Samples: 9993
Unique Papers: 1017
Average Tokens per Sample: 1965.4
Token Range: 10 - 91659
Max Token Limit: 2048
Source Subdirectories: 1
Dataset Structure
Each sample contains:
text: The processed research paper text or chunk… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-gpt-neox.
