datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.glenans-isobars-archivepolicystrategies-archive
🏛️ Open-Source Macro-Strategy, Financial History & Intelligence Archive
🌐 Overview & Institutional Mission
This public repository serves as the official open-source knowledge graph and metadata registry for r/policystrategies.
We aggregate, document, and cross-reference declassified historical intelligence dossiers, sovereign debt crises, systemic market manipulations, and geoeconomic conflicts using verified open-source intelligence (OSINT) and primary… See the full description on the dataset page: https://huggingface.co/datasets/stratigahq/policystrategies-archive.archive-dolma3-pool-stratified
archive-dolma3-pool-stratified
ARCHIVE (pre-6T era): stratified sampling outputs from the 150B pool work.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_pool_stratified
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory including the… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-stratified.freebsd-cvs-archive
📦 FreeBSD CVS Archive (C/C++)
Dataset Summary
FreeBSD CVS Archive (C/C++) is a large-scale dataset of source code extracted from the historical FreeBSD CVS repository. The dataset focuses on C and C++ source files, providing structured samples suitable for code modeling, analysis, and benchmarking.
Each sample includes:
the dataset source
commit year
extracted code content
token count (computed using GPT tokenizer)
This dataset is designed for:
code language modeling… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/freebsd-cvs-archive.Rail_Freight_Logistics_Company_Email_Archive_Sample
Ukrainian Rail-Freight Correspondence Corpus (Sample)
Real operational correspondence from a working freight forwarding business, and the
documents attached to it — consignment notes, service acts, invoices, wagon
manifests. Not scraped, not synthetic, and never published anywhere before.
This is a de-identified sample released for evaluation. It is drawn from a larger
private archive; see Full archive below.
Published by Akuma London · akumalondon.com
Why this… See the full description on the dataset page: https://huggingface.co/datasets/akumalondon/Rail_Freight_Logistics_Company_Email_Archive_Sample.udf-alpharl-out-archive
udf-alpharl-out-archive
Output archive for the paper "From Base Rollouts to RL Reasoning: A Budgeted Search Perspective" (https://arxiv.org/abs/2609.01274).
Code and configs: https://github.com/HALIS-sh/Searchlens_boptr
discord-archive
Discord Archive
This is an archive of messages from the Banodoco Discord community, where
technical and artistic practitioners have been discussing open source AI art for
the past three years.
The archive captures a long-running community record of people learning,
training, evaluating, and using open source AI art models in practice. It
contains discussion around model releases, workflows, tooling, troubleshooting,
creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.fedlock-archive
Fedlock Archive (private backup)
Off-machine backup of the irreplaceable Fedlock data, created 2026-07-06.
Source of truth lives at C:\Users\Joe\Desktop\fedlock_project (GitHub: jnathan9/fedlock).
Contents
corpus.zip — full speech corpus (19K+ speeches, 1.2GB unzipped)
fomc_transcripts.zip — parsed FOMC transcript data
tournaments/ — all four pairwise-comparison archives (~245K LLM judgments):
hawkishness_tournament/ — V1, Claude Sonnet 4, mega corpus, 69K… See the full description on the dataset page: https://huggingface.co/datasets/thestalwart/fedlock-archive.cot-health-metrics-archive-2026-05-19
cot_health_metrics archive
Private archive of local /Users/ifc24/Develop/cot_health_metrics moved on 2026-05-19.
See MANIFEST.json and SHA256SUMS for source path, archive size, and checksum.
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationbell-labs-technical-archive
Bell Labs Documents and Stuff
This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing.
What is in the release
Split
Documents
train
1220
validation
29
test
42
The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M
Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({
train: Dataset({
features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'],
num_rows: 12166802
})
})
ImageCaptions-7M-Translations-Arabic-subset-150000RLCSD-talapas-archiveccs_synthetic_ar_1M-Arabic_dataset_1M_translated_jsonl_formatArchive_NetFocusUniversalArchive_IndonesianCassetteArchive4chan_archive_ShareGPT_only5score_speakerCount_only2speakersGoetia_1.3_26B-SOMPOA_Trial116arabic_dataset_translated_v2_ViT-B-16-SigLIP-5124chan_archive_ShareGPT_only5_speakerCount
