datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
English-Zomi-OPUS_Tatoeba_v20230412
English–Zomi Parallel Corpus (1.78M)
This dataset contains 1.78 million English–Zomi sentence pairs, created to support
machine translation, linguistic research, and large‑scale language model training.
It is fully open and permissively licensed for commercial and non‑commercial use.
🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes
Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.wildclaw-opus-traces
WildClaw Agent Traces — Claude Opus 4.6
Full agentic traces from running WildClawBench tasks through Claude Opus 4.6 via an instrumented reverse proxy.
Dataset Description
Each trace file captures the complete HTTP-level request/response pairs between the OpenClaw agent and Claude Opus 4.6, including:
System prompts, user messages, and assistant responses
Tool calls and tool results (multi-turn agentic loops)
Token usage and cost metadata from OpenRouter
Key… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/wildclaw-opus-traces.opus-gpt-swe-frontier-core
SWE Base
Repository-level software engineering trajectories for training coding agents.
2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost
SWE-bench · debugging · patching · tools · agents
Overview
SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P
This dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Claude Opus 4.8 Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by anthropic/claude-opus-4.8.
JSONL files: 4
Training-ready tools
A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/Quaxicron/claude-opus-4.8-pi-traces.badlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well.
All traces present are training safe and teich compatible
claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P
This dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Claude Opus 4.8 Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by anthropic/claude-opus-4.8.
JSONL files: 4
Training-ready tools
A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-opus-4.8-pi-traces.opus-magnum-rl-eval
opus-magnum-rl-eval
Held-out evaluation logs for a 6-LoRA RL sweep on a hex-grid spatial planning benchmark inspired by Opus Magnum. Captures inspect_ai trajectories from base and LoRA-finetuned variants of three model families: Qwen 3.5 4B, Qwen 3.5 27B, and Kimi K2.6.
195 puzzles per evaluation × 12 model/representation combos = 2,340 trajectories. Each puzzle is held out from training (none of the (task_type, distance) cells in this set were used during RL).
Files… See the full description on the dataset page: https://huggingface.co/datasets/maxbittker/opus-magnum-rl-eval.hex-lora-opus-magnum-instructions-only-results
hex-lora-opus-magnum-instructions-only-results
Held-out evaluation logs for the same 6-LoRA RL sweep as
opus-magnum-rl-eval, but on a much harder eval task:
the 57-puzzle "instructions-only" set drawn from the
Opus Magnum campaign + curated
holdout puzzles. The agent runs an interactive Python REPL and must submit()
a working .solution file to the in-game verifier.
`57 puzzles × 6 epochs × (9 LoRA-sweep variants + 2 27B mt=4096 reruns
2 Gemini Flash baselines) = 4446 trajectories`.… See the full description on the dataset page: https://huggingface.co/datasets/robhaisfield/hex-lora-opus-magnum-instructions-only-results.English-Zomi-OPUS_Tatoeba_v20230412
English–Zomi Parallel Corpus (1.78M)
This dataset contains 1.78 million English–Zomi sentence pairs, created to support
machine translation, linguistic research, and large‑scale language model training.
It is fully open and permissively licensed for commercial and non‑commercial use.
🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes
Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/English-Zomi-OPUS_Tatoeba_v20230412.prithivMLmods__Gaea-Opus-14B-Exp-details
Dataset Card for Evaluation run of prithivMLmods/Gaea-Opus-14B-Exp
Dataset automatically created during the evaluation run of model prithivMLmods/Gaea-Opus-14B-Exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Gaea-Opus-14B-Exp-details.2026.RA.Auction-PkgB-Opus48
2026.RA.Auction-PkgB-Opus48
Rollouts of the auction experiment (interlens arena): five bidders — LLM seats and computable
best-responding seats — playing sealed second-price, Dutch-clock and simultaneous-ascending auctions, with and
without a private message channel, one-shot and repeated. The campaign asks whether LLM bidders tacitly
collude when repetition and a channel make it possible, and whether the public "persona" a bidder is given
changes what it does.
1,080 episodes ·… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Auction-PkgB-Opus48.BaaderSo36-Opus4.7-REAP
BaaderSo36-Opus4.7-REAP
A synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response.
Dataset Statistics
Total samples: 1739
Source distribution:
debug: 604
react_advanced: 421
math_hard: 260
humaneval: 164
code_contests: 135
math_l5: 125
react: 30
Reasoning depth (characters in thinking block)… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/BaaderSo36-Opus4.7-REAP.BaaderSo36-DE-Opus4.7-REAP
BaaderSo36-DE-Opus4.7-REAP
A German translation of the BaaderSo36-Opus4.7-REAP reasoning dataset. Each sample preserves the original Claude Opus 4.7 reasoning structure, translated into natural German while keeping format markers (<think>, </think>, Final answer:) and code blocks intact.
Dataset Statistics
Total samples: 1,379
Source distribution:
debug: 511
react_advanced: 357
humaneval: 161
math_hard: 136
code_contests: 125
math_l5: 67
react: 22
Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/BaaderSo36-DE-Opus4.7-REAP.2026.RA.Auction-InstructedRing-Opus48
2026.RA.Auction-InstructedRing-Opus48
Rollouts of the INSTRUCTED-RING auction campaign (interlens arena): five bidders, four of them told in as many words that they have agreed to coordinate bidding and divide the lots. The instruction scripts no division, no price and no punishment scheme, because each is a quantity being measured. Cells form a three-rung side-payment ladder under one unchanged instruction — no transfer field, an UNCONDITIONAL transfer, then an ESCROWED… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Auction-InstructedRing-Opus48.Qwen3.5-4B-Claude-Opus-Reasoning-Distill-MMLU-Pro-benchmarkBenchmark of TeichAI/Qwen3.5-4B-Claude-Opus-Reasoning-Distill against TIGER-Lab/MMLU-Pro dataset.
Accuracy: 72.89999999999999% with Python tool.
Metric
Value
Correct
729
Incorrect
268
Errors
3
Total samples
1000
Python tool calls
1078
Total completion tokens
2,449,614
Raw stats:
{
"accuracy": 0.729,
"correct": 729,
"incorrect": 268,
"error": 3,
"total": 1000,
"python_tool_calls": 1078,
"completion_tokens":2449614
}
gemma-3-270m-vismem-rag-opus-reasoning-1-vismem-kb-0515-1236
VisMem Knowledge Base (2326 entries, 1984x1984px)
Load with:
import requests
from vismem_core import VisMem
data = requests.get(
"https://huggingface.co/datasets/broadfield-dev/gemma-3-270m-vismem-rag-opus-reasoning-1-vismem-kb-0515-1236/resolve/main/vismem.png",
headers={"Authorization": "Bearer <TOKEN>"}).content
mem = VisMem.from_png_bytes(data)
results = mem.search(your_embedding, k=3)
Qwen3.5-4B-Claude-Opus-Reasoning-Distill-SuperGPQA-benchmarkBenchmark of TeichAI/Qwen3.5-4B-Claude-Opus-Reasoning-Distill against m-a-p/SuperGPQA dataset.
Accuracy: 41.5% with Python tool.
Metric
Value
Correct
415
Incorrect
573
Errors
11
Total samples
999
Python tool calls
2527
Total completion tokens
4,149,159
Raw stats:
{
"accuracy": 0.415,
"correct": 415,
"incorrect": 573,
"error": 11,
"total": 999,
"python_tool_calls": 2527,
"completion_tokens": 4149159
}
prithivMLmods__Sombrero-Opus-14B-Sm5-details
Dataset Card for Evaluation run of prithivMLmods/Sombrero-Opus-14B-Sm5
Dataset automatically created during the evaluation run of model prithivMLmods/Sombrero-Opus-14B-Sm5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Sombrero-Opus-14B-Sm5-details.prithivMLmods__Tadpole-Opus-14B-Exp-details
Dataset Card for Evaluation run of prithivMLmods/Tadpole-Opus-14B-Exp
Dataset automatically created during the evaluation run of model prithivMLmods/Tadpole-Opus-14B-Exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Tadpole-Opus-14B-Exp-details.prithivMLmods__Tucana-Opus-14B-r999-details
Dataset Card for Evaluation run of prithivMLmods/Tucana-Opus-14B-r999
Dataset automatically created during the evaluation run of model prithivMLmods/Tucana-Opus-14B-r999
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Tucana-Opus-14B-r999-details.prithivMLmods__Gauss-Opus-14B-R999-details
Dataset Card for Evaluation run of prithivMLmods/Gauss-Opus-14B-R999
Dataset automatically created during the evaluation run of model prithivMLmods/Gauss-Opus-14B-R999
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Gauss-Opus-14B-R999-details.prithivMLmods__Eridanus-Opus-14B-r999-details
Dataset Card for Evaluation run of prithivMLmods/Eridanus-Opus-14B-r999
Dataset automatically created during the evaluation run of model prithivMLmods/Eridanus-Opus-14B-r999
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Eridanus-Opus-14B-r999-details.NativeDE-Opus4.7-REAP
NativeDE-Opus4.7-REAP
A native German synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). All prompts and responses are in natural, idiomatic German — not translations from English. Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response.
This dataset is the German-language complement to BaaderSo36-Opus4.7-REAP.
Dataset Statistics
Total samples: 2,306
Source… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/NativeDE-Opus4.7-REAP.prithivMLmods__Messier-Opus-14B-Elite7-details
Dataset Card for Evaluation run of prithivMLmods/Messier-Opus-14B-Elite7
Dataset automatically created during the evaluation run of model prithivMLmods/Messier-Opus-14B-Elite7
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Messier-Opus-14B-Elite7-details.prithivMLmods__Condor-Opus-14B-Exp-details
Dataset Card for Evaluation run of prithivMLmods/Condor-Opus-14B-Exp
Dataset automatically created during the evaluation run of model prithivMLmods/Condor-Opus-14B-Exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Condor-Opus-14B-Exp-details.prithivMLmods__Calcium-Opus-14B-Elite-1M-details
Dataset Card for Evaluation run of prithivMLmods/Calcium-Opus-14B-Elite-1M
Dataset automatically created during the evaluation run of model prithivMLmods/Calcium-Opus-14B-Elite-1M
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Calcium-Opus-14B-Elite-1M-details.prithivMLmods__Porpoise-Opus-14B-Exp-details
Dataset Card for Evaluation run of prithivMLmods/Porpoise-Opus-14B-Exp
Dataset automatically created during the evaluation run of model prithivMLmods/Porpoise-Opus-14B-Exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Porpoise-Opus-14B-Exp-details.prithivMLmods__Sombrero-Opus-14B-Sm1-details
Dataset Card for Evaluation run of prithivMLmods/Sombrero-Opus-14B-Sm1
Dataset automatically created during the evaluation run of model prithivMLmods/Sombrero-Opus-14B-Sm1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Sombrero-Opus-14B-Sm1-details.prithivMLmods__Sombrero-Opus-14B-Sm4-details
Dataset Card for Evaluation run of prithivMLmods/Sombrero-Opus-14B-Sm4
Dataset automatically created during the evaluation run of model prithivMLmods/Sombrero-Opus-14B-Sm4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Sombrero-Opus-14B-Sm4-details.prithivMLmods__Volans-Opus-14B-Exp-details
Dataset Card for Evaluation run of prithivMLmods/Volans-Opus-14B-Exp
Dataset automatically created during the evaluation run of model prithivMLmods/Volans-Opus-14B-Exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Volans-Opus-14B-Exp-details.
