datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P
This dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Claude Opus 4.8 Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by anthropic/claude-opus-4.8.
JSONL files: 4
Training-ready tools
A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/Quaxicron/claude-opus-4.8-pi-traces.claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P
This dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Claude Opus 4.8 Pi Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by anthropic/claude-opus-4.8.
JSONL files: 4
Training-ready tools
A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-opus-4.8-pi-traces.v3-2k-traj-claude-opus-4.7v4-4k-traj-claude-opus-4.7swebench-multilingual-claude-opus-4.7swebench-verified-claude-opus-4.7Qwen3.5-4B-Claude-Opus-Reasoning-Distill-MMLU-Pro-benchmarkBenchmark of TeichAI/Qwen3.5-4B-Claude-Opus-Reasoning-Distill against TIGER-Lab/MMLU-Pro dataset.
Accuracy: 72.89999999999999% with Python tool.
Metric
Value
Correct
729
Incorrect
268
Errors
3
Total samples
1000
Python tool calls
1078
Total completion tokens
2,449,614
Raw stats:
{
"accuracy": 0.729,
"correct": 729,
"incorrect": 268,
"error": 3,
"total": 1000,
"python_tool_calls": 1078,
"completion_tokens":2449614
}
Qwen3.5-4B-Claude-Opus-Reasoning-Distill-SuperGPQA-benchmarkBenchmark of TeichAI/Qwen3.5-4B-Claude-Opus-Reasoning-Distill against m-a-p/SuperGPQA dataset.
Accuracy: 41.5% with Python tool.
Metric
Value
Correct
415
Incorrect
573
Errors
11
Total samples
999
Python tool calls
2527
Total completion tokens
4,149,159
Raw stats:
{
"accuracy": 0.415,
"correct": 415,
"incorrect": 573,
"error": 11,
"total": 999,
"python_tool_calls": 2527,
"completion_tokens": 4149159
}
qgqa-gpqa-migrate-20260219-141149-claude-opus-4-6qgqa-unified-processed-claude-opus-4-6-20260213-033913Qwen3.5-4B-Claude-Opus-Reasoning-Distill-GPQA-Diamond-benchmarkBenchmark of TeichAI/Qwen3.5-4B-Claude-Opus-Reasoning-Distill against fingertap/GPQA-Diamond dataset.
Accuracy: 60.099999999999994% with Python tool.
Metric
Value
Correct
119
Incorrect
78
Errors
1
Total samples
198
Python tool calls
189
Total completion tokens
871,865
Raw stats:
{
"accuracy": 0.601,
"correct": 119,
"incorrect": 78,
"error": 1,
"total": 198,
"python_tool_calls": 189,
"completion_tokens":871865
}
qgqa-claude-opus-4-6-20260213-041708res_gptoss120b_original_1_high_0.7_16000_claude-opus-4-7res_gptoss120b_original_1_xhigh_0.7_32000_claude-opus-4-7query-evaluation-full_model_eval_claude_4_opusqgqa-testing-claude-opus-4-6
