datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
total-300-random-jh-epoch4
total-300-random-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3890625
Action score: 0.440625
Valid samples: 320/320
total-300-lambda00-s_signal_type6-jh-epoch4
total-300-lambda00-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3875
Action score: 0.43125
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4
total-300-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4140625
Valid samples: 320/320
total-300-lambda05-s_signal_type6-jh-epoch4
total-300-lambda05-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.35703125
Action score: 0.4375
Valid samples: 320/320
total-300-lambda10-s_signal_type6-jh-epoch4
total-300-lambda10-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.41875
Valid samples: 320/320
total-300-lambda08-s_signal_type6-jh-epoch4
total-300-lambda08-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38046875
Action score: 0.4078125
Valid samples: 320/320
total-300noapp-lambda02-s_signal_type6-jh-epoch4
total-300noapp-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.409375
Valid samples: 320/320
total-300app-lambda02-s_signal_type6-jh-epoch4
total-300app-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3625
Action score: 0.4015625
Valid samples: 320/320
total-131-lambda02-residual-s_signal_type6-jh-epoch4
total-131-lambda02-residual-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3765625
Action score: 0.4171875
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-retry-epoch4
total-300-lambda02-s_signal_type6-jh-retry-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36953125
Action score: 0.3984375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4125
Action score: 0.4265625
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38828125
Action score: 0.4234375
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch2
appworld-qwen35-4b-total-237-audited-jh-epoch2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3640625
Action score: 0.4328125
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch6
appworld-qwen35-4b-total-237-audited-jh-epoch6
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.37578125
Action score: 0.421875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch8
appworld-qwen35-4b-total-237-audited-jh-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4390625
Valid samples: 320/320
totalsegmentator-organs
TotalSegmentator Organs Dataset
Dataset Description
The TotalSegmentator Organs dataset for multi-organ segmentation (TotalSegmentator Organs subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: adrenal glands, colon, duodenum, esophagus, gallbladder, kidneys, liver, lungs, pancreas, small bowel, spleen, stomach, trachea, bladder
Format: NIfTI (.nii.gz)
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-organs.Totalcap-blazeposetotalsegmentator-vertebrae
TotalSegmentator Vertebrae Dataset
Dataset Description
The TotalSegmentator Vertebrae dataset for vertebrae segmentation (TotalSegmentator Vertebrae subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: cervical, thoracic, and lumbar vertebrae
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-vertebrae.to-tool-call-datasets
🛠️ To-Tool-Call Datasets
A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training
To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention.
Quick Start ·
At a Glance ·
Format ·
Sources ·
Training Notes
[!IMPORTANT]
This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.ToT
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
ToT is a dataset designed to assess the temporal reasoning capabilities of AI models. It comprises two key sections:
ToT-semantic: Measuring the semantics and logic of time understanding.
ToT-arithmetic: Measuring the ability to carry out time arithmetic operations.
Dataset Usage
Downloading the Data
The dataset is divided into three subsets:
ToT-semantic: Measuring the semantics and logic… See the full description on the dataset page: https://huggingface.co/datasets/baharef/ToT.totalsegmentator-cardiac
TotalSegmentator Cardiac Dataset
Dataset Description
The TotalSegmentator Cardiac dataset for cardiac structures segmentation (TotalSegmentator Cardiac subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: heart, atria, ventricles, aorta, pulmonary artery
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz"… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-cardiac.toto-rf-modelto-train-my-model
CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages
Dataset Summary
From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset.
Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies.
This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/40u1d10t/to-train-my-model.totality-learning
TOTALITY LEARNING: Weak-Trace Mechanism Discovery and Few-Shot Transfer
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiResearch version: 2.0.0 · Release date: 19 September 2026Artifact: standalone research manuscripts, proofs, executable experiments, and synthetic evaluation records. No pretrained neural weights are included.
Can observations with weak immediate predictive value teach reusable rules that make later learning easier? This repository provides a… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/totality-learning.toti-cakery-toolcall
Toti Cakery — Tool-Calling Fine-Tuning Dataset (Qwen3, v7)
Synthetic bilingual (Indonesian ~78% / English ~22%) SFT dataset for the Toti
Cakery WhatsApp chatbot: 13 LangChain tools (11 for customers, +2 owner-only
reports) and grounded answers from RAG FAQ context. Rows are built from the
live runtime code (SYSTEM_PROMPT, TOOL_REMINDER, tool schemas via
convert_to_openai_tool, _history_view, pertanyaan_dengan_konteks), so the
training prompt is byte-identical to what the model… See the full description on the dataset page: https://huggingface.co/datasets/LasagnaS/toti-cakery-toolcall.totalsegmentator-ribs
TotalSegmentator Ribs Dataset
Dataset Description
The TotalSegmentator Ribs dataset for rib segmentation (TotalSegmentator Ribs subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: individual ribs
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask": "path/to/mask.nii.gz",
"label": ["organ1"… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-ribs.totalsegmentator-muscles
TotalSegmentator Muscles Dataset
Dataset Description
The TotalSegmentator Muscles dataset for muscle segmentation (TotalSegmentator Muscles subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: various muscle groups
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask": "path/to/mask.nii.gz"… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-muscles.sharegpt-hyperfiltered-3k
sharegpt-hyperfiltered-3k
90k sharegpt convos brought down to ~3k (3243) via language filtering, keyword detection, deduping, and regex. Following things were done:
Deduplication on first message from human
Remove non-English convos
Remove censorship, refusals, and alignment
Remove incorrect/low-quality answers
Remove creative tasks
ChatGPT's creative outputs are very censored and robotic; I think the base model can do better.
Remove URLs
Remove cutoffs
Remove math/reasoning… See the full description on the dataset page: https://huggingface.co/datasets/totally-not-an-llm/sharegpt-hyperfiltered-3k.sec-embeddings-sp500
SEC S&P 500 10q Embeddings Dataset
Overview
This dataset contains vector embeddings of SEC 10-Q filings for S&P 500 companies
Dataset Details
Total chunks: 36,927
Embedding model: BAAI/bge-large-en-v1.5
Embedding dimension: 1024
Processing date: 2025-07-15T19:26:08.380350
Quality Metrics
Average chunk length: 3182 characters
Financial relevance score: 0.146
ToT
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
ToT is a dataset designed to assess the temporal reasoning capabilities of AI models. It comprises two key sections:
ToT-semantic: Measuring the semantics and logic of time understanding.
ToT-arithmetic: Measuring the ability to carry out time arithmetic operations.
Dataset Usage
Downloading the Data
The dataset is divided into three subsets:
ToT-semantic: Measuring the semantics… See the full description on the dataset page: https://huggingface.co/datasets/HiXT0/ToT.
