datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NextSearch-1-Trajectories
NextSearch-1 Trajectories
The supervised training trajectories behind the
NextSearch-1 web research agents:
complete research episodes — reasoning, tool calls, live-web tool results,
and final answers — for every task in the companion
NextSearch-1-Tasks
SFT configs. Directly trainable: each row is a prompt (messages) plus a
target trajectory (target) with per-message reasoning and OpenAI-format
tool calls.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Trajectories.NextSearch-1-Tasks
NextSearch-1 Tasks
The task pools behind the NextSearch-1
web research agents: every row is a research question with its reference
answer and grading spec — the sft-tasks configs are the tasks behind the
supervised corpora, the rl-tasks configs the verified prompt+gold pools
used for reinforcement learning. Full trajectories for the SFT configs are in the companion
NextSearch-1-Trajectories.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Tasks.nexttoken-model-2-dataset-sft
NextToken Model 2 (SAM) SFT dataset
Training data for
somasekhar-dev/NextToken-model-2
(SAM -- Small Action Model), a banking tool-calling assistant. This is the
sep18_round2 dataset (kashyap/task-1/dataset.jsonl) that trained the
current best checkpoint.
Files
dataset.jsonl (4,914 rows) -- chat-format (messages: system/user/assistant),
each row's system message embeds the tool schema subset shown for that
example (see below).
manifest.json -- the actual training… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-2-dataset-sft.next_token
Supreme Court of India Judgments Dataset (1950-2025)
Dataset Description
This dataset contains a comprehensive collection of judgments and orders from the Supreme Court of India, spanning from its inception in 1950 up to early 2025.
Dataset Summary
Total Documents: 26,688
Total Tokens: ~196.9 Million (counted using cl100k_base encoding)
Format: JSONL (JSON Lines)
Language: English
Time Range: 1950 - 2025
Data Fields
Each entry in the .jsonl file… See the full description on the dataset page: https://huggingface.co/datasets/psychopenguin/next_token.
