datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NextSearch-1-Trajectories
NextSearch-1 Trajectories
The supervised training trajectories behind the
NextSearch-1 web research agents:
complete research episodes — reasoning, tool calls, live-web tool results,
and final answers — for every task in the companion
NextSearch-1-Tasks
SFT configs. Directly trainable: each row is a prompt (messages) plus a
target trajectory (target) with per-message reasoning and OpenAI-format
tool calls.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Trajectories.clean-PD-16000-books3
📚 clean-PD-16000-books3
A treasure trove of ~16,000 high-quality, public domain books in English language — nicely cleaned, with rich metadata, and ready for language modeling.
✨ What Makes This Dataset Special?
This isn’t just another dump of dusty old text files.
clean-PD-16000-books3 is the result of a rigorous cleaning and curation process applied to a large collection of public domain literature, including:
✅ Readable prose — paragraphized prose, without unnatural… See the full description on the dataset page: https://huggingface.co/datasets/next-token/clean-PD-16000-books3.NextSearch-1-Tasks
NextSearch-1 Tasks
The task pools behind the NextSearch-1
web research agents: every row is a research question with its reference
answer and grading spec — the sft-tasks configs are the tasks behind the
supervised corpora, the rl-tasks configs the verified prompt+gold pools
used for reinforcement learning. Full trajectories for the SFT configs are in the companion
NextSearch-1-Trajectories.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Tasks.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.Tumbuka_Continuous_Next-Token_Prediction_Datasetnext_token
Supreme Court of India Judgments Dataset (1950-2025)
Dataset Description
This dataset contains a comprehensive collection of judgments and orders from the Supreme Court of India, spanning from its inception in 1950 up to early 2025.
Dataset Summary
Total Documents: 26,688
Total Tokens: ~196.9 Million (counted using cl100k_base encoding)
Format: JSONL (JSON Lines)
Language: English
Time Range: 1950 - 2025
Data Fields
Each entry in the .jsonl file… See the full description on the dataset page: https://huggingface.co/datasets/psychopenguin/next_token.
