next-token
Next_Token_Prediction_datasetNextSearch-1-Trajectories
NextSearch-1 Trajectories
The supervised training trajectories behind the
NextSearch-1 web research agents:
complete research episodes — reasoning, tool calls, live-web tool results,
and final answers — for every task in the companion
NextSearch-1-Tasks
SFT configs. Directly trainable: each row is a prompt (messages) plus a
target trajectory (target) with per-message reasoning and OpenAI-format
tool calls.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Trajectories.clean-PD-16000-books3
📚 clean-PD-16000-books3
A treasure trove of ~16,000 high-quality, public domain books in English language — nicely cleaned, with rich metadata, and ready for language modeling.
✨ What Makes This Dataset Special?
This isn’t just another dump of dusty old text files.
clean-PD-16000-books3 is the result of a rigorous cleaning and curation process applied to a large collection of public domain literature, including:
✅ Readable prose — paragraphized prose, without unnatural… See the full description on the dataset page: https://huggingface.co/datasets/next-token/clean-PD-16000-books3.NextSearch-1-Tasks
NextSearch-1 Tasks
The task pools behind the NextSearch-1
web research agents: every row is a research question with its reference
answer and grading spec — the sft-tasks configs are the tasks behind the
supervised corpora, the rl-tasks configs the verified prompt+gold pools
used for reinforcement learning. Full trajectories for the SFT configs are in the companion
NextSearch-1-Trajectories.
Technical report: nexttoken.co/research/nextsearch-1 ·
Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Tasks.nexttoken-model-1-dataset-sft
NextToken Model 1 SFT dataset (v4)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
846 scheme/product source documents across 57 schemes/products, chunked
into 1,445 passages.
v4 vs v3: v3 merged in a second batch (6,664 rows) without… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-1-dataset-sft.llava-next-qwen-format-400k-tokenized
