CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NextTokenAI /NextSearch-1-Trajectories NextSearch-1 Trajectories The supervised training trajectories behind the NextSearch-1 web research agents: complete research episodes — reasoning, tool calls, live-web tool results, and final answers — for every task in the companion NextSearch-1-Tasks SFT configs. Directly trainable: each row is a prompt (messages) plus a target trajectory (target) with per-message reasoning and OpenAI-format tool calls. Technical report: nexttoken.co/research/nextsearch-1 · Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Trajectories.texttext-generation1K<n<10K0 likes110 downloads1mo agoHugging Face02next-token /clean-PD-16000-books3 📚 clean-PD-16000-books3 A treasure trove of ~16,000 high-quality, public domain books in English language — nicely cleaned, with rich metadata, and ready for language modeling. ✨ What Makes This Dataset Special? This isn’t just another dump of dusty old text files. clean-PD-16000-books3 is the result of a rigorous cleaning and curation process applied to a large collection of public domain literature, including: ✅ Readable prose — paragraphized prose, without unnatural… See the full description on the dataset page: https://huggingface.co/datasets/next-token/clean-PD-16000-books3.text10K<n<100K6 likes101 downloads1y agoHugging Face03NextTokenAI /NextSearch-1-Tasks NextSearch-1 Tasks The task pools behind the NextSearch-1 web research agents: every row is a research question with its reference answer and grading spec — the sft-tasks configs are the tasks behind the supervised corpora, the rl-tasks configs the verified prompt+gold pools used for reinforcement learning. Full trajectories for the SFT configs are in the companion NextSearch-1-Trajectories. Technical report: nexttoken.co/research/nextsearch-1 · Harness and evals:… See the full description on the dataset page: https://huggingface.co/datasets/NextTokenAI/NextSearch-1-Tasks.textquestion-answering10K<n<100K2 likes77 downloads1mo agoHugging Face04somasekhar-dev /nexttoken-pmkisan-domain-sft-data NextToken pmkisan domain SFT data (v1) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from ~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA, banking products, insurance, savings instruments, etc.). Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.tabularquestion-answering1K<n<10K0 likes53 downloads7d agoHugging Face05Mwanzau /Tumbuka_Continuous_Next-Token_Prediction_Datasettext1K<n<10K0 likes22 downloads2mo agoHugging Face06psychopenguin /next_token Supreme Court of India Judgments Dataset (1950-2025) Dataset Description This dataset contains a comprehensive collection of judgments and orders from the Supreme Court of India, spanning from its inception in 1950 up to early 2025. Dataset Summary Total Documents: 26,688 Total Tokens: ~196.9 Million (counted using cl100k_base encoding) Format: JSONL (JSON Lines) Language: English Time Range: 1950 - 2025 Data Fields Each entry in the .jsonl file… See the full description on the dataset page: https://huggingface.co/datasets/psychopenguin/next_token.texttext-generation10K<n<100K0 likes18 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.