datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openswe-harbor
OpenSWE-Harbor — NeMo Gym ready
⚠️ Read this before training: there are TWO sets in here — full (45,316) and filtered (8,875).
Set
Tasks
Where it is
When to use
Full
45,316
routing/openswe_oss.jsonl, routing/openswe_other.jsonl, all of tasks/
Eval-only, dataset analysis, sweeps where you don't care about RL signal quality
Filtered (RL default)
8,875
routing/openswe_oss_filtered.jsonl, routing/openswe_other_filtered.jsonl, filtered_ids.txt
Use this for RL… See the full description on the dataset page: https://huggingface.co/datasets/ritvik-sarvam/openswe-harbor.Sarvam-105b-Distill-100k
Sarvam 105B Distill 100K
Dataset Summary
Reasoning dataset distilled from Sarvam 105B packaged with / tags in thinking/sharegpt/chatml/simple_qa schemas.
Source
Input JSONL: distillation_pipeline\dataset_final_p1_100k\full_dataset.jsonl
Generated at: 2026-05-17T11:24:23.743584+00:00
Splits
Train: 92040
Validation: 1917
Test: 1918
Distribution Counts
Domain
coding_computer_science: 16667
creative_planning_openended: 3809… See the full description on the dataset page: https://huggingface.co/datasets/abhinav0231/Sarvam-105b-Distill-100k.sarvam-30b-audit-prompts
Sarvam-30B Responsible-AI Audit — Pre-Registered Prompt Manifest
120 prompts across 5 categories, sampled deterministically (seed = 42) and pre-registered
as the eval contract for a public responsible-AI audit of
Sarvam-30B, India's sovereign-built
reasoning LLM.
This dataset is the eval contract committed to git before any prompt was sent to the model.
Reviewers can verify every prompt by going to the cited source and pulling that exact row.
Composition
#… See the full description on the dataset page: https://huggingface.co/datasets/procodec/sarvam-30b-audit-prompts.sarvam-m-testseagle3-sarvam-30b-training-data
Eagle3 Sarvam-30B Training Data
Training data used to build the Eagle3 draft model for Sarvam-30B.
Dataset Description
This dataset contains 90,000 prompt-response pairs used to train an Eagle3 speculative decoding draft model for the Sarvam-30B language model.
Each sample consists of a prompt and its corresponding response generated by the Sarvam-30B base model. During training, the model also consumes hidden state features extracted from auxiliary layers of the base… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/eagle3-sarvam-30b-training-data.sarvamai__OpenHathi-7B-Hi-v0.1-Base-details
Dataset Card for Evaluation run of sarvamai/OpenHathi-7B-Hi-v0.1-Base
Dataset automatically created during the evaluation run of model sarvamai/OpenHathi-7B-Hi-v0.1-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sarvamai__OpenHathi-7B-Hi-v0.1-Base-details.sequence-results-sarvam-mseattask-results-sarvam-mSarvam-m-Batch
