datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.nepse-market-data
NEPSE Market Data
Daily and tick-level data from the Nepal Stock Exchange, captured by an
open-source pipeline and validated at every layer boundary.
Daily prices reach back to 1995-07-20. Tick data starts 2026-08-19. That gap
is a property of the source, not a backlog - see Coverage below.
Tables
Path
Grain
Source
Coverage
eod_history_ext/
symbol x day
ShareSansar (second source)
1995-07-20 -> present, 653 symbols
eod_price/
symbol x day
NEPSE… See the full description on the dataset page: https://huggingface.co/datasets/AnkitSanjyal/nepse-market-data.cc100-nepali
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text derived from
CC-100, suitable for transformer pretraining.
Originally published at himalaya-ai/cc100-nepali.Dataset contents replaced with the cleaned version from Titung/cc100-nepali-cleaned.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Token Statistics (train split)
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/cc100-nepali.G1_Dex3_ToastedBread_DatasetThis dataset was created using LeRobot and converted to
3.0.
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in Part 5 of AVP Teleoperation Documentation.
Data collection is not completed in a single session, and variations between data entries exist. Ensure these variations are accounted for during model training.
Dataset Structure
meta/info.json:
{… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/G1_Dex3_ToastedBread_Dataset.can_clean_final_30fps
can_clean_final_30fps
Whole-body teleoperation on a Unitree G1, recorded 2026-08-25. The robot picks a can off a low
table and places it on a white table.
This is can_clean_final resampled
from 50 fps to 30 fps. Nothing else was changed: same 105 episodes, same task, same review, same
observation.state and action values on the frames that were kept.
episodes
105
frames
127,395 (was 212,290)
fps
30 (was 50)
duration
70.8 min (unchanged)
codebase version
v3.0… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/can_clean_final_30fps.nepali-tokenizer-corpuscan_clean_final
can_clean_final
Whole-body teleoperation on a Unitree G1, recorded 2026-08-25. The robot picks a can off a low
table and places it on a white table.
This is can_to_martino_2 and
can_to_martino_3 combined, with
every episode reviewed by hand and the failures removed. Prefer this over either source.
episodes
105
frames
212,290
fps
50
codebase version
v3.0
task
Bring the can to the white table
Schema
feature
dtype
shape
contents… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/can_clean_final.unitree_box_move_blue_fullThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unitree_g1",
"total_episodes": 550,
"total_frames": 475206,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:550"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/unitree_box_move_blue_full.can_to_martino_3
can_to_martino_3
Whole-body teleoperation on a Unitree G1, recorded the evening of 2026-08-25. The robot picks a
can off a low table and places it on a white table. Same rig, task and pipeline as
nepyope/can_to_martino_2, recorded
in three later sessions and released separately rather than merged into it.
episodes
29
frames
57,651
fps
50
codebase version
v3.0
task
Bring the can to the white table
Schema
Identical to can_to_martino_2, so… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/can_to_martino_3.can_to_martino
can_to_martino
Whole-body teleoperation on a Unitree G1: pick a can off a low table and hand it to a person.
Recorded 2026-08-23 with the SONIC teleop stack (PICO full-body tracking driving the balance
controller, Damiao CAN grippers on both hands), then reshaped into the state/action layout
π₀.₅ trains on.
robot
unitree_g1
episodes
44
frames
98,516
fps
50
task
Move the can from the low table to Martino
cameras
ego_view, left_wrist, right_wrist — 480×640… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/can_to_martino.can_to_martino_2
can_to_martino_2
Whole-body teleoperation on a Unitree G1, recorded 2026-08-25. The robot picks a can off a low
table and places it on a white table. Successor to
nepyope/can_to_martino (44 episodes),
recorded on the same rig with the same pipeline.
episodes
95
frames
187,224
fps
50
codebase version
v3.0
task
Bring the can to the white table
Schema
feature
dtype
shape
contents
observation.state
float32
[31]
29 body joints in… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/can_to_martino_2.nepal_flood_2026
Nepal Flood 2026, Upper Trishuli and Bhote Koshi Corridor
Building inventory for the 26 August 2026 flash flood on the Nepal-China border.
AOI: 1 km buffer around the Bhote Koshi and Trishuli river centrelines, 132.93 km² across Nuwakot and Rasuwa districts. Tasking Manager project 62904.
Contents
upperstream/buildings.geojson (+ .parquet): 13,663 footprints
upperstream/building_density_h3_r8.geojson (+ .parquet): buildings per H3 res 8 cell (~0.7 km²)… See the full description on the dataset page: https://huggingface.co/datasets/hotosm/nepal_flood_2026.ipfs_nepal_laws_ir
Nepal legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_nepal_laws (revision 22395665af98b03f562994b5b7aca768b7e34b54) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Nepal prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_nepal_laws_ir.nepali-honorific-benchnepali-textbooks-corpus
Nepali Textbooks Corpus for Grades 1-12
This dataset contains OCR-extracted, chapter-first, chunked text from Nepali school textbooks.
Summary
Samples: 5634
Grades: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]
Subjects: ['Civic_Education', 'Civic_Science', 'Economics', 'Education', 'Enterprenuership_and_Technology', 'Health_Physcial_and_Creative_Arts', 'Health_Physical_and_Creative_Arts', 'Health_and_Physical_Education', 'Math', 'My_Math', 'My_Nepali'… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali-textbooks-corpus.unitree_box_pushThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 16,
"total_frames": 65095,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/unitree_box_push.nepali-bias-dataset
Nepali Bias Language Dataset
Dataset Description
A synthetic dataset of Nepali sentences labeled for
bias categories including gender, religion, caste,
regional, appearance, social status, political, age,
and disability bias. Sentences were first labeled by
LLMs (ChatGPT, Grok) prompted with real Nepali news
context, then manually reviewed and corrected by human
annotators.
Dataset Summary
Split
Examples
Train
1,362
Validation
292… See the full description on the dataset page: https://huggingface.co/datasets/ios-ioe/nepali-bias-dataset.nepali-tokenizer-corpusaibharat_nepali_dataset_finalnepali-tokenizer-corpust-shirt_pick_and_place_cleanThis dataset was created using LeRobot.
Dataset Description
A filtered version of nepyope/t-shirt_pick_and_place
containing only the episodes that a human reviewer did not mark as failed.
47 episodes / 101,254 frames kept (from 79 episodes / 167,887 frames)
32 episodes dropped
Episodes are renumbered 0..46; all other fields, features and encoder settings are unchanged.
Statistics in meta/stats.json are recomputed over the kept episodes only.
Dropped episodes… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/t-shirt_pick_and_place_clean.gorkhapatra-nepali-epaper
Gorkhapatra Nepali E-Paper Corpus
Per-article text extracted from PDF e-papers published on
epaper.gorkhapatraonline.com, covering 11 newspaper
slugs (gorkhapatra, risingnepal, friday-suppliment, madhuparka, muna, nayanepal,
loksewa, saturday, yuwamunch, gorkhapatra-125, other).
Extraction is layout-aware (geometry + font size, no ML model/fixed template) and reconstructs
article boundaries — headline, dateline, and paragraphs in reading order — directly from the PDF's… See the full description on the dataset page: https://huggingface.co/datasets/Aananda-giri/gorkhapatra-nepali-epaper.universalml__NepaliGPT-2.0-details
Dataset Card for Evaluation run of universalml/NepaliGPT-2.0
Dataset automatically created during the evaluation run of model universalml/NepaliGPT-2.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/universalml__NepaliGPT-2.0-details.Sakalti__Neptuno-Alpha-details
Dataset Card for Evaluation run of Sakalti/Neptuno-Alpha
Dataset automatically created during the evaluation run of model Sakalti/Neptuno-Alpha
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Neptuno-Alpha-details.cc100-nepali-cleaned
CC-100 Nepali — Cleaned & Deduplicated
Cleaned, language-filtered, and deduplicated Nepali monolingual text from
CC-100 suitable for transformer pretraining.
Statistics
Split
Sentences
train
4,736,157
validation
48,328
test
48,329
total
4,832,814
Created: 2026-04-02
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact… See the full description on the dataset page: https://huggingface.co/datasets/Titung/cc100-nepali-cleaned.walk_back_and_forthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.ego_view": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/walk_back_and_forth.unitree_boxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unitree_g1",
"total_episodes": 4,
"total_frames": 304189,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/unitree_box.rakshak-nepali-toxicity-final
🛡️ RakshakAI — Augmented Nepali Toxicity Dataset
The full augmented training dataset used to train the RakshakAI toxicity detection models. Contains 4,716 samples expanded from the curated 1,574 sample dataset through back-translation augmentation via English and Hindi as intermediate languages. For the clean curated dataset only, see rakshak-all-data-combined.
📄 Paper: RakshakAI: Multi-Label Toxicity Detection for Low-Resource Nepali Social Media Content
Why this… See the full description on the dataset page: https://huggingface.co/datasets/biraj-bhusal/rakshak-nepali-toxicity-final.t-shirt_pick_and_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 50,
"features": {
"observation.images.ego_view": {
"dtype": "video",
"shape": [
480,
640,
3
],
"names": [
"height",
"width",
"channel"
],
"info": {… See the full description on the dataset page: https://huggingface.co/datasets/nepyope/t-shirt_pick_and_place.gemma4-e2b-nepali-sft-pairs
Nepali SFT pairs for Gemma 4 E2B
468 (English prompt -> Nepali answer) pairs, the exact training data behind
saliltambe/gemma-4-E2B-it-nepali-lora.
Published so the training notebook can skip a ~13 minute generation step and so anyone
reproducing it evaluates on the same held-out split.
Provenance
Prompts: English conversation openers from
OpenAssistant/oasst1 (Apache-2.0,
human-written), filtered to role == "prompter", parent_id is None, lang == "en".
Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.
