datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kiiteitte
Kiiteitte history
Kiiteitte が収集した、今までの選曲履歴。
1時間おきに更新されます。
型
{
// 動画ID
"video_id": "sm44670499",
// タイトル
"title": "library->w4nderers / 足立レイ、つくよみちゃん",
// 投稿者
"author": "名無し。",
// サムネイルのURL
"thumbnail": "https://nicovideo.cdn.nimg.jp/thumbnails/44670499/44670499.91820835",
// 選曲日時
"date": "2025-02-22 12:51:51",
// 新しく増えたお気に入り数。不明の場合は null
"new_faves": 5,
// 回ったユーザーの数。不明の場合は null
"spins": 13,
// イチ押しリストのユーザーのURL。イチ押しリスト以外から選曲された場合は null… See the full description on the dataset page: https://huggingface.co/datasets/sevenc-nanashi/kiiteitte.c4-nanochatbpe-10B
c4-nanochatbpe-10B
C4 (en) (from allenai/c4), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
10,000,000,000
val.bin
val
168,272,017
train and val are disjoint held-out partitions. Each .bin is a raw
little-endian uint16 stream (no header); token count = filesize / 2, and
train.meta.json / val.meta.json carry the full metadata. The tokenizer/files… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/c4-nanochatbpe-10B.fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.Pluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.picotron_bench
Wrapup results:
compute mfu for each results
change status of jobs
Push to hub
add scripts reproductible
add topology
bandwidth etc
FineWeb-Nano
FineWeb-Nano
Dataset Description
FineWeb-Nano is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb.
How "The Best" Was Determined
This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors:
High language_score (if provided by the upstream extraction).
Optimal document length (penalizing abnormally short snippets and excessively… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/FineWeb-Nano.fineweb-nanochatbpe-20B
fineweb-nanochatbpe-20B
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
20,000,000,000
val.bin
val
52,336,096
train and val are disjoint held-out partitions. Each .bin is a raw
little-endian uint16 stream (no header); token count = filesize / 2, and
train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.nan-nli
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
Natural Language Inference
Text Classification
Languages
en
Dataset Structure
Data Instances
Data Fields
premise:
hypothesis:
label:
Data Splits
Evaluation: 258 samples
Dataset Creation
Curation Rationale
Extracting samples corresponding to different linguistics constructions of… See the full description on the dataset page: https://huggingface.co/datasets/joey234/nan-nli.nangang_sports_centernanochat-climbmix-hq
Nanochat climbmix dataset filtered
This repository contains a filtered version of the climbmix dataset for efficient use with Andrej Karpathy’s Nanochat project.
local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.nanofold-public
NanoFold Public
NanoFold Public is the public train/validation portion of the nanoFold protein-folding benchmark. It packages a compact, fixed, auditable subset of OpenProteinSet/OpenFold-derived protein structure training data for fast iteration on data-efficient folding models.
The dataset has 10000 train chains and 1000 public validation chains. Each row is one protein chain. The original processed .npz tensors are unrolled into Hugging Face Dataset columns so users can load… See the full description on the dataset page: https://huggingface.co/datasets/ChrisHayduk/nanofold-public.f-actor-behavior-sd-nanocodec
F-Actor Nano-Codec Dataset
This repository contains the data accompanying the paper
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models.
The data consists of the Behavior-SD dataset, encoded using nvidia/nemo-nano-codec-22khz-0.6kbps-12.5fps, and augmented with a different narrative.
About our work:
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-nanocodec.nanochat-climbmix-annotated
Summary
A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats.
Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code.
Dataset Structure
Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page: https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.jetson_orin_nano_super_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 1,
"total_frames": 590,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Ruth011/jetson_orin_nano_super_1.glossapi-greek-nanochat-pretraining-dataset-v2
GlossAPI Greek pretraining corpus v2
HPLT filtering method
The HPLT component is HPLT/ell_Grek_ge8_no_mt_clean60. It retains HPLT quality bins 8, 9, 10 (GE8), uses the pre-applied no-MT/register filter, requires greek_badness_score <= 60, and applies Wave4 Greek re-cleaning and normalization. The standalone filtered slice contained 48,728,774 documents; 48,629,460 remain after corpus-wide deduplication.
GlossAPI datasets and token counts
GlossAPI… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset-v2.DIF-Dataset-Comparison
DIF Dataset Comparison & Collected Data
A curated collection of educational assessment datasets evaluated for Differential Item Functioning (DIF) analysis, with actual downloaded data where available.
What's Here
📊 Comparison Spreadsheets
File
Description
DIF_Dataset_Comparison_VERIFIED.xlsx
Main Excel with 3 sheets: Master Comparison, Column Inventory, DIF Readiness Checklist
DIF_Dataset_Master_Comparison.csv
Dataset-level comparison (responses… See the full description on the dataset page: https://huggingface.co/datasets/NanaSomuah0233/DIF-Dataset-Comparison.R3C-Universal-Nanofabrication
R3C — Reservoir-Rank and Reaction-Repair Compiler
Reservoir-rank engineering and finite-bandwidth reaction repair toward programmable nanofabrication
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiVersion: 1.0.0 — 17 September 2026Repository: PureOne/R3C-Universal-NanofabricationResource type: theoretical research report + reproducible software + entirely synthetic datasetsScientific status: conditional finite-model theory; no physical… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/R3C-Universal-Nanofabrication.kanitts2-fr-nanocodecgenerations-nemotron-nano-9b-v2-simnpo-gentle-igm-10bstoichforge-universal-nanofabrication-v1
STOICHFORGE
Deferred-Dissipation Reaction Compilation for Universal Nanofabrication
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific version: v1.0.0 · Hub distribution: hf.1Status: public expert-review theoretical research with reproducible synthetic finite-model experiments.
Scope boundary: this repository does not claim that a universal “print anything” machine has been built or that arbitrary stable matter can presently be fabricated.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/stoichforge-universal-nanofabrication-v1.Nano-SFT-SWE-Gym-gemini-2.5-flashnemotron-nano-eval-logs-and-scoresgenerations-nemotron-nano-9b-v2-simnpo-gentle-baselineManipuriGPT-Corpus-v1.0
ManipuriGPT Corpus v1.0
ManipuriGPT Corpus v1.0 is a research-grade, multi-script, deduplicated, and quality-scored corpus specifically engineered for pretraining Manipuri (Meiteilon) language foundation models.
Quick Summary
Total Sequences: 147,956
Total Tokens (ManipuriGPT-Tokenizer-v1.0): 4,347,075
Total Characters: 16,019,401
Pipeline Version: 5.6
Release Version: v1.0.0
Build Timestamp: 2026-07-25T09:17:21.960438Z
Primary Writing Systems… See the full description on the dataset page: https://huggingface.co/datasets/nanskong/ManipuriGPT-Corpus-v1.0.nanobubbleeval
NanoBubbleEval v1.0
⚠ For NeurIPS reviewers — use this Croissant URL
Please do NOT use the URL exposed by the "Use this dataset → Croissant"
button at the top-right of this page. That URL triggers a known bug in
mlcroissant==1.0.16 (the version pinned by the
NeurIPS Croissant validator Space)
and produces a FilterFiles error that does not reflect a problem with the
dataset itself.
Use this URL instead — copy the line below verbatim into the validator's
"URL Input" tab:… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/nanobubbleeval.Lerobot_datagenerations-nemotron-nano-9b-v2-simnpo-baselinegenerations-nemotron-nano-9b-v2-pre_vallm-eval-results-Kquant03-Nanashi-2x7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Nanashi-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Nanashi-2x7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private.
