CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sevenc-nanashi /kiiteitte Kiiteitte history Kiiteitte が収集した、今までの選曲履歴。 1時間おきに更新されます。 型 { // 動画ID "video_id": "sm44670499", // タイトル "title": "library->w4nderers / 足立レイ、つくよみちゃん", // 投稿者 "author": "名無し。", // サムネイルのURL "thumbnail": "https://nicovideo.cdn.nimg.jp/thumbnails/44670499/44670499.91820835", // 選曲日時 "date": "2025-02-22 12:51:51", // 新しく増えたお気に入り数。不明の場合は null "new_faves": 5, // 回ったユーザーの数。不明の場合は null "spins": 13, // イチ押しリストのユーザーのURL。イチ押しリスト以外から選曲された場合は null… See the full description on the dataset page: https://huggingface.co/datasets/sevenc-nanashi/kiiteitte.image100K<n<1M2 likes4.3k downloads30m agoHugging Face02alexkstern /c4-nanochatbpe-10B c4-nanochatbpe-10B C4 (en) (from allenai/c4), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 10,000,000,000 val.bin val 168,272,017 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry the full metadata. The tokenizer/files… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/c4-nanochatbpe-10B.tabularn<1K0 likes3k downloads4mo agoHugging Face03alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.2k downloads3mo agoHugging Face04ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes1.3k downloads3mo agoHugging Face05nanotron /picotron_bench Wrapup results: compute mfu for each results change status of jobs Push to hub add scripts reproductible add topology bandwidth etc tabularn<1K2 likes1k downloads2y agoHugging Face06ray0rf1re /FineWeb-Nano FineWeb-Nano Dataset Description FineWeb-Nano is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb. How "The Best" Was Determined This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors: High language_score (if provided by the upstream extraction). Optimal document length (penalizing abnormally short snippets and excessively… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/FineWeb-Nano.tabular1M<n<10M0 likes878 downloads6mo agoHugging Face07alexkstern /fineweb-nanochatbpe-20B fineweb-nanochatbpe-20B FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 20,000,000,000 val.bin val 52,336,096 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.tabularn<1K0 likes866 downloads4mo agoHugging Face08joey234 /nan-nli Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards Natural Language Inference Text Classification Languages en Dataset Structure Data Instances Data Fields premise: hypothesis: label: Data Splits Evaluation: 258 samples Dataset Creation Curation Rationale Extracting samples corresponding to different linguistics constructions of… See the full description on the dataset page: https://huggingface.co/datasets/joey234/nan-nli.tabulartext-classificationn<1K0 likes818 downloads4y agoHugging Face09czl /nangang_sports_centertabular100K<n<1M0 likes689 downloads15h agoHugging Face10ddudek /nanochat-climbmix-hq Nanochat climbmix dataset filtered This repository contains a filtered version of the climbmix dataset for efficient use with Andrej Karpathy’s Nanochat project. tabular10M<n<100M0 likes664 downloads6mo agoHugging Face11nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes445 downloads7d agoHugging Face12ChrisHayduk /nanofold-public NanoFold Public NanoFold Public is the public train/validation portion of the nanoFold protein-folding benchmark. It packages a compact, fixed, auditable subset of OpenProteinSet/OpenFold-derived protein structure training data for fast iteration on data-efficient folding models. The dataset has 10000 train chains and 1000 public validation chains. Each row is one protein chain. The original processed .npz tensors are unrolled into Hugging Face Dataset columns so users can load… See the full description on the dataset page: https://huggingface.co/datasets/ChrisHayduk/nanofold-public.tabularfeature-extraction10K<n<100K17 likes392 downloads3mo agoHugging Face13maikezu /f-actor-behavior-sd-nanocodec F-Actor Nano-Codec Dataset This repository contains the data accompanying the paper F-Actor: Controllable Conversational Behaviour in Full-Duplex Models. The data consists of the Behavior-SD dataset, encoded using nvidia/nemo-nano-codec-22khz-0.6kbps-12.5fps, and augmented with a different narrative. About our work: Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/f-actor-behavior-sd-nanocodec.tabular100K<n<1M1 likes375 downloads8mo agoHugging Face14ddudek /nanochat-climbmix-annotated Summary A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats. Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code. Dataset Structure Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page: https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.tabulartext-classification10M<n<100M0 likes355 downloads6mo agoHugging Face15Ruth011 /jetson_orin_nano_super_1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 1, "total_frames": 590, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Ruth011/jetson_orin_nano_super_1.tabularrobotics10K<n<100K0 likes267 downloads1y agoHugging Face16fffoivos /glossapi-greek-nanochat-pretraining-dataset-v2 GlossAPI Greek pretraining corpus v2 HPLT filtering method The HPLT component is HPLT/ell_Grek_ge8_no_mt_clean60. It retains HPLT quality bins 8, 9, 10 (GE8), uses the pre-applied no-MT/register filter, requires greek_badness_score <= 60, and applies Wave4 Greek re-cleaning and normalization. The standalone filtered slice contained 48,728,774 documents; 48,629,460 remain after corpus-wide deduplication. GlossAPI datasets and token counts GlossAPI… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset-v2.tabular10M<n<100M0 likes254 downloads1mo agoHugging Face17NanaSomuah0233 /DIF-Dataset-Comparison DIF Dataset Comparison & Collected Data A curated collection of educational assessment datasets evaluated for Differential Item Functioning (DIF) analysis, with actual downloaded data where available. What's Here 📊 Comparison Spreadsheets File Description DIF_Dataset_Comparison_VERIFIED.xlsx Main Excel with 3 sheets: Master Comparison, Column Inventory, DIF Readiness Checklist DIF_Dataset_Master_Comparison.csv Dataset-level comparison (responses… See the full description on the dataset page: https://huggingface.co/datasets/NanaSomuah0233/DIF-Dataset-Comparison.tabular100K<n<1M0 likes252 downloads5mo agoHugging Face18PureOne /R3C-Universal-Nanofabrication R3C — Reservoir-Rank and Reaction-Repair Compiler Reservoir-rank engineering and finite-bandwidth reaction repair toward programmable nanofabrication Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiVersion: 1.0.0 — 17 September 2026Repository: PureOne/R3C-Universal-NanofabricationResource type: theoretical research report + reproducible software + entirely synthetic datasetsScientific status: conditional finite-model theory; no physical… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/R3C-Universal-Nanofabrication.documentn<1K0 likes231 downloads6d agoHugging Face19Rcarvalo /kanitts2-fr-nanocodectabular100K<n<1M0 likes215 downloads6mo agoHugging Face20arpandeepk /generations-nemotron-nano-9b-v2-simnpo-gentle-igm-10btabular10K<n<100K0 likes164 downloads5mo agoHugging Face21PureOne /stoichforge-universal-nanofabrication-v1 STOICHFORGE Deferred-Dissipation Reaction Compilation for Universal Nanofabrication Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific version: v1.0.0 · Hub distribution: hf.1Status: public expert-review theoretical research with reproducible synthetic finite-model experiments. Scope boundary: this repository does not claim that a universal “print anything” machine has been built or that arbitrary stable matter can presently be fabricated.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/stoichforge-universal-nanofabrication-v1.tabularothern<1K0 likes163 downloads4d agoHugging Face22ASSERT-KTH /Nano-SFT-SWE-Gym-gemini-2.5-flashtabular1K<n<10K1 likes161 downloads1y agoHugging Face23twinkle-ai /nemotron-nano-eval-logs-and-scorestabular100K<n<1M0 likes160 downloads7mo agoHugging Face24unlearning-cleanslate /generations-nemotron-nano-9b-v2-simnpo-gentle-baselinetabular10K<n<100K0 likes139 downloads5mo agoHugging Face25nanskong /ManipuriGPT-Corpus-v1.0 ManipuriGPT Corpus v1.0 ManipuriGPT Corpus v1.0 is a research-grade, multi-script, deduplicated, and quality-scored corpus specifically engineered for pretraining Manipuri (Meiteilon) language foundation models. Quick Summary Total Sequences: 147,956 Total Tokens (ManipuriGPT-Tokenizer-v1.0): 4,347,075 Total Characters: 16,019,401 Pipeline Version: 5.6 Release Version: v1.0.0 Build Timestamp: 2026-07-25T09:17:21.960438Z Primary Writing Systems… See the full description on the dataset page: https://huggingface.co/datasets/nanskong/ManipuriGPT-Corpus-v1.0.tabulartext-generation100K<n<1M0 likes138 downloads2mo agoHugging Face26EliasHossain /nanobubbleeval NanoBubbleEval v1.0 ⚠ For NeurIPS reviewers — use this Croissant URL Please do NOT use the URL exposed by the "Use this dataset → Croissant" button at the top-right of this page. That URL triggers a known bug in mlcroissant==1.0.16 (the version pinned by the NeurIPS Croissant validator Space) and produces a FilterFiles error that does not reflect a problem with the dataset itself. Use this URL instead — copy the line below verbatim into the validator's "URL Input" tab:… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/nanobubbleeval.tabularquestion-answering10K<n<100K0 likes124 downloads5mo agoHugging Face27nekko-nanode /Lerobot_datatabular10K<n<100K0 likes123 downloads6d agoHugging Face28unlearning-cleanslate /generations-nemotron-nano-9b-v2-simnpo-baselinetabular10K<n<100K0 likes121 downloads5mo agoHugging Face29unlearning-cleanslate /generations-nemotron-nano-9b-v2-pre_valtabular10K<n<100K0 likes120 downloads5mo agoHugging Face30nyu-dice-lab /lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private Dataset Card for Evaluation run of Kquant03/Nanashi-2x7B-bf16 Dataset automatically created during the evaluation run of model Kquant03/Nanashi-2x7B-bf16 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private.tabular100K<n<1M0 likes112 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.