CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Pclanglais /Nanochattext10M<n<100M8 likes7.9k downloads10mo agoHugging Face02msr-spare-1 /nemotron-3-nano-30b-20260719-spare-games-envs Nemotron-3-Nano-30B SPARE Self-Play Environments (run_20260719_final) This dataset packages the self-play generated game environments produced by a live SPARE (Self-Play with Adaptive cuRriculum Extension) training run of NVIDIA-Nemotron-3-Nano-30B-A3B. It is a raw-data export for another agent to pick up, replay, and build its own visualization / weave log from. Provenance Run: run_20260719_final Source Ray job: spare_nemotron_games_mtpg768_1784556397 (the live… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/nemotron-3-nano-30b-20260719-spare-games-envs.textn<1K0 likes6k downloads2mo agoHugging Face03sentence-transformers /NanoBEIR-entext10K<n<100K4 likes5.6k downloads10mo agoHugging Face04zeta-alpha-ai /NanoMSMARCOtexttext-retrieval1K<n<10K4 likes5.3k downloads2y agoHugging Face05zeta-alpha-ai /NanoQuoraRetrievaltexttext-retrieval1K<n<10K1 likes4.8k downloads2y agoHugging Face06hakari-bench /NanoMTEB-Scandinavian NanoMTEB-Scandinavian This dataset is a Nano-style retrieval dataset for HAKARI-bench. NanoMTEB-Scandinavian is a compact retrieval benchmark for Scandinavian-language MTEB-style task families. It includes Danish, Norwegian, and Swedish retrieval tasks spanning fact verification, question answering, news, encyclopedic content, FAQ retrieval, and social-media retrieval. Usage from datasets import load_dataset dataset_id = "hakari-bench/NanoMTEB-Scandinavian" split… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Scandinavian.text10K<n<100K0 likes3.9k downloads3mo agoHugging Face07Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes3.7k downloads1mo agoHugging Face08alexkstern /c4-nanochatbpe-10B c4-nanochatbpe-10B C4 (en) (from allenai/c4), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 10,000,000,000 val.bin val 168,272,017 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry the full metadata. The tokenizer/files… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/c4-nanochatbpe-10B.tabularn<1K0 likes3.2k downloads4mo agoHugging Face09zeta-alpha-ai /NanoFiQA2018texttext-retrieval1K<n<10K1 likes3.2k downloads2y agoHugging Face10alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.3k downloads3mo agoHugging Face11sentence-transformers-testing /NanoBEIR-detext10K<n<100K0 likes2.1k downloads10mo agoHugging Face12zeta-alpha-ai /NanoSciFacttexttext-retrieval1K<n<10K1 likes2.1k downloads2y agoHugging Face13zeta-alpha-ai /NanoNFCorpustexttext-retrieval1K<n<10K0 likes1.9k downloads2y agoHugging Face14nanotron /minipile_100_samplestextn<1K2 likes1.9k downloads2y agoHugging Face15zeta-alpha-ai /NanoNQtexttext-retrieval1K<n<10K0 likes1.8k downloads2y agoHugging Face16JoyboyBrian /nano-omni-vlmimage1M<n<10M0 likes1.7k downloads2y agoHugging Face17zeta-alpha-ai /NanoHotpotQAtexttext-retrieval1K<n<10K0 likes1.5k downloads2y agoHugging Face18zeta-alpha-ai /NanoArguAnatexttext-retrieval1K<n<10K0 likes1.4k downloads2y agoHugging Face19LiquidAI /nanobeir-multilingual-extended NanoBEIR Multilingual Extended Dataset This dataset extends the NanoBEIR multilingual collection with Japanese and Korean translations. Dataset Structure Each configuration follows the pattern <BASE>_<LANG> with splits: corpus: Document corpus queries: Search queries qrels: Query relevance judgments (when available) Languages Arabic (ar), German (de), English (en), Spanish (es), French (fr) Italian (it), Norwegian (no), Portuguese (pt), Swedish (sv)… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/nanobeir-multilingual-extended.text100K<n<1M10 likes1.3k downloads10mo agoHugging Face20huseinzol05 /mosaic-nanot5-512textn<1K0 likes1.3k downloads2y agoHugging Face21Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes1.2k downloads1mo agoHugging Face22zeta-alpha-ai /NanoDBPediatexttext-retrieval1K<n<10K0 likes1.1k downloads2y agoHugging Face23zeta-alpha-ai /NanoFEVERtexttext-retrieval1K<n<10K0 likes1.1k downloads2y agoHugging Face24inaciose /bagaco3_nanochatpt bagaco3_nanochatpt Dataset convertido a partir de duarteocarmo/bagaco3 para utilização no nanochat. Estatísticas Documentos: 34,234,368 Caracteres: 123,453,219,263 Caracteres/documento: 3,606.12 Tokens estimados: 29,393,623,634 Caracteres/token usados na estimativa: 4.2 Documentos vazios: 0 Mínimo de caracteres/documento: 47 Máximo de caracteres/documento: 64,123,869 Formato Formato: Parquet Coluna: text Rows por shard: 86,016 Rows por row group:… See the full description on the dataset page: https://huggingface.co/datasets/inaciose/bagaco3_nanochatpt.text10M<n<100M0 likes1.1k downloads2mo agoHugging Face25zeta-alpha-ai /NanoSCIDOCStexttext-retrieval1K<n<10K0 likes1.1k downloads2y agoHugging Face26zeta-alpha-ai /NanoClimateFEVERtexttext-retrieval1K<n<10K0 likes1.1k downloads2y agoHugging Face27zeta-alpha-ai /NanoTouche2020texttext-retrieval1K<n<10K0 likes1.1k downloads2y agoHugging Face28alexkstern /fineweb-nanochatbpe-20B fineweb-nanochatbpe-20B FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. file split tokens train.bin train 20,000,000,000 val.bin val 52,336,096 train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.tabularn<1K0 likes967 downloads4mo agoHugging Face29nanotron /picotron_bench Wrapup results: compute mfu for each results change status of jobs Push to hub add scripts reproductible add topology bandwidth etc tabularn<1K2 likes947 downloads2y agoHugging Face30hakari-bench /NanoMTEB-Dutch NanoMTEB-Dutch This dataset is a Nano-style retrieval dataset for HAKARI-bench. NanoMTEB-Dutch is a compact Dutch retrieval benchmark containing the MTEB-NL retrieval-family splits. It combines Dutch BEIR-style tasks, legal and public-domain QA, news, tender, web FAQ, Wikipedia, and cross-lingual Belebele retrieval splits. Usage from datasets import load_dataset dataset_id = "hakari-bench/NanoMTEB-Dutch" split = "argu_ana_nl" queries = load_dataset(dataset_id… See the full description on the dataset page: https://huggingface.co/datasets/hakari-bench/NanoMTEB-Dutch.text100K<n<1M0 likes872 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.