CoolFace
20 results

nix

brianmcgee /nix-cache-dataset Nix Cache Dataset This repository contains several datasets relating to the contents of https://cache.nixos.org: Getting started To make it easier to explore the datasets, a [Nix] devshell is provided in shell.nix. To enter it, run ❯ nix develop -f shell.nix [brian@saturn:~/Development/com/github/numtide/nix-cache-dataset]$ If you are a [Direnv] user, you can also run direnv allow to automatically load the devshell: ❯ direnv allow direnv: loading… See the full description on the dataset page: https://huggingface.co/datasets/brianmcgee/nix-cache-dataset.tabular1B<n<10B0 likes775 downloads8mo agoHugging FaceNix-ai /cat-v3xxxxl-plus 🐱 cat-v3xxxxl-plus (XXXXL-Plus) Part of the cat-v3 dataset family — synthetic instruction-tuning data that teaches models to be helpful, accurate, and delightfully cat-flavoured. About cat-v3xxxxl-plus (XXXXL-Plus) The largest variant in the cat-v3 family at 13,083,399 rows — 4.91725× the XXXXL dataset. Designed for pre-training-scale instruction tuning on the full topic distribution. Stored in snappy-compressed Parquet shards of 500,000 rows. Streaming is strongly… See the full description on the dataset page: https://huggingface.co/datasets/Nix-ai/cat-v3xxxxl-plus.texttext-generation10M<n<100M0 likes637 downloads6mo agoHugging Faceyamankushwah51 /nix-core-weights1 likes468 downloads5h agoHugging Facenixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes341 downloads2y agoHugging Facenixiesearch /ms-marco-dummy MS MARCO dummy+test dataset Used for testing nixietune: a dummy dataset of random 1000 queries from MS MARCO. The format is the following: { "query": ")what was the immediate impact of the success of the manhattan project?", "positive": [ "The presence of communication amid scientific minds was equally important to the success of the Manhattan Project as scientific intellect was. The only cloud hanging over the impressive achievement of the atomic researchers and engineers… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/ms-marco-dummy.textsentence-similarity1K<n<10K0 likes229 downloads3y agoHugging Facenixiesearch /bfhnd Big Hard Negatives Dataset A dataset for training embedding models for semantic search. TODO: add desc A dataset in a nixietune compatible format: { "query": ")what was the immediate impact of the success of the manhattan project?", "pos": [ "The presence of communication amid scientific minds was equally important to the success of the Manhattan Project as scientific intellect was. The only cloud hanging over the impressive achievement of the atomic researchers and… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/bfhnd.textsentence-similarity1M<n<10M1 likes179 downloads3y agoHugging Face