CoolFace
21 results

star

bigcode /starcoderdatagated StarCoder Training Dataset Dataset description This is the dataset used for training StarCoder and StarCoderBase. It contains 783GB of code in 86 programming languages, and includes 54GB GitHub Issues + 13GB Jupyter notebooks in scripts and text-code pairs, and 32GB of GitHub commits, which is approximately 250 Billion tokens. Dataset creation The creation and filtering of The Stack is explained in the original dataset, we additionally decontaminate and… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoderdata.texttext-generation100M<n<1B545 likes42k downloads3y agoHugging FaceLab-MSP /libritts-r-stark1 likes13k downloads2mo agoHugging Facesonglab /gpn-star-scores GPN-Star genome-wide scores Genome-wide, mutation-rate-calibrated GPN-Star constraint and variant scores for eight score sets covering human, mouse, chicken, D. melanogaster, C. elegans, and A. thaliana. Canonical scores are available as chromosome-sharded Parquet. Hugging Face hosts 72 BigWigs; a multi-assembly UCSC track hub references the 64 logo/LLR views. Overview and quick links Resource Link Files Browse all dataset files Default Dataset… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-star-scores.tabular10B<n<100B4 likes12k downloads2mo agoHugging Facegeodesic-research /pa-warm-start-sft-heavy-25b-mix geodesic-research/pa-warm-start-sft-heavy-25b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-heavy-25b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-heavy-25b-mix.tabular10M<n<100M0 likes8.1k downloads21d agoHugging Facestanford-star /the-join The Join A broad collection of relational databases spanning many domains (academic, e-commerce, finance, sports, biomedical, government, text2sql, and more), ported to the RelBench manifest format. The Join is built for pretraining relational/tabular foundation models: each database is self-describing and tasks ship labels as-is for large-scale pretraining rather than held-out benchmarking. Each dataset lives in its own subdirectory in the self-describing manifest layout (plain… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/the-join.tabular10K<n<100K5 likes7.6k downloads2mo agoHugging Facestanford-star /relbench-v1 RelBench v1 databases The original RelBench v1 relational databases and tasks, each database in its own subdirectory in the self-describing manifest layout (plain parquet + manifest.yaml): <dataset>/ manifest.yaml # tables, primary keys, foreign-key graph, val/test timestamps schema.svg # ER diagram db/*.parquet # relational tables (plain parquet) tasks/<task>/manifest.yaml # task spec (+ duckdb SQL for `forecast`… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/relbench-v1.tabularn<1K1 likes7.5k downloads28d agoHugging Face

Projects