CoolFace
20 results

corpora

instruction-pretrain /general-instruction-augmented-corpora Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.texttext-classification24 likes38k downloads7mo agoHugging FaceGuyHam /c-sac-corpora C-SAC LibriTTS-R training subset Deterministically selected and resampled speech from mythicinfinity/libritts_r for the C-SAC causal speech-codec program. The package retains source revision, Parquet shard, row, utterance, transcript, and content hashes. LibriTTS-R is distributed under CC BY 4.0; downstream users remain responsible for attribution. Only prefixes with a hash-bound _COMPLETE.json sentinel are admissible. audiotext-to-speech0 likes14k downloads2d agoHugging Facelucius1022 /DeMix_Corpora Dataset Card for DeMix Corpora DeMix 📄 Paper: Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training 🤗 Dataset: DeMix Corpora 🐱 Github: Demix Dataset Details Dataset Description DeMix Corpora (15T original tokens and 22T mixture tokens) serves as a comprehensive, high-quality, large-scale, and carefully mixed resource that can be directly employed for pre-training. (2026.2.7: This is an… See the full description on the dataset page: https://huggingface.co/datasets/lucius1022/DeMix_Corpora.tabularn<1K3 likes3.4k downloads7mo agoHugging FaceSotirisLegkas /kalamaki_corporatabular100M<n<1B0 likes2.4k downloads1y agoHugging Facecais /wmdp-corpora Dataset Card for WMDP Corpora The Weapons of Mass Destruction Proxy (WMDP) Corpora includes all of the corpora used to perform unlearning on WMDP-Bio and WMDP-Cyber. See our paper, website, and GitHub for more details! The corpora are also available at the following mirrors with password wmdpcorpora: 1, 2 The bio forget corpus must be requested separately; please visit this form. cyber-retain-corpus and cyber-forget-corpus The forget and retain corpora consist of… See the full description on the dataset page: https://huggingface.co/datasets/cais/wmdp-corpora.texttext-generation10K<n<100K5 likes2.3k downloads2y agoHugging Facemteb /legalbench_corporate_lobbying LegalBenchCorporateLobbying An MTEB dataset Massive Text Embedding Benchmark The dataset includes bill titles and bill summaries related to corporate lobbying. Task category t2t Domains Legal, Written Reference https://huggingface.co/datasets/nguha/legalbench/viewer/corporate_lobbying How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_corporate_lobbying.texttext-retrievaln<1K0 likes889 downloads7mo agoHugging Face