CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01maximoss /follie Dataset Card for FOLLIE dataset Dataset Details Dataset Description FOLLIE (First-Order Logic for Language Inference and Entailment) is the first dataset for French with natural language sentences and their corresponding first-order logic (FOL) formulas. The sentences in the dataset were drawn from all the existing French Natural Language Inference (NLI) datasets, namely DACCORD, FraCaS-FR, GQNLI-FR, RTE3-FR (dev and test), SICK-FR, and XNLI… See the full description on the dataset page: https://huggingface.co/datasets/maximoss/follie.texttext-classification10K<n<100K0 likes276 downloads16d agoHugging Face02Lots-of-LoRAs /task1148_maximum_ascii_value Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1148_maximum_ascii_value Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1148_maximum_ascii_value.texttext-generationn<1K0 likes191 downloads2y agoHugging Face03Maximiliano-Flores-Dev /agentlans-combined-roleplay_Dataset Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.texttext-generation1M<n<10M0 likes64 downloads2d agoHugging Face04Maximiliano-Flores-Dev /QuixiAI-dolphin_DatasetDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.texttext-generation1M<n<10M1 likes55 downloads2d agoHugging Face05Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes49 downloads2d agoHugging Face06maximuspowers /muat-pca-10-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods pca Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 8 to 10 Neurons per… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-pca-10-medium.texttext-generation10K<n<100K0 likes45 downloads10mo agoHugging Face07MaximoLopezChenlo /OncoAgent-Clinical-266K 🧬 OncoAgent Clinical Dataset — 266K Curated Multi-Source Oncology Training Dataset AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0 Dataset Description This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks. Composition Source Samples Description PMC-Patients ~100,000 Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/MaximoLopezChenlo/OncoAgent-Clinical-266K.texttext-generation100K<n<1M0 likes44 downloads5mo agoHugging Face08maximuspowers /muat-fourier-5-large Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods fourier Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 8 to 10 Neurons… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-fourier-5-large.texttext-generation10K<n<100K0 likes31 downloads10mo agoHugging Face09maximuspowers /muat-pca-15 Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods pca Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6 Neurons… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-pca-15.texttext-generation1K<n<10K0 likes28 downloads10mo agoHugging Face10maximuspowers /muat-pca-5 Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods pca Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6 Neurons… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-pca-5.texttext-generation10K<n<100K0 likes26 downloads10mo agoHugging Face11maximuspowers /muat-mean-std Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std.texttext-generation10K<n<100K0 likes25 downloads10mo agoHugging Face12maximuspowers /muat-pca-10 Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods pca Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6 Neurons… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-pca-10.texttext-generation1K<n<10K0 likes24 downloads10mo agoHugging Face13maximuspowers /muat-mean-std-pca-10-fourier-5 Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std, pca, fourier Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-pca-10-fourier-5.texttext-generation10K<n<100K0 likes24 downloads10mo agoHugging Face14Maximofn /short-jokes-dataset Dataset Card for "short-jokes-dataset" Dataset from amoudgl short-jokes-dataset More Information needed texttext-generation100K<n<1M3 likes23 downloads2y agoHugging Face15maximuspowers /muat-mean-std-fourier-5-pca-10-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std, pca, fourier Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-fourier-5-pca-10-medium.texttext-generation10K<n<100K0 likes22 downloads10mo agoHugging Face16maximuspowers /muat-fourier-5-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods fourier Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 6 to 8 Neurons… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-fourier-5-medium.texttext-generation10K<n<100K0 likes19 downloads10mo agoHugging Face17maximuspowers /muat-sigs-with-input-correlations Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std, fourier, input_correlations, pre_activation_mean, pre_activation_std Prompt Format separate Signature Dataset… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-sigs-with-input-correlations.texttext-generation10K<n<100K0 likes19 downloads8mo agoHugging Face18Maximilianzxp /wikipedia-cn-20230720-filtered本数据集基于中文维基2023年7月20日的dump存档。作为一项以数据为中心的工作,本数据集仅保留了 254,547条 质量较高的词条内容。具体而言: 过滤了Template, Category, Wikipedia, File, Topic, Portal, MediaWiki, Draft, Help等特殊类型的词条 使用启发式的方法和自有的NLU模型过滤了一部分质量较低的词条 过滤了一部分内容较为敏感或存在争议性的词条。 进行了简繁转换和习惯用词转换,确保符合中国大陆地区的习惯用词。 This dataset is based on the Chinese Wikipedia dump archive from July 20th, 2023. As a data-centric effort, the dataset retains 254,574 high-quality entries. Specifically: Entries of special types such as Template, Category, Wikipedia, File, Topic… See the full description on the dataset page: https://huggingface.co/datasets/Maximilianzxp/wikipedia-cn-20230720-filtered.texttext-generation100K<n<1M0 likes18 downloads5mo agoHugging Face19maximuspowers /muat-mean-std-medium Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 6 to 8… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-medium.texttext-generation10K<n<100K0 likes17 downloads10mo agoHugging Face20maximuspowers /hypernet_validated Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std, fourier, input_correlations, pre_activation_mean, pre_activation_std Prompt Format separate Signature Dataset… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/hypernet_validated.texttext-generation10K<n<100K0 likes17 downloads7mo agoHugging Face21Nick-Maximillien /medgate-compiler-data MedGate Compiler Training Data The Logic-Grounding Corpus for Computable Medical Law This dataset is a hand-curated collection of 606 high-fidelity mappings designed to train Structural Compilers. It facilitates the translation of unstructured clinical guidelines and medical policy prose into machine-executable symbolic logic (JSON). Dataset Summary The MedGate Compiler Training Data provides the ground truth for Nexus Forensic – Layer 0 (Protocol Vault). Each record… See the full description on the dataset page: https://huggingface.co/datasets/Nick-Maximillien/medgate-compiler-data.texttext-generationn<1K0 likes14 downloads7mo agoHugging Face22maximuspowers /muat-fourier-10 Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods fourier Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-fourier-10.texttext-generation10K<n<100K0 likes11 downloads10mo agoHugging Face23maximuspowers /muat-fourier-3 Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods fourier Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-fourier-3.texttext-generation10K<n<100K0 likes11 downloads10mo agoHugging Face24maximuspowers /muat-mean-std-large Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods mean, std Prompt Format separate Signature Dataset configs/dataset_gen/signature_dataset.json Model Architecture Number of Layers 8 to 10… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-mean-std-large.texttext-generation1K<n<10K0 likes10 downloads10mo agoHugging Face25maximuspowers /muat-fourier-5 Subject Models for Interpretability Training These examples are intended for training an interpreter to: Identify what patterns a model classifies as positive based on an activation signature, with examples of: trained model + signature → pattern identification. Signature Extraction Neuron Profile Methods fourier Prompt Format separate Signature Dataset dataset_generation/exp_1/signature_dataset.json Model Architecture Number of Layers 4 to 6… See the full description on the dataset page: https://huggingface.co/datasets/maximuspowers/muat-fourier-5.texttext-generation10K<n<100K0 likes8 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.