CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01puwaer /dlsite-jp-v1 puwaer/dlsite-jp-v1 This dataset consists of text extracted exclusively in Japanese from dlsite.com and is structured as JSON files. The files are categorized based on the type of URL. このデータセットは、dlsite.comより日本語データのみを抽出したテキストで、jsonファイルで構成されます。 urlの種類によってファイル分けされています。 texttext-generation1M<n<10M4 likes174 downloads2y agoHugging Face02dl3239491 /clara-stage2-data Clara Stage 2 Training Data Training data for Clara Stage 2 (Compression Instruction Tuning). Dataset Description This dataset contains high-quality QA pairs with single documents for training Clara's decoder adapter to generate answers from compressed document representations. Data Format Each record contains: question: The query/question answer: Gold answer docs: List containing 1 document meta: Source description metadata: Additional metadata (repo, scope… See the full description on the dataset page: https://huggingface.co/datasets/dl3239491/clara-stage2-data.textquestion-answering1K<n<10K0 likes26 downloads8mo agoHugging Face03mengxiayu /AIRC-DL-Intro Dataset Card for Dataset Name The dataset provides educator-generated multiple-choice quiz questions from lectures in real-world classrooms in Computer Science. This is an subset containing the following course: DL-Intro: an undergraduate-level course about various basic concepts and topics in Deep Learning. Dataset Details Uses from datasets import load_dataset data = load_dataset('mengxiayu/AIRC-DL-Intro', split='test') print(data[0]) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mengxiayu/AIRC-DL-Intro.texttext-generationn<1K0 likes8 downloads1y agoHugging Face04Dltha-Labs /dltha_reasoning_v1.jsonl DLTHA Reasoning Dataset v1 Description This dataset is the first release from DLTHA Labs, focused on enhancing the logical reasoning and step-by-step problem-solving capabilities of Large Language Models (LLMs). At DLTHA, we believe that the path to AGI (Artificial General Intelligence) requires high-fidelity synthetic data that mimics complex human thought processes. This dataset provides a structured "Chain-of-Thought" (CoT) format for technical and logical queries.… See the full description on the dataset page: https://huggingface.co/datasets/Dltha-Labs/dltha_reasoning_v1.jsonl.texttext-generationn<1K1 likes8 downloads8mo agoHugging Face05dlewicki /neocortirrhea-lexicon Dataset Card: Neocortirrhea Lexicon Entry Summary This dataset entry defines and contextualizes the psychological, neurological, and somatic neologism Neocortirrhea. Dataset Structure JSON Lines Representation (data.jsonl) { "term": "Neocortirrhea", "part_of_speech": "noun", "phonetic": "/ˌniː.oʊˌkɔːr.tɪˈriː.ə/", "etymology": "Neocortex (higher-order cognitive processing) + -rrhea (Greek rhoia: abnormal/excessive flow or… See the full description on the dataset page: https://huggingface.co/datasets/dlewicki/neocortirrhea-lexicon.texttext-generationn<1K0 likes6 downloads1mo agoHugging Face06Hardeep /fso-dataset-dlm fso-dataset-dlm Dataset generated by FusionX ModelsX. Format: sharegpt Samples: 806 Created: 2025-12-26T11:18:39.011263+00:00 Description A dataset to fine tune a DLM on fso terms based on Financial Services Domain. texttext-generationn<1K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.