datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dlsite-jp-v1
puwaer/dlsite-jp-v1
This dataset consists of text extracted exclusively in Japanese from dlsite.com and is structured as JSON files. The files are categorized based on the type of URL.
このデータセットは、dlsite.comより日本語データのみを抽出したテキストで、jsonファイルで構成されます。
urlの種類によってファイル分けされています。
clara-stage2-data
Clara Stage 2 Training Data
Training data for Clara Stage 2 (Compression Instruction Tuning).
Dataset Description
This dataset contains high-quality QA pairs with single documents for training Clara's decoder adapter to generate answers from compressed document representations.
Data Format
Each record contains:
question: The query/question
answer: Gold answer
docs: List containing 1 document
meta: Source description
metadata: Additional metadata (repo, scope… See the full description on the dataset page: https://huggingface.co/datasets/dl3239491/clara-stage2-data.AIRC-DL-Intro
Dataset Card for Dataset Name
The dataset provides educator-generated multiple-choice quiz questions from lectures in real-world classrooms in Computer Science.
This is an subset containing the following course:
DL-Intro: an undergraduate-level course about various basic concepts and topics in Deep Learning.
Dataset Details
Uses
from datasets import load_dataset
data = load_dataset('mengxiayu/AIRC-DL-Intro', split='test')
print(data[0])
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mengxiayu/AIRC-DL-Intro.dltha_reasoning_v1.jsonl
DLTHA Reasoning Dataset v1
Description
This dataset is the first release from DLTHA Labs, focused on enhancing the logical reasoning and step-by-step problem-solving capabilities of Large Language Models (LLMs).
At DLTHA, we believe that the path to AGI (Artificial General Intelligence) requires high-fidelity synthetic data that mimics complex human thought processes. This dataset provides a structured "Chain-of-Thought" (CoT) format for technical and logical queries.… See the full description on the dataset page: https://huggingface.co/datasets/Dltha-Labs/dltha_reasoning_v1.jsonl.neocortirrhea-lexicon
Dataset Card: Neocortirrhea Lexicon Entry
Summary
This dataset entry defines and contextualizes the psychological, neurological, and somatic neologism Neocortirrhea.
Dataset Structure
JSON Lines Representation (data.jsonl)
{
"term": "Neocortirrhea",
"part_of_speech": "noun",
"phonetic": "/ˌniː.oʊˌkɔːr.tɪˈriː.ə/",
"etymology": "Neocortex (higher-order cognitive processing) + -rrhea (Greek rhoia: abnormal/excessive flow or… See the full description on the dataset page: https://huggingface.co/datasets/dlewicki/neocortirrhea-lexicon.fso-dataset-dlm
fso-dataset-dlm
Dataset generated by FusionX ModelsX.
Format: sharegpt
Samples: 806
Created: 2025-12-26T11:18:39.011263+00:00
Description
A dataset to fine tune a DLM on fso terms based on Financial Services Domain.
