datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NSFW-Stories-JsonLConverted to JsonL from: bluuwhale/nsfwstory2
Nifty-Gay-Category
Siterip of articles from the LGBT literature digital archive nifty.org (excluding any beastiality).
Datasets is currently Work-In-Progress.
Nifty-Authoritarian-ScrapeData Scrape from LGBT Literature Archive Nifty.Org
-Category: Authoritarian
finance-legal-mrc_merged-table
데이터셋 설명
shchoice/finance-legal-mrc 데이터 중 병합된 테이블만 추출한 뒤 이미지와 함께 저장한 데이터입니다.
slm-bilanco-financial-qa-tr
slm-bilanco-financial-qa-tr
Turkce finansal analiz sorulari ve detayli cevaplarindan olusan instruction-tuning veri seti.
BIST sirketlerinin bilanco verileri uzerinden olusturulmustur.
Onemli Uyarilar
Bu veri seti yatirim tavsiyesi icermez. Analizler yalnizca egitim ve arastirma amaclidir.
Icerideki bilgiler ve bilanco degerleri gercek bilgiler olmayabilir. Yapay zeka tarafindan uretilmis veya sentetik olarak olusturulmus veriler icerebilir.
Herhangi bir yatirim karari almak icin… See the full description on the dataset page: https://huggingface.co/datasets/mrcuren/slm-bilanco-financial-qa-tr.finance-legal-mrc-chat-template
📊 Finance-Legal-MRC Chat Template
이 데이터셋은 문서 내 표 이미지와 문맥 정보를 기반으로, LLM이 description, title, summary, key_entities를 생성하도록 학습하기 위해 구성된 Chat Template 형식의 Instruction Dataset입니다.질문은 user가 주고, 정형화된 답변은 assistant가 응답하는 구조이며, 일부 메시지에는 표 이미지(image_url)가 포함됩니다.
📁 Dataset Overview
데이터 수: 1,197개 (train/test 8:2 split)
언어: Korean (한국어)
형식: OpenAI/ChatGPT-style messages 구조
입력 정보: 문맥(context) + 표 이미지(image_url)
출력 목표:
description: 자연어로 푼 설명
title: 표를 대표하는 제목
summary: 주요 요약… See the full description on the dataset page: https://huggingface.co/datasets/didi0di/finance-legal-mrc-chat-template.finance-legal-mrc-with-images
🧾 finance-legal-mrc-with-images (tableqa/test)
Multimodal-ready table image dataset designed for TIG (Table Information Generation) inputTIG 입력 전용 테이블 이미지 데이터셋 (VLM 활용 가능)
📦 Dataset Overview | 데이터셋 개요
Feature
Description (EN)
설명 (KR)
🧩 Split
tableqa / test
tableqa / test 스플릿
📄 Total Rows
1,197 (unique tables only)
총 1,197건 (중복 제거된 고유 테이블 기준)
🖼️ Image Format
PNG (rendered from raw HTML tables)
원본 HTML 테이블을 PNG 이미지로 렌더링
🔗 Source
Derived from… See the full description on the dataset page: https://huggingface.co/datasets/didi0di/finance-legal-mrc-with-images.klue-mrc-ko-rag-cot
데이터셋 설명
iamjoon/klue-mrc-ko-rag-dataset을 활용해 답변을 생성하는 과정을 CoT로 보강한 데이터셋입니다.
검색된 문서 수가 3개인 경우와 5개인 경우를 나눠서 데이터셋을 구성하였습니다.
데이터셋 구조
question: 사용자의 질문
search_result: 검색 결과
최소 1개~최대 5개까지 다양하게 구성.
answer: 사용자의 질문과 검색 결과를 바탕으로 답변합니다.
extracted_ref_numbers: 검색 결과 중 실제 정답으로 사용된 문서의 번호
최소 0개~최대 5개까지 다양하게 구성.
llm_result : 위 데이터를 활용해 LLM으로 생성한 답변.
구체적으로는 질문에 대해 검색된 문서들을 참고해 적절한 답변을 'answer' 블록에 생성하고, 그 과정을 추론하게 한 결과를 'reasoning' 블록에 작성, 최종적으로 답변 생성에 참고한 문서의 번호를 'doc_num'에 작성하도록… See the full description on the dataset page: https://huggingface.co/datasets/didi0di/klue-mrc-ko-rag-cot.SD-Prompt-DPOLuminous-OpusConverted from ChaoticNeutrals/Luminous_Opus
