datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sage-pretrain-corpus
Sage Pretrain Corpus (v0.8)
License-audited Korean/English pretraining corpus for the Sage Korean LLM project.
Records: id, text, source, license, meta (JSON of original fields).
⚠️ Mixed licenses — license: other. Comply with each source's license individually. CC-BY / ODC-BY require attribution; CC-BY-SA carries ShareAlike.
Sources (v0.8)
source
docs
license
origin
fineweb2_ko
46,470,574
ODC-BY-1.0
FineWeb-2 Korean (HuggingFaceFW/fineweb-2 kor_Hang)… See the full description on the dataset page: https://huggingface.co/datasets/seongchaeae/sage-pretrain-corpus.korean-assembly-minutes
대한민국 국회 회의록 아카이브
국회 회의록 원문(record.assembly.go.kr) PDF 에서 본문을 뽑아 모은 것이다.
본회의와 각 위원회 회의록이 모두 들어 있다.
수록 기간: 1948~1993
회의 수: 1,951건
본문 분량: 65,306,444자
구성
연도별 JSONL(gzip) 한 덩이다.
from datasets import load_dataset
ds = load_dataset("seoulraphaellee/korean-assembly-minutes", split="train")
ds = load_dataset("seoulraphaellee/korean-assembly-minutes", data_files="data/2026.jsonl.gz", split="train")
필드
이름
설명
meeting_key
회의 식별자 (record:<id>)… See the full description on the dataset page: https://huggingface.co/datasets/seoulraphaellee/korean-assembly-minutes.global-seo-knowledgeCHUNKY-tulu3-SFT-25k-attributes-full
SURF Attributes (Full)
Complete dataset for SURF research and extension.
Paper: Chunky Post-Training
Quick Start
For running SURF, use the minimal dataset: seoirsem/CHUNKY-tulu3-SFT-25k-attributes
uv run -m surf.cli.main sweep \
--attributes seoirsem/CHUNKY-tulu3-SFT-25k-attributes \
--rubric rubrics/rebuttal.yaml \
-o results/
Dataset Fields
prompt: The query text
response: The model response (if available)
attributes: Raw extracted attributes… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/CHUNKY-tulu3-SFT-25k-attributes-full.MAVIS
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
📖 Paper | 💻 Evaluation
Dataset Summary
MAVIS is a new dataset for open-domain, long-form visual question answering, characterized by three key features: (1) the questions incorporate input images, requiring visual understanding to correctly interpret the user’s intent; (2) the desired answers are long-form, necessitating the retrieval and synthesis of diverse information rather… See the full description on the dataset page: https://huggingface.co/datasets/seokwon99/MAVIS.turkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.turkish-seo-reasoning
Turkish SEO Reasoning
Bu veri seti, küçük parametreli bir dil modeline Türkçe SEO vakalarında kanıta dayalı karar verme becerisi kazandırmak ve aynı senaryoda özel bir benchmark oluşturmak için hazırlanmıştır.
Projenin kapsamı
Bu sürümde tool-call eğitimi yoktur. Modelden araç seçmesi veya araç çağrısı üretmesi beklenmez.
Hedeflenen davranış şudur:
Verilen SEO kanıtını okumak
İlgili Google Search Central ilkesini uygulamak
Kısa ve denetlenebilir bir gerekçe… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning.CHUNKY-tulu3-SFT-25k-attributes
SURF Attributes
Minimal dataset for running SURF (Surfacing Unintended Response Failures).
Paper: Chunky Post-Training
Usage
uv run -m surf.cli.main sweep \
--attributes seoirsem/CHUNKY-tulu3-SFT-25k-attributes \
--rubric rubrics/rebuttal.yaml \
-o results/
Fields
prompt: The query text
sae_attributes: List of semantic attribute cluster summaries
How it works
Each prompt was analyzed to extract 10 raw attributes describing its content… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/CHUNKY-tulu3-SFT-25k-attributes.spanish-programmatic-seo-services-dataset
Spanish Programmatic SEO & Services Dataset (1,249 Tracks)
Este dataset de alta densidad contiene 1,249 trayectorias de agentes sintéticos diseñadas específicamente para el entrenamiento (fine-tuning) de modelos de lenguaje (LLMs) en tareas de razonamiento local, intenciones de búsqueda transaccionales y generación de estructuras SEO avanzadas para el mercado de España.
Estructura del Dataset
Cada registro sigue el formato de instrucción tuning estándar… See the full description on the dataset page: https://huggingface.co/datasets/rgjj30/spanish-programmatic-seo-services-dataset.opsd-plain-4b-rollouts
opsd-plain-4b-rollouts
This dataset contains rollout generations collected during training.
Source experiment
method: opsd-plain
model_size: 4b
experiment_dir: /home/irteam/outputs/opsd_plain_4b
Format
Each row contains:
step
sample_index
prompt
completion
method
model_size
source_file
Viewer structure
all: all rollout rows together
step_<N>: only one rollout step, easier to inspect in the dataset viewer
Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-4b-rollouts.sea-product-listing-seo-sample
SEA Multilingual Product Listing SEO Sample
This public sample contains 1,000 synthetic, AI-generated marketplace-style
product listing examples for Southeast Asian e-commerce workflows.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, listing-copy prototyping, SEO
keyword experiments, and multilingual catalog workflow testing.
Important Limitations… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-product-listing-seo-sample.opsd-plain-8b-rollouts
opsd-plain-8b-rollouts
This dataset contains rollout generations collected during training.
Source experiment
method: opsd-plain
model_size: 8b
experiment_dir: /home/irteam/outputs/opsd_plain_8b
Format
Each row contains:
step
sample_index
prompt
completion
method
model_size
source_file
Viewer structure
all: all rollout rows together
step_<N>: only one rollout step, easier to inspect in the dataset viewer
Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-8b-rollouts.global-seo-knowledgehuman_eval_1
