CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01seongchaeae /sage-pretrain-corpus Sage Pretrain Corpus (v0.8) License-audited Korean/English pretraining corpus for the Sage Korean LLM project. Records: id, text, source, license, meta (JSON of original fields). ⚠️ Mixed licenses — license: other. Comply with each source's license individually. CC-BY / ODC-BY require attribution; CC-BY-SA carries ShareAlike. Sources (v0.8) source docs license origin fineweb2_ko 46,470,574 ODC-BY-1.0 FineWeb-2 Korean (HuggingFaceFW/fineweb-2 kor_Hang)… See the full description on the dataset page: https://huggingface.co/datasets/seongchaeae/sage-pretrain-corpus.texttext-generation100M<n<1B1 likes2.9k downloads4mo agoHugging Face02seoulraphaellee /korean-assembly-minutes 대한민국 국회 회의록 아카이브 국회 회의록 원문(record.assembly.go.kr) PDF 에서 본문을 뽑아 모은 것이다. 본회의와 각 위원회 회의록이 모두 들어 있다. 수록 기간: 1948~1993 회의 수: 1,951건 본문 분량: 65,306,444자 구성 연도별 JSONL(gzip) 한 덩이다. from datasets import load_dataset ds = load_dataset("seoulraphaellee/korean-assembly-minutes", split="train") ds = load_dataset("seoulraphaellee/korean-assembly-minutes", data_files="data/2026.jsonl.gz", split="train") 필드 이름 설명 meeting_key 회의 식별자 (record:<id>)… See the full description on the dataset page: https://huggingface.co/datasets/seoulraphaellee/korean-assembly-minutes.tabulartext-generation10K<n<100K0 likes427 downloads15d agoHugging Face03metehan777 /global-seo-knowledgetexttext-generation1K<n<10K3 likes200 downloads1y agoHugging Face04seoirsem /CHUNKY-tulu3-SFT-25k-attributes-full SURF Attributes (Full) Complete dataset for SURF research and extension. Paper: Chunky Post-Training Quick Start For running SURF, use the minimal dataset: seoirsem/CHUNKY-tulu3-SFT-25k-attributes uv run -m surf.cli.main sweep \ --attributes seoirsem/CHUNKY-tulu3-SFT-25k-attributes \ --rubric rubrics/rebuttal.yaml \ -o results/ Dataset Fields prompt: The query text response: The model response (if available) attributes: Raw extracted attributes… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/CHUNKY-tulu3-SFT-25k-attributes-full.texttext-generation100K<n<1M0 likes113 downloads8mo agoHugging Face05seokwon99 /MAVIS MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering 📖 Paper | 💻 Evaluation Dataset Summary MAVIS is a new dataset for open-domain, long-form visual question answering, characterized by three key features: (1) the questions incorporate input images, requiring visual understanding to correctly interpret the user’s intent; (2) the desired answers are long-form, necessitating the retrieval and synthesis of diverse information rather… See the full description on the dataset page: https://huggingface.co/datasets/seokwon99/MAVIS.imagequestion-answeringn<1K1 likes104 downloads8mo agoHugging Face06berkbirkan /turkish-seo-reasoning-benchmark-results Turkish SEO Reasoning Benchmark Results Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir. Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora Sonuç Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti. Mutlak artış: +10,28 puan Göreli artış: %85,97 Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.tabulartext-generationn<1K0 likes55 downloads2mo agoHugging Face07berkbirkan /turkish-seo-reasoning Turkish SEO Reasoning Bu veri seti, küçük parametreli bir dil modeline Türkçe SEO vakalarında kanıta dayalı karar verme becerisi kazandırmak ve aynı senaryoda özel bir benchmark oluşturmak için hazırlanmıştır. Projenin kapsamı Bu sürümde tool-call eğitimi yoktur. Modelden araç seçmesi veya araç çağrısı üretmesi beklenmez. Hedeflenen davranış şudur: Verilen SEO kanıtını okumak İlgili Google Search Central ilkesini uygulamak Kısa ve denetlenebilir bir gerekçe… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning.textquestion-answering1K<n<10K0 likes51 downloads2mo agoHugging Face08seoirsem /CHUNKY-tulu3-SFT-25k-attributes SURF Attributes Minimal dataset for running SURF (Surfacing Unintended Response Failures). Paper: Chunky Post-Training Usage uv run -m surf.cli.main sweep \ --attributes seoirsem/CHUNKY-tulu3-SFT-25k-attributes \ --rubric rubrics/rebuttal.yaml \ -o results/ Fields prompt: The query text sae_attributes: List of semantic attribute cluster summaries How it works Each prompt was analyzed to extract 10 raw attributes describing its content… See the full description on the dataset page: https://huggingface.co/datasets/seoirsem/CHUNKY-tulu3-SFT-25k-attributes.texttext-generation100K<n<1M0 likes36 downloads8mo agoHugging Face09rgjj30 /spanish-programmatic-seo-services-dataset Spanish Programmatic SEO & Services Dataset (1,249 Tracks) Este dataset de alta densidad contiene 1,249 trayectorias de agentes sintéticos diseñadas específicamente para el entrenamiento (fine-tuning) de modelos de lenguaje (LLMs) en tareas de razonamiento local, intenciones de búsqueda transaccionales y generación de estructuras SEO avanzadas para el mercado de España. Estructura del Dataset Cada registro sigue el formato de instrucción tuning estándar… See the full description on the dataset page: https://huggingface.co/datasets/rgjj30/spanish-programmatic-seo-services-dataset.texttext-generation1K<n<10K0 likes32 downloads13d agoHugging Face10SeongryongJung /opsd-plain-4b-rollouts opsd-plain-4b-rollouts This dataset contains rollout generations collected during training. Source experiment method: opsd-plain model_size: 4b experiment_dir: /home/irteam/outputs/opsd_plain_4b Format Each row contains: step sample_index prompt completion method model_size source_file Viewer structure all: all rollout rows together step_<N>: only one rollout step, easier to inspect in the dataset viewer Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-4b-rollouts.tabulartext-generationn<1K0 likes23 downloads4mo agoHugging Face11nwchang /sea-product-listing-seo-sample SEA Multilingual Product Listing SEO Sample This public sample contains 1,000 synthetic, AI-generated marketplace-style product listing examples for Southeast Asian e-commerce workflows. Languages English Chinese Malay Indonesian Formats CSV JSONL Intended Use Use this sample for inspection, evaluation, listing-copy prototyping, SEO keyword experiments, and multilingual catalog workflow testing. Important Limitations… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-product-listing-seo-sample.texttext-generation1K<n<10K0 likes22 downloads4mo agoHugging Face12SeongryongJung /opsd-plain-8b-rollouts opsd-plain-8b-rollouts This dataset contains rollout generations collected during training. Source experiment method: opsd-plain model_size: 8b experiment_dir: /home/irteam/outputs/opsd_plain_8b Format Each row contains: step sample_index prompt completion method model_size source_file Viewer structure all: all rollout rows together step_<N>: only one rollout step, easier to inspect in the dataset viewer Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-8b-rollouts.tabulartext-generationn<1K0 likes12 downloads4mo agoHugging Face13YellowJack /global-seo-knowledgetexttext-generation1K<n<10K0 likes7 downloads9mo agoHugging Face14Seongyun /human_eval_1texttext-generationn<1K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.