CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01seonjeongh /science_reasoning science_reasoning Mistral-7B의 과학 지식·추론 능력 향상을 위해 6개 공개 과학 객관식 QA 데이터셋을 통일 포맷으로 변환하고, ARC-Challenge test와의 오염을 제거한 데이터셋입니다. 원본 데이터셋 allenai/sciq allenai/openbookqa (main) allenai/qasc allenai/quartz allenai/ai2_arc (ARC-Easy / ARC-Challenge) nguyen-brat/worldtree 전처리 포맷 통일: 각 데이터셋의 서로 다른 스키마를 unique_id, orig_id, source, question, choices, answer, support 필드로 변환. support는 근거 문단/문장으로, 데이터셋별 원본 필드(support/fact/para/cot)에서 구성하거나 없으면 빈 문자열.… See the full description on the dataset page: https://huggingface.co/datasets/seonjeongh/science_reasoning.textmultiple-choice10K<n<100K0 likes328 downloads2mo agoHugging Face02seonglae /wikipedia-256This is Wikidedia passages dataset for ODQA retriever. Each passages have 256~ tokens splitteed by gpt-4 tokenizer using tiktoken. Token count {'~128': 1415068, '128~256': 1290011, '256~512': 18756476, '512~1024': 667, '1024~2048': 12, '2048~4096': 0, '4096~8192': 0, '8192~16384': 0, '16384~32768': 0, '32768~65536': 0, '65536~128000': 0, '128000~': 0} Text count {'~512': 1556876,'512~1024': 6074975, '1024~2048': 13830329, '2048~4096': 49, '4096~8192': 2, '8192~16384': 3, '16384~32768': 0… See the full description on the dataset page: https://huggingface.co/datasets/seonglae/wikipedia-256.textquestion-answering10M<n<100M0 likes295 downloads3y agoHugging Face03metehan777 /global-seo-knowledgetexttext-generation1K<n<10K3 likes200 downloads1y agoHugging Face04seokwon99 /MAVIS MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering 📖 Paper | 💻 Evaluation Dataset Summary MAVIS is a new dataset for open-domain, long-form visual question answering, characterized by three key features: (1) the questions incorporate input images, requiring visual understanding to correctly interpret the user’s intent; (2) the desired answers are long-form, necessitating the retrieval and synthesis of diverse information rather… See the full description on the dataset page: https://huggingface.co/datasets/seokwon99/MAVIS.imagequestion-answeringn<1K1 likes104 downloads8mo agoHugging Face05berkbirkan /turkish-seo-reasoning-benchmark-results Turkish SEO Reasoning Benchmark Results Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir. Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora Sonuç Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti. Mutlak artış: +10,28 puan Göreli artış: %85,97 Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.tabulartext-generationn<1K0 likes55 downloads2mo agoHugging Face06berkbirkan /turkish-seo-reasoning Turkish SEO Reasoning Bu veri seti, küçük parametreli bir dil modeline Türkçe SEO vakalarında kanıta dayalı karar verme becerisi kazandırmak ve aynı senaryoda özel bir benchmark oluşturmak için hazırlanmıştır. Projenin kapsamı Bu sürümde tool-call eğitimi yoktur. Modelden araç seçmesi veya araç çağrısı üretmesi beklenmez. Hedeflenen davranış şudur: Verilen SEO kanıtını okumak İlgili Google Search Central ilkesini uygulamak Kısa ve denetlenebilir bir gerekçe… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning.textquestion-answering1K<n<10K0 likes51 downloads2mo agoHugging Face07seongs /dell-qa-en-to-ko-translated-by-ke-t5-base Dell QA English to Korean Translation Dataset Dataset Description This dataset, dell-qa-en-to-ko-translated-by-ke-t5-base, is a Korean translation of the original English Dell QA dataset. Source The original dataset, dell_qa, is designed for question-answering tasks and contains questions and answers related to Dell technologies. This translated version extends the utility to Korean language tasks. Dataset Structure Data Fields input… See the full description on the dataset page: https://huggingface.co/datasets/seongs/dell-qa-en-to-ko-translated-by-ke-t5-base.textquestion-answering10K<n<100K1 likes20 downloads1y agoHugging Face08YellowJack /global-seo-knowledgetexttext-generation1K<n<10K0 likes7 downloads9mo agoHugging Face09seooyxx /GUI-Libra-81K-SFT-extendedgated GUI-Libra-81K-SFT-extended This dataset contains: data/images/: split image archives (*.tar.gz.part-*) data/annotations/: original GUI-Libra JSON annotation files new_annotation/: LaRA-GUI structured annotation files stored as raw JSON files image_t1_feature_cache/: sharded cached image_t1 visual features for transition supervision Compared with the original GUI-Libra dataset release, this extended variant also records: image_t1 paths in the JSON samples when a valid next frame… See the full description on the dataset page: https://huggingface.co/datasets/seooyxx/GUI-Libra-81K-SFT-extended.textquestion-answering100K<n<1M0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.