CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ginigen-ai /smol-worldcup 🏟️ Smol AI WorldCup — SHIFT Benchmark The world's first 5-axis evaluation framework for small language models. Not just "how smart?" — but "how honest? how fast? how small? how efficient?" 🏟️ Leaderboard huggingface.co/spaces/ginigen-ai/smol-worldcup 📊 Dataset huggingface.co/datasets/ginigen-ai/smol-worldcup 🏅 ALL Bench huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard 🏆 Official Ranking: WCS (WorldCup Score) WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.tabulartext-generationn<1K47 likes425 downloads7mo agoHugging Face02GinkgoQ /LongBench LongBench Dataset Summary LongBench is a bilingual, multitask benchmark for evaluating long-context understanding in large language models. It covers long-text application scenarios including single-document question answering, multi-document question answering, summarization, few-shot learning, synthetic long-context tasks, and code completion. This Hugging Face dataset repository repackages locally downloaded LongBench JSONL files into a clean, typed, data-only… See the full description on the dataset page: https://huggingface.co/datasets/GinkgoQ/LongBench.tabularquestion-answering1K<n<10K1 likes120 downloads4mo agoHugging Face03gineven /GeneralScience-MLLM-22K GeneralScience-MLLM-22K Dataset Summary GeneralScience-MLLM-22K is a unified general-science multiple-choice QA collection built from local snapshots of SciQ, AI2 ARC, and ScienceQA. It follows the same release style as a subject-specific MLLM dataset: every sample is stored as one JSONL record, text-only and image-text examples share one schema, and ScienceQA images are exported as standalone files referenced by relative paths. The release contains 22,661… See the full description on the dataset page: https://huggingface.co/datasets/gineven/GeneralScience-MLLM-22K.imagequestion-answering10K<n<100K1 likes61 downloads2mo agoHugging Face04ginigen-ai /Metacognition-Bench Metacognition-Bench "Not whether a model knows the answer — but whether it knows when it might be wrong, and can correct itself." Metacognition-Bench is a curated benchmark of 300 metacognitive-trap problems that measure functional metacognition in Large Language Models: the ability to detect and recover from one's own reasoning errors, rather than final-answer accuracy alone. Every problem embeds a hidden_trap — a seductive but wrong reasoning path that makes even capable… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/Metacognition-Bench.text-generationn<1K31 likes50 downloads3mo agoHugging Face05ginigen /Korean-Hallucination-Bench Korean Hallucination Benchmark (한국어 환각 진단 벤치마크) 한국어 LLM의 환각(hallucination) 저항성을 평가하는 4지선다 벤치마크입니다. 법령·특허·행정·의료·금융 5개 전문 도메인에서, 한국어 특화 환각 유형 10종을 다룹니다. 각 문항은 사실 정답 1개와 전문가도 속을 만큼 그럴듯한 환각 오답 3개로 구성됩니다. 통계 총 10,167문항 (4지선다) 도메인(5): 법령 1,901 · 특허 2,190 · 행정 1,864 · 의료 2,113 · 금융 2,099 환각 유형(10): 수치·날짜 오류, 조항 왜곡, 근거 없는 추론, 한자어·신조어 혼재, 과잉 일반화, 멀티턴 맥락 붕괴, 존댓말·반말 역전, 출처 날조, 용어 왜곡, 사실 날조 구축 방법 문항 생성: Darwin-398B-JGOS (문제·선택지) 정답 검수·교정: Claude (Anthropic)… See the full description on the dataset page: https://huggingface.co/datasets/ginigen/Korean-Hallucination-Bench.textquestion-answering10K<n<100K3 likes33 downloads3mo agoHugging Face06ginipick /awesome-chatgpt-prompts🧠 Awesome ChatGPT Prompts [CSV dataset] This is a Dataset Repository of Awesome ChatGPT Prompts View All Prompts on GitHub License CC-0 textquestion-answeringn<1K0 likes22 downloads11mo agoHugging Face07gingdev /llama_vi_52k Llama 2 Vietnamese dataset Bộ dữ liệu Alpaca được dịch sang tiếng Việt theo chuẩn Llama 2 Prompt. Prompt template <s>[INST] <<SYS>> {system_message} <</SYS>> {user_message_1} [/INST] {model_reply_1}</s><s>[INST] {user_message_2} [/INST] Tác giả Iambestfeed Alex Nguyen Thanh Trần textquestion-answering10K<n<100K0 likes17 downloads3y agoHugging Face08yensonalvi6 /ginecologia-venezuela Ginecología Venezuela Dataset Descripción del Dataset Este dataset contiene 250 instrucciones especializadas en ginecología y obstetricia, enfocadas específicamente en el contexto de la salud pública venezolana. Está diseñado para entrenar modelos de lenguaje en el dominio médico ginecológico con consideraciones específicas del sistema de salud venezolano. Contenido Tamaño: 250 ejemplos de instrucciones Idioma: Español (Venezuela) Dominio: Ginecología y… See the full description on the dataset page: https://huggingface.co/datasets/yensonalvi6/ginecologia-venezuela.texttext-generationn<1K1 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.