CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01basicv8vc /SimpleQA SimpleQA A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions. Sources openai/simple-evals Introducing SimpleQA Measuring short-form factuality in large language models textquestion-answering1K<n<10K32 likes3.7k downloads2y agoHugging Face02TICK666 /Basic-Math-Chinese-1M-V1.1比较于上一个版本 ·1.新增了乘方和开方(二次方根)的题目 ·2.新增生成比例: 四则运算45% 一元一次方程30% 实际问题15% 乘方与开方10% ·3.新增四则运算变异:生成时有20%的几率在后面问“这个数(加,减,乘,除)a等于几?”(可堆叠) 联系方式:qq:2981447942 bilibili:一髅子Tick textquestion-answering1M<n<10M6 likes119 downloads3y agoHugging Face03hac541309 /basic_korean_dict Dataset Card for "basic_korean_dict" This dataset is a NLP learnable form of Korean Basic Dictionary(한국어기초사전). It follows the original copyright policy (cc-by-sa-2.0) Some words have usage examples in other languages, effectively rendering this into a parallel corpus. This version is built from xls_20230601 한국어 기초 사전을 학습 가능한 형태로 처리한 데이터입니다. 한국어 기초 사전의 저작권을 따릅니다. 여러 언어로 이루어진 표제어들이 있어 병렬 말뭉치의 기능이 있습니다. xls_20230601으로부터 생성되었습니다. texttable-question-answering10K<n<100K6 likes63 downloads3y agoHugging Face04ktiyab /cooking-knowledge-basics Comprehensive Cooking Knowledge Q&A Dataset This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.textquestion-answering1K<n<10K6 likes53 downloads2y agoHugging Face05leninangelov /basic-chat-model-datasettextquestion-answering1K<n<10K0 likes24 downloads1d agoHugging Face06pthinc /prometech_inc_basic_coder Prometech Inc Basic Coder Dataset Dataset Overview Filename: prometech_inc_basic_coder.jsonlTotal Entries: 263,903File Size: ~854 MBProvider: Prometech Bilgisayar Bilimleri AŞ This dataset is a unified collection of high-quality coding instruction-following records, designed for fine-tuning Large Language Models (LLMs) or for use in Retrieval-Augmented Generation (RAG) systems. It aggregates data from multiple open-source high-quality datasets, synthetic documentation… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/prometech_inc_basic_coder.tabulartext-generation100K<n<1M0 likes22 downloads9mo agoHugging Face07iamjry /ai-basic-law-dataset 台灣人工智慧基本法 訓練資料集 Taiwan AI Basic Law (人工智慧基本法) Q&A dataset for LLM finetuning. Files File Description Entries train.jsonl Full training dataset with oversampling ~5000 fulltext.jsonl Clean article fulltext (20 articles) 38 Data Composition Category Unique Repeat Purpose Article Fulltext Q&A ~157 x15 Verbatim article text with topic anchors Alias Recognition ~109 x10 「基本法」「AI基本法」→ 人工智慧基本法 Legislative Reasons ~35 x3 Background… See the full description on the dataset page: https://huggingface.co/datasets/iamjry/ai-basic-law-dataset.textquestion-answering1K<n<10K0 likes22 downloads7mo agoHugging Face08duarteocarmo /portugal-basic-qa-ptcore Portugal Basic QA PTCORE 50 very basic multiple-choice questions about Portugal, written in European Portuguese. Schema: question: question in Portuguese choices: three answer options label: correct option as a, b, or c answer: answer text This dataset is used by nanochat/ptcore_eval.py as the portugal-basic-qa-pt PTCORE task. textquestion-answeringn<1K0 likes14 downloads4mo agoHugging Face09Visualforma /portugal-basic-qa-ptcore Portugal Basic QA PTCORE 50 very basic multiple-choice questions about Portugal, written in European Portuguese. Schema: question: question in Portuguese choices: three answer options label: correct option as a, b, or c answer: answer text This dataset is used by nanochat/ptcore_eval.py as the portugal-basic-qa-pt PTCORE task. textquestion-answeringn<1K0 likes12 downloads2mo agoHugging Face10CJHauser /basic-general-use-dataset Basic General Use Dataset This is a dataset that has just general things for training a small ai textquestion-answeringn<1K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.