datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SimpleQA
SimpleQA
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Sources
openai/simple-evals
Introducing SimpleQA
Measuring short-form factuality in large language models
Basic-Math-Chinese-1M-V1.1比较于上一个版本
·1.新增了乘方和开方(二次方根)的题目
·2.新增生成比例:
四则运算45%
一元一次方程30%
实际问题15%
乘方与开方10%
·3.新增四则运算变异:生成时有20%的几率在后面问“这个数(加,减,乘,除)a等于几?”(可堆叠)
联系方式:qq:2981447942
bilibili:一髅子Tick
basic_korean_dict
Dataset Card for "basic_korean_dict"
This dataset is a NLP learnable form of Korean Basic Dictionary(한국어기초사전).
It follows the original copyright policy (cc-by-sa-2.0)
Some words have usage examples in other languages, effectively rendering this into a parallel corpus.
This version is built from xls_20230601
한국어 기초 사전을 학습 가능한 형태로 처리한 데이터입니다.
한국어 기초 사전의 저작권을 따릅니다.
여러 언어로 이루어진 표제어들이 있어 병렬 말뭉치의 기능이 있습니다.
xls_20230601으로부터 생성되었습니다.
cooking-knowledge-basics
Comprehensive Cooking Knowledge Q&A Dataset
This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.Harm_Reduction_QA_dataset_basic
>>>In Development<<<
The description below aims to provide an overview of HRIP-Basic benchmark.
This project is also in development; some inconsistencies between the information below and the actual dataset are expected.
If you have any questions about the dataset, please contact the author
Harm Reduction Information Provision Benchmark (HRIPBench)
Dataset Card Contact
Kaixuan Wang
kw215@st-andrews.ac.uk
Dataset Description
HRIP-Basic is a… See the full description on the dataset page: https://huggingface.co/datasets/RayK/Harm_Reduction_QA_dataset_basic.ai-basic-law-dataset
台灣人工智慧基本法 訓練資料集
Taiwan AI Basic Law (人工智慧基本法) Q&A dataset for LLM finetuning.
Files
File
Description
Entries
train.jsonl
Full training dataset with oversampling
~5000
fulltext.jsonl
Clean article fulltext (20 articles)
38
Data Composition
Category
Unique
Repeat
Purpose
Article Fulltext Q&A
~157
x15
Verbatim article text with topic anchors
Alias Recognition
~109
x10
「基本法」「AI基本法」→ 人工智慧基本法
Legislative Reasons
~35
x3
Background… See the full description on the dataset page: https://huggingface.co/datasets/iamjry/ai-basic-law-dataset.prometech_inc_basic_coder
Prometech Inc Basic Coder Dataset
Dataset Overview
Filename: prometech_inc_basic_coder.jsonlTotal Entries: 263,903File Size: ~854 MBProvider: Prometech Bilgisayar Bilimleri AŞ
This dataset is a unified collection of high-quality coding instruction-following records, designed for fine-tuning Large Language Models (LLMs) or for use in Retrieval-Augmented Generation (RAG) systems. It aggregates data from multiple open-source high-quality datasets, synthetic documentation… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/prometech_inc_basic_coder.portugal-basic-qa-ptcore
Portugal Basic QA PTCORE
50 very basic multiple-choice questions about Portugal, written in European Portuguese.
Schema:
question: question in Portuguese
choices: three answer options
label: correct option as a, b, or c
answer: answer text
This dataset is used by nanochat/ptcore_eval.py as the portugal-basic-qa-pt PTCORE task.
basic-mathportugal-basic-qa-ptcore
Portugal Basic QA PTCORE
50 very basic multiple-choice questions about Portugal, written in European Portuguese.
Schema:
question: question in Portuguese
choices: three answer options
label: correct option as a, b, or c
answer: answer text
This dataset is used by nanochat/ptcore_eval.py as the portugal-basic-qa-pt PTCORE task.
basic-general-use-dataset
Basic General Use Dataset
This is a dataset that has just general things for training a small ai
basic-chat-model-dataset
