clue
Datasets
All datasets matching “clue”clue
Dataset Card for "clue"
Dataset Summary
CLUE, A Chinese Language Understanding Evaluation Benchmark
(https://www.cluebenchmarks.com/) is a collection of resources for training,
evaluating, and analyzing Chinese language understanding systems.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
afqmc
Size of downloaded dataset files: 1.20 MB
Size… See the full description on the dataset page: https://huggingface.co/datasets/clue/clue.ClueAnchor
ClueAnchor Dataset
This repository provides the datasets used and generated in the paper ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation.
Code: https://github.com/thunlp/ClueAnchor
Model: https://huggingface.co/MethaneChen222/ClueAnchor
Abstract
Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge to improve factuality. However, existing RAG systems frequently… See the full description on the dataset page: https://huggingface.co/datasets/MethaneChen222/ClueAnchor.ClueCorpusSmallDatasetwikipedia-ipa
Wikipedia IPA Audio and Symbols
Normalized IPA consonant and vowel inventories with reference
audio, derived from the Wikipedia IPA chart pages.
One record per IPA symbol, tagged with the phonological
dimensions used on those chart pages.
Contents
base/consonant/sound.jsonl: one record per pulmonic
consonant, with symbol, place, manner, and voicing.
base/vowel/sound.jsonl: one record per vowel, with
symbol, height, backness, and roundedness.
base/consonant/audio/*.wav… See the full description on the dataset page: https://huggingface.co/datasets/cluesurf/wikipedia-ipa.clue-ner
CLUE-NER 命名实体识别数据集
字段说明
text: 文本
entities: 文本中包含的实体
id: 实体 id
entity: 实体对应的字符串
start_offset: 实体开始位置
end_offset: 实体结束位置的下一位
label: 实体对应的开始位置
urdu_afsana_audiobook_dataset
