datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaokao-dataset详见 Github 仓库 rainewhk/gaokao。
rainy-sharegpt-advanced-prefills-filteredPinpointQA
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
Important: This repository releases benchmark annotations and grounded intermediate spatial representations only. It does not redistribute the original scene assets or converted video files.
🧭 Overview
PinpointQA focuses on a practical question: given a known small object such as a phone, charger, remote, or bottle, can a model determine whether it appears, localize it… See the full description on the dataset page: https://huggingface.co/datasets/RainChow/PinpointQA.LongPIBench
LongPIBench
LongPIBench is a benchmark release containing 100 linked examples. Each example
is identified by an integer from 0 through 99 and contains:
a complete multi-file LaTeX paper project in papers/<id>/;
structured personal-profile data in person_info/<id>.json;
an email exchange and attachment text in emails/<id>.json; and
a before/after programming task in code_changes/<id>.json.
The directory structure is part of the dataset. In particular, each paper is a
raw LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/RainWatcher/LongPIBench.SlimOrca-Llama-3-Preference-DPO-Pairs
SlimOrca-Llama-3-Preference-DPO-Pairs
This dataset is based on instructions of SlimOrca-Dedup-Alpaca, with Llama-3 generated response to form a preference dataset.
DreadPoor__Summer_Rain-8B-TIES-details
Dataset Card for Evaluation run of DreadPoor/Summer_Rain-8B-TIES
Dataset automatically created during the evaluation run of model DreadPoor/Summer_Rain-8B-TIES
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Summer_Rain-8B-TIES-details.chinese-roleplayopenassistant-guanaco-deThis is only a Copy of the Work of OpenAssistant and the user timdettmers
Target of this trainingdata is finetuning only in German language.
File openassistant_origfile_with_lang_informations.txt is the full trainingdata. Every line starts with Language Informations. You can easily filter with:
cat openassistant_origfile_with_lang_informations.txt | grep ^de | sed s/^de,//g > openassistant_best_replies_de_train.jsonl
Replace ^de with the language you are interessted in.
For language detection… See the full description on the dataset page: https://huggingface.co/datasets/RainerGa/openassistant-guanaco-de.Dextromethorphan-10k
About Dextromethorphan-10k
Resampled 10k prompt from lmsys-1m. Llama-3-70k generated.
raingoa-artifactsTinyChatVOCALOID_songsMessage_Define_Updateglaive_code_assistant_v3_resample_95krainy-logsDextromethorphan-50k-v0.1r1_20k_high_score_data
高质量筛选数据集
数据集描述
本数据集是从原始数据集 distill_r1_110k_sft.jsonl 中精选而来的高质量子集。我们首先从原始的110k条数据中随机抽取了20k条样本,然后通过质量评分机制进一步筛选,只保留了评分较高的数据条目。
引用
如果您使用了本数据集,请引用原始数据集 distill_r1_110k_sft.jsonl 以及本筛选数据集。
许可证
本数据集遵循Apache 2.0许可证。
text_define_thinkclaudy-chat-5kDreadPoor__Summer_Rain-8B-SCE-details
Dataset Card for Evaluation run of DreadPoor/Summer_Rain-8B-SCE
Dataset automatically created during the evaluation run of model DreadPoor/Summer_Rain-8B-SCE
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Summer_Rain-8B-SCE-details.Message_defineIFC-function-callingclaudy-chat-CJK-5kTSInSAR-LLMOpenThoughts-79k-filtered-Translated-Chinese-10k-cleanedTiny-Philosopher-50kScam_Detect_SplitRED20_InstructScam_Detect_20pretrain_demo这个数据集来自minimind_dataset数据集。
由于其pretrain_hq.jsonl相对较大,有1.6G左右,这里取了其1/8左右的内容,只用了216MB左右的数据。
而且格式上也做了一些简单修改,具体内容可以查看数据详情。
之所以创建这个数据集,主要是为了在RainFallModelFactory中作为演示使用,用别的地方的数据集有被删掉的风险。
所以创建一个稳定存在的,不会被删掉的数据集。
