datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb_2_500k_both_deduplicatedsharegpt-deduplicated
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset is a deduplicated version of sharegpt4.
The deduplication process has two steps:
The literal duplicates (both input and outputs) are removed
The remaining (5749) instances are embedded with the SentenceTransformer library ("paraphrase-multilingual-mpnet-base-v2" model).
Then, we compute the cosine similarity among all the possible pairs, and consider paraphrases those pairs with a… See the full description on the dataset page: https://huggingface.co/datasets/CaterinaLac/sharegpt-deduplicated.nl2sql-deduplicated
NL2SQL Deduplicated Training Dataset
A curated and deduplicated Text-to-SQL training dataset with 683,015 unique examples from 4 high-quality sources.
📊 Dataset Summary
Total Examples: 683,015 unique question-SQL pairs
Sources: Spider, SQaLe, Gretel Synthetic, SQL-Create-Context
Deduplication Strategy: Input-only (question-based) with conflict resolution via quality priority
Conflicts Resolved: 2,238 cases where same question had different SQL
SQL Dialect: Standard SQL… See the full description on the dataset page: https://huggingface.co/datasets/AsadIsmail/nl2sql-deduplicated.master-ebook-library-deduplicatedClustering_deduplicated_reasoning
Clustering_deduplicated_reasoning
数据集描述
Clustering deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category
文件结构
clustering_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
# 加载数据集
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Clustering_deduplicated_reasoning.100K_deduplicated_ner_indexes_name_country_alpaca_format_json_response_all_casesSemantic_similarity_deduplicated_reasoning_data_english
Semantic_similarity_deduplicated_reasoning_data_english
数据集描述
Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category
文件结构
semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.50K_deduplicated_ner_indexes_name_country_alpaca_format_json_responsehash_deduplicated_reasoning_data_english
hash_deduplicated_reasoning_data_english
数据集描述
Hash deduplicated reasoning data filtered from OpenThoughts2-1M, 72710 examples in total
文件结构
hash_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
# 加载数据集
dataset = load_dataset("Ibisbill/hash_deduplicated_reasoning_data_english")… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/hash_deduplicated_reasoning_data_english.vhdl_github_deduplicatedifc-bim-qa-deduplicated_2deduplicated_cot_fs_noopt_train.jsonlDataset Summary
This is a processed and deduplicated version of the Flan V2 dataset.
I'm not affiliated with the creators, I'm just releasing the files in an easier-to-access format after processing.
Data Instances
Chain-of-thought (cot)
cot_fs_noopt_train.jsonl
deduplicated_PersianQuADpwc-github-links-deduplicatedtest_deduplicated_datasetdeduplicated_datasettulu-3-deduplicated
