Unfiltered
ShareGPT_Vicuna_unfilteredFurther cleaning done. Please look through the dataset and ensure that I didn't miss anything.
Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c
Two choices:
Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered/blob/main/ShareGPT_V3_unfiltered_cleaned_split_no_imsorry.json
Has instances of "I'm sorry, but":… See the full description on the dataset page: https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered.ShareGPT_Vicuna_unfiltered
Dataset Card
This is a reupload of this dataset that was further cleaned by gozfarb.
Vero-2.5M-unfiltered
Vero-2.5M-unfiltered
[!Note]
This repository contains the full unfiltered dataset used to construct Vero-600k and Vero-1.6M, before question and answer filtering.
Note that task categories are not balanced in this dataset.
Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with vision-language models. This repository contains the Vero-2.5M-unfiltered dataset, a curation of 2.5M reinforcement learning samples… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/Vero-2.5M-unfiltered.Function_Calling_Unfilteredsynth-cc-unfilteredembeddings-fine-tuning-multilingual-unfiltered
Overview
This dataset provides multilingual and code retrieval data for fine-tuning text embedding models. It is composed of high quality data sources with mined documents annotated with bi-encoder scores.
For each query, the 2048 closest documents are mined with snowflake-arctic-embed-l-v2.0 for MIRACL and MLDR and with gte-modernbert-base for CodeEditSearchTrain, and annotated with their bi-encoder similarity score. No false-negative filtering or cross-encoder annotation is… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered.
