datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sni-each-convertedConverted Super-NaturalInstructions to jsonl
https://github.com/allenai/natural-instructions
converted_narrative_qa
Dataset Card for "converted_narrative_qa"
More Information needed
VR-egodex-annotation-converted-v6.0
VR-egodex-annotation-converted-v6.0
EgoDex converted from LeRobot v2.1 into the Layer-1 v0.6.0 annotation schema, with
per-clip narration included as language sidecars.
314,839 clips · 78,282,306 frames · 724.8 hours @ 30 fps · 129 tasks
100% narration coverage (1 sidecar per clip)
71 GB annotations + 2.3 GB narratives
Videos are NOT included. This release contains annotations and narration only. Source
video lives in griffinlabs/EgoDex-LeRobot-v3.0;
orig_id in the manifest… See the full description on the dataset page: https://huggingface.co/datasets/VR-VLA/VR-egodex-annotation-converted-v6.0.Toucan-converted-datasetcommonvoice-12.0-arabic-voice-converted
Dataset Card for Voice Converted Arabic Common Voice 12.0
This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.flan_v2_convertedThis is a converted version of the Flan dataset into Tulu SFT training format.
The conversion script can be found in our open-instruct repo.
The conversion took the following parameters:
apply_keyword_filters: True
apply_empty_message_filters: True
push_to_hub: True
hf_entity: ai2-adapt-dev
converted_dataset_name: flan_v2_converted
local_save_dir: ./data/sft/flan
The original FLAN dataset needs extensive efforts to be regenerated, so we are using a reproduced version by the OpenOrca… See the full description on the dataset page: https://huggingface.co/datasets/ai2-adapt-dev/flan_v2_converted.II-Medical-Reasoning-SFT-Convertedpersona_instruct_2shot_convertedaya_collection_language_split-standard_malay-ConvertedOpenCodeInstruct-ConvertedMegaScience-dataset-ConvertedOpenCodeReasoning-split_1-ConvertedxP3x-ind_Latn-Convertedpubhealth-converted
PUBHEALTH
PUBHEALTH is a public-health fact-checking dataset introduced in:
Neema Kotonya and Francesca Toni. 2020. Explainable Automated Fact-Checking for Public Health Claims.
This repository is a scriptless Parquet conversion of the original PUBHEALTH TSV files. It is intended to load with the Hugging Face datasets library without requiring deprecated remote dataset scripts.
The deprecated script-based dataset was available at:… See the full description on the dataset page: https://huggingface.co/datasets/Jezzarax/pubhealth-converted.riddle_sense_convertedconverted_cvssNemotron-Post-Training-Dataset-v2-chat-Convertedconverted_mixed_pickandplace_datasetpersona_math_2shot_convertedSYNTHETIC-2-SFT-verified-ConvertedFRED-CONVERTEDaya_collection_language_split-central_khmer-Convertedindonesian-reasoning-Convertedaya_collection_language_split-thai-ConvertedTrendyol-Cybersecurity-Instruction-Tuning-Dataset-Convertedpersona_instruction_following_convertedDeepWriting-20K-Convertedrefchartqa_converted_hfCapybara-Converted
This is the Official Capybara dataset. Over 10,000 multi-turn examples.
Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others.
The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.no_robots_convertedThis is a converted version of the no_robots dataset into Tulu SFT training format.
The conversion script can be found in our open-instruct repo.
The conversion took the following parameters:
apply_keyword_filters: False
apply_empty_message_filters: False
push_to_hub: True
hf_entity: ai2-adapt-dev
converted_dataset_name: no_robots_converted
local_save_dir: ./data/sft/no_robots
Please refer to the original dataset for more information about this dataset and the license.
