datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.KM-Medallion
🧬 K. marxianus Omniscient Digital Twin v8.2.0
The most comprehensive public dataset for Kluyveromyces marxianus worldwide
This dataset implements the "Omniscient Data Ingestion Protocol" with zero tolerance for false negatives, covering 4000+ repositories, 1200+ historical strain names, and complete provenance tracking.
🔬 Taxonomic Coverage
This dataset handles all historical synonyms for K. marxianus:
Candida kefyrKluyveromyces fragilis
Saccharomyces marxianus… See the full description on the dataset page: https://huggingface.co/datasets/Milad96/KM-Medallion.scientific-multitask-instructions
Scientific Multitask Instructions
A multi-task scientific instruction-following dataset created for
supervised fine-tuning and preference-optimization experiments.
Dataset summary
The dataset contains 1,576 conversational scientific examples across
eight task types.
Split
Examples
Train
1,260
Validation
158
Test
158
Total
1,576
Task distribution
Task
Examples
Scientific question answering
256
Summarization
220… See the full description on the dataset page: https://huggingface.co/datasets/Miladsaeedi70/scientific-multitask-instructions.alpaca
Dataset Card for Alpaca
Dataset Summary
Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better.
The authors built on the data generation pipeline from Self-Instruct framework and made the following modifications:
The text-davinci-003 engine to generate the instruction data instead… See the full description on the dataset page: https://huggingface.co/datasets/Milabench/alpaca.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/Milabench/alpaca-cleaned.
