datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/guidelines.guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-b1/guidelines.task879_schema_guided_dstc8_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task879_schema_guided_dstc8_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task879_schema_guided_dstc8_classification.br-sovereign-llm-corpus
BR Sovereign LLM Corpus
Status
This public repository is an audited corpus protocol and initial validation
snapshot. Public release 0.1.1 contains only a small, explicitly identified
corpus sample for pipeline verification. It is not a target-scale pretraining
corpus and does not support model-quality claims.
The scientific corpus remains under construction. Aggregate counts, source
shares, and token counts are not reported until a content-addressed snapshot… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/br-sovereign-llm-corpus.smolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
hulk_dataset_0.1This dataset is AFAIK (12 january 2024) the biggest ready to use open source dataset to finetune LLMs. It contains more than 3.8 million chat samples.
Its a collection of multiple different datasets. Some of them have been built using GPT4 or using scraped data. Here is the list:
gathnex/Gath_baize
teknium/openhermes
nomic-ai/gpt4all-j-prompt-generations
teknium/dataforge-economics
Anthropic/hh-rlhf: we kept only the selected prompts
teknium1_GPTeacher_codegen… See the full description on the dataset page: https://huggingface.co/datasets/guigux/hulk_dataset_0.1.prepware_study_guide-dataset
Prepware_Study_Guide Dataset
Generated by DocParserEngine.
Field
Value
Documents
1
Records
1
Schema
full
Usage
from datasets import load_dataset
ds = load_dataset("Remixonwin/prepware_study_guide-dataset")
tabiji-travel-safety-guides
Tabiji Travel & Safety Guides
AI-curated travel data from tabiji.ai: destination profiles, day-by-day itineraries, head-to-head comparisons, safety profiles, country-level travel advisories, and city-level scam guides — sourced from Reddit, government advisories (US State Dept., UK FCDO), and editorial curation.
What's in here
Config
Records
Description
destinations
6,498
Global destination catalog: climate, currency, language, plug type, tap-water safety… See the full description on the dataset page: https://huggingface.co/datasets/tabiji/tabiji-travel-safety-guides.astro_qa_fr_0.1
Astrophysics french QA
The "Astrophysics french QA" dataset is an innovative collection combining scraped articles from the web with ChatGPT-generated question and answer pairs, offering a unique blend of information and interactive learning in the field of astrophysics. It contains almost 5k prompt / response generated by ChatGPT. It can be used to train / finetune / evaluate LLMs on astro subjects.
task880_schema_guided_dstc8_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task880_schema_guided_dstc8_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task880_schema_guided_dstc8_classification.pulaar_corpus
Ndimaagu Pulaar Corpus
Dataset Description
This dataset contains the full transcription of the Pulaar folktale "Ndimaagu" (Nobleness/Dignity). It follows the story of Daado, Yero, and the challenges they face regarding honor and loyalty.
[cite_start]Source: ndimaagu.pdf [cite: 384, 544]
[cite_start]Language: Pulaar (ff) [cite: 384]
Format: Apache Parquet
Structure
Each entry in the dataset represents a narrative segment or a dialogue:
[cite_start]id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/guizme/pulaar_corpus.taiwan-epilepsy-diagnostic-guidelines-qa
Taiwan Epilepsy Diagnostic Guidelines QA
This dataset contains question-answer pairs related to Taiwan's epilepsy diagnostic guidelines, generated using a combination of PDF partitioning and AI-powered question-answering. The dataset is designed to facilitate the development of language models for epilepsy diagnosis and clinical decision support.
Dataset Creation
The dataset was created using a combination of PDF partitioning and AI-powered question-answering,
based on… See the full description on the dataset page: https://huggingface.co/datasets/leeannielee/taiwan-epilepsy-diagnostic-guidelines-qa.smolified-bengali-local-food-guide
🤏 smolified-bengali-local-food-guide
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-bengali-local-food-guide.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 638d3b25)
Records: 1050
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
