datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
humans-top
humans.top — LIVE Global ranking of influential people (open dataset)
This dataset ranks real, named living people by global influence — e.g. #1
Donald Trump, #2 Xi Jinping, #3 Vladimir Putin, alongside figures like Elon Musk,
Narendra Modi and Lionel Messi. Every row is a person: their live influence
rank, a concise biography in 15 languages, and Wikidata / Wikipedia links.
Published from the website humans.top (.top is the
domain name).
Available on (identical CC0… See the full description on the dataset page: https://huggingface.co/datasets/dsfox/humans-top.genomes-v4-genome_set-humans-intervals-v15_256_128genomes-v4-genome_set-humans-intervals-v1_256_128-id0.3_cov0.3genomes-v4-genome_set-humans-intervals-v1_254_127-id0.3_cov0.3genomes-v4-genome_set-humans-intervals-v5_256_128genomes-v4-genome_set-humans-intervals-v1_256_128genomes-v5-genome_set-humans-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v1_255_128
Humans promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
226,162 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v1_255_128.genomes-v4-genome_set-humans-intervals-v16_254_127-id0.3_cov0.3genomes-v2-genome_set-humans-intervals-v1_512_256genomes-v2-genome_set-humans-intervals-v2_512_256genomes-v5-genome_set-humans-intervals-v15_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v15_255_128
Humans downstream-of-CDS (v15) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
66,290 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v15_255_128.genomes-v3-genome_set-humans-intervals-v1_512_256genomes-v5-genome_set-humans-intervals-v17_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v17_255_128
Humans cCRE enhancers (v17) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
3,447,988 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included).… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v17_255_128.genomes-v5-genome_set-humans-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v5_255_128
Humans CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
539,732 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an exact… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v5_255_128.genomes-v5-genome_set-humans-intervals-v18_255_128
bolinas-dna/genomes-v5-genome_set-humans-intervals-v18_255_128
Humans conserved cCRE enhancers (v18) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
752,548 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-humans-intervals-v18_255_128.genomes-v3-genome_set-humans-intervals-v3_512_256Human-Style-Answers
Human Style Answers
This Datasets contains question and answers on different topics in Human style. (For Chatbots training)
This Datasets is build using TOP AI like (GPT4, Claude3 , Command R+, etc.)
Dataset Details
Description
The Human Style Response Dataset is a rich collection of question-and-answer pairs, meticulously crafted in a human-like style. It serves as a valuable resource for training chatbots and conversational AI models. Let's dive into the… See the full description on the dataset page: https://huggingface.co/datasets/innova-ai/Human-Style-Answers.humansync
RLHF dataset for Large Language Model alignment
Overview
This dataset is designed for training reinforcement learning agents through human feedback, allowing for a more human-aligned response generation. It is in JSON Lines (JSONL) format and contains three key fields: prompt, chosen, and rejected. Each entry provides a prompt followed by a preferred response (chosen) and a less preferred response (rejected), allowing models to learn from human preference patterns.… See the full description on the dataset page: https://huggingface.co/datasets/ParamTh/humansync.Human_StoriesWe took this dataset [https://huggingface.co/datasets/Gryphe/Opus-WritingPrompts] and then downloaded the entire subreddit, made a script to search through it and compiled a new dataset that has human writings in contrast to AI.
We used this dataset to train a classifier which has 95% accuracy you can find here [https://huggingface.co/nothingiisreal/open-gpt-3.5-detector]
However, the main goal of this is instead to remove the watermarks imposed by OpenAI and other AI companies including… See the full description on the dataset page: https://huggingface.co/datasets/nothingiisreal/Human_Stories.HumanSupportSystemMade with Meta Llama 3 🤦
HumanSupportSystem
MAN! Being a human is hard.
Proof of concept on how LIMv01 can be used. Keep licences in mind though.
The instructions and followup were generated using LIM and Llama3-8B generated the responses.
Code example (how it was made)
Llama3 is great at keeping the conversation going, but has limited use for creating datasets that can be used to train models that aren't Llama3. I suppose appending If the instruciton is unclear you… See the full description on the dataset page: https://huggingface.co/datasets/trollek/HumanSupportSystem.genomes-v3-genome_set-humans-intervals-v2_512_256HumanStoryHumanStudyData
