datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
atlas-crispr-10k-benchmark
🧬 ATLAS CRISPR 10k Benchmark
Contribution communauté LeWorldModel — Benchmark CRISPR 10k guides ARNUtilisé pour fine-tuner aguennoune17/negenWM-jepa-v2 — ATLAS NWM Sprint 3Self-supervised I-JEPA · Encodage téléologique (κ, τ, λ)
Description
10 000 guides ARN Cas9 de 20 nucléotides consolidés depuis 12 études expérimentales
de criblage CRISPR génomique à grande échelle. Ce dataset est le benchmark officiel du
Sprint 3 ATLAS NWM v2 — entraînement I-JEPA… See the full description on the dataset page: https://huggingface.co/datasets/aguennoune17/atlas-crispr-10k-benchmark.synthetic-nsclc-10kGameLabel-10kGameLabel-10k Dataset Card
This dataset contains was created in collaboration with the game developers of Armchair Commander. It contains 9800 human preferences over pairs of Flux-Schnell generated images, with over 6800 unique prompts. All labels were crowdsourced from Armchair Commander players.
Usage Example
from datasets import load_dataset
from PIL import Image
import base64
from io import BytesIO
dataset = load_dataset("Jonathan-Zhou/GameLabel-10k")
# For some reason, when using… See the full description on the dataset page: https://huggingface.co/datasets/Jonathan-Zhou/GameLabel-10k.ahsanaseer_top-rated-tmdb-movies-10k
TMDB Movies Dataset
Dataset of 10k top rated TMDB movies for text preprocessing (NLP)
Dataset Info
Source: Kaggle
Original Size: 1.43 MB
Kaggle Downloads: 8,035
Files: 1
Files
top10K-TMDB-movies.csv
Mirrored from Kaggle
DPO_ID-Wiki_10kTesting
HOW TO WRANGLING THIS DATASET TO DPO & CHATML FORMAT
def return_prompt_and_responses(samples) -> dict[str, str, str]:
return {
"prompt": [
"<|im_start|>user\n" + i + "<|im_end|>\n"
for i in samples["PROMPT"]
],
"chosen": [
"<|im_start|>assistant\n" + j + "<|im_end|>"
for j in samples["CHOSEN"]
],
"rejected": [
"<|im_start|>assistant\n" + k + "<|im_end|>"
for k in… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/DPO_ID-Wiki_10kTesting.ifcb-10k10K_bug_reportsOutput-features_10kgnuvid_10kbooks_10kefdtest-10Kclassified_data_10K
path: filepath and unique identifier of the file
company: unique company identifier
year: year of the filing
filename: filename of the 8-K, 10-K, 10-Q given by us in the downloading process
date: date of the filing
paragraph: text paragraph that is analyzed with the LLMs
num_paragraphs: number of paragraphs that the entire filing contained
num_words: number of words that the entire filing contained
Storm, Flood, Heatwave, Drought, Wildfire, Coldwave, physical risk: Indicates 1 if the text… See the full description on the dataset page: https://huggingface.co/datasets/extreme-weather-impacts/classified_data_10K.sp500-synthetic-10k-groundedchelsa_10kThis is a benchmark dataset for regression against a variety of climate variables from the following dataset: https://www.chelsa-climate.org/datasets/chelsa-trace21k-centennial-bioclim. It consists of 10k uniform-at-random sampled points on landmasses, each of which has 8 associated bioclimactic variables from the CHELSA dataset.
All values for bioclim variables are raw values; if using as a regression benchmark, we would recommend min-max normalization.
synthetic_recipes_10k10k_prompts_ranked_allsynthetic_binary_classification_10k10K_MELD_Plus_v1.0
Synthetic MELD-Plus (10K Patients)
Watch a demo
This dataset contains 10,000 synthetic patients inspired by the published MELD-Plus study (a collboration between Massachusetts General Hospital and IBM Research). Each row corresponds to a single admission, with demographics, labs, comorbidities, medications, derived scores (MELD, MELD-Na, MELD-Plus), and the binary outcome Death_Within_90_Days.
All data are artificially generated and contain no identifiable patient records.… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/10K_MELD_Plus_v1.0.gamehistory_10kasf5ytd-10Ksynthetic_ehr_full_10ksen-stu-10ksen-stu-10k2dataset-resto-10kSemantiBench_Dataset_10K
