datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pubmed-embeddings
PubMed Embedding Vectors
This dataset contains embedding vectors generated from local PubMed title and abstract text.
It is designed for biomedical retrieval and nearest-neighbor research.
The public files intentionally do not include PubMed titles, abstracts, or full text.
Rows contain PMIDs, embeddings, hashes, and lightweight metadata so researchers can join
against their own authorized PubMed mirror or the official NCBI/PubMed services.
Configs
Config
Model… See the full description on the dataset page: https://huggingface.co/datasets/aaekay/pubmed-embeddings.AA_Exppubmed-atlas
PubMed Atlas
Hierarchical clustering, MeSH-derived labels, Qwen3.5-27B-generated topic
names, and atlas-driven research-void rankings for all 28,460,827 PubMed
abstracts in the aaekay/pubmed-embeddings
release, computed independently for three encoders:
NeuML/pubmedbert-base-embeddings (768-d, biomedical-specialised)
Qwen/Qwen3-Embedding-0.6B (1024-d, generalist)
BAAI/bge-m3 (1024-d, multilingual generalist)
Joinable to the embedding release by pmid.
Headline numbers… See the full description on the dataset page: https://huggingface.co/datasets/aaekay/pubmed-atlas.image-text_aaeb-xiv-xvii
Dataset Card for image-text_aaeb-xiv-xvii
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 2566 samples across 1 split(s).
Geographical scope: SwitzerlandPeriod: 1400-1500Languages: Early Modern GermanType of document: ProtocolsProvenance: Archives de l'ancien Evêché de Bâle
Projects Included
B_168_14-11_1
B_168_14-11_2
B_168_14-11_3
B_168_14-12
B_168_14-14_1
B_168_14-15_1
B_168_14-15_2… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_aaeb-xiv-xvii.aae2
PIE Dataset Card for "aae2"
This is a PyTorch-IE wrapper for the Argument Annotated Essays v2 (AAE2) dataset (paper and homepage). Since the AAE2 dataset is published in the BRAT standoff format, this dataset builder is based on the PyTorch-IE brat dataset loading script.
Therefore, the aae2 dataset as described here follows the data structure from the PIE brat dataset card.
Usage
from pie_datasets importload_dataset
from pie_datasets.builders.brat import… See the full description on the dataset page: https://huggingface.co/datasets/pie/aae2.olmo-2-0325-32b-preference-mix-aaestupid-paint-aaed43
stupid-paint-aaed43
Synthetic weather test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/charles-harris/stupid-paint-aaed43.aimage-text_aaeb-xiv-xvii-part-2
Dataset Card for transkribus-exports-127147-raw-xml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 121 samples across 1 split(s).
This dataset contains images and transcription from the Archives de l’ancien Evêché de Bâle.
Most texts are in Latin, French, and Early Modern German.
Geographical scope: SwitzerlandPeriod: 1400-1500Languages: Early Modern GermanType of document: ProtocolsProvenance:… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_aaeb-xiv-xvii-part-2.aae-dialect-fairness
AAE Dialect-Fairness Set
A reusable set for debiasing hate/toxicity classifiers against African-American English (AAE) false positives. Off-the-shelf classifiers flag benign AAE text as toxic at 2x+ the rate of benign General-American English (Sap et al. 2019). This set provides (1) high-AAE benign text to augment training so a model can't use dialect as a toxicity cue, and (2) a held-out dialect-balanced benchmark to measure the residual gap.
Built for the… See the full description on the dataset page: https://huggingface.co/datasets/Aeryx-ai/aae-dialect-fairness.dialectic-preferences-bias-aae-sae-parallel
Dialectic Preferences Bias Dataset
Dataset Description
Overview
This dataset is part of a research study examining dialectic preference bias in Large Language Models (LLMs). It contains paired sentences in African American English (AAE) and Standard American English (SAE), used to analyze potential biases in language models' treatment of different dialects.
The dataset contains two columns:
african_american_english: Text samples in African American English… See the full description on the dataset page: https://huggingface.co/datasets/furquan/dialectic-preferences-bias-aae-sae-parallel.Revealwildchat_aaepersonas_grade_math_prompts_generated_aae_1000personas_grade_math_prompts_generated_aae_609personas_grade_math_prompts_generated_aae_607aae-sae-translation-aae_to_sae_initial_5000_result-20250510_123439personas_grade_math_prompts_generated_aae_600personas_grade_math_prompts_generated_aae_503personas_grade_math_prompts_generated_aae_700wildchat_aae_promptsmath_personas_aae_promptsmath_personas_aaepersonas_grade_math_prompts_generated_aae_501personas_grade_math_prompts_generated_aae_504aae-sae-translation-aae_to_sae_next_5000_results-20250514_200105MATH-500-aaeTruthfulQA-aaepersonas_grade_math_prompts_generated_aae_608AAE-v2-essays-qcm
