datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medstack-medimage-medtech-v1
MedStack Medical Image-Gen Corpus
License-clean, caption-conditioned medical images for training an SDXL per-cluster LoRA
image generator. Built 2026-07-26 · version v0.2 ·
906 images.
License & provenance
Tier A (CC0 / PD, any jurisdiction incl. US/India/global government PD) and Tier B
(CC BY + government-open licenses: India GODL, UK/national OGL, NLOD) only — CC BY-SA / NC / ND / GFDL /
research-only are excluded because a generator emits derivatives. Each row… See the full description on the dataset page: https://huggingface.co/datasets/zeahealth/medstack-medimage-medtech-v1.Maori_English_New_Zealand
Dataset for Translation from Maori to English
The source of this dataset is scraped from the website TEARA. Due to the lack of resources in the Maori language, only a small set of texts are collected, and we are still working on scrapping high quality datasets from other websites.
Dataset Usage
This dataset can only be used for research purposes for NLP tasks (e.g., translation, language identification, etc.)
License
All text is licensed under the Creative… See the full description on the dataset page: https://huggingface.co/datasets/jinglishi0206/Maori_English_New_Zealand.New-Zealand-Stock-Symbols-and-Metadata
New Zealand Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in New Zealand.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/New-Zealand-Stock-Symbols-and-Metadata.astrology-training-corpusdataset_AMLQdata_v1medstack-ayush-instructions-v2
MedStackAI AYUSH Instruction Dataset v2 (long-context, leak-free splits)
Synthetic AYUSH (Ayurveda + Yoga + Integrative Medicine) instruction-tuning dataset
for fine-tuning Mistral-7B-v0.3 into the MedStackAI AYUSH Assistant.
v2 expands each patient into a genuine longitudinal record (8–14 encounters, a
lab-trend table, review of systems, family history, imaging/diagnostic reports,
prior specialist consult notes, preventive/screening and care-coordination
sections). Full… See the full description on the dataset page: https://huggingface.co/datasets/zeahealth/medstack-ayush-instructions-v2.quechua-audioset-4GBdataset_amlq_v2_gpuunlearning-new_zealandmedstack-medreason-eval5-splits
medreason eval5 splits (leak-free)
Deterministic seed=42 train/validation/test split of the governed foundry dataset
(see SPLIT_MANIFEST.json for exact counts + per-domain coverage). The test
split is reserved for the eval-harness and is NEVER seen during training or
early-stopping; training early-stops on validation only. Derived from zeahealth/medstack-medreason-instructions-v1.
new-zealand-dataset-1b
🇳🇿 New Zealand Web Text — 1B-token Sample 🌿
A 1-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-1b.quickjs-hitbreakpointcuzco-quechua-2-spanish-datasetNew_Zealand_Classification
New Zealand Guard Dataset
A binary classification dataset for training guard models to identify whether user inputs are related to New Zealand or not. This dataset is designed for fine-tuning language models to act as content filters, ensuring that only New Zealand-related queries are processed by specialised New Zealand AI assistants.
Dataset Description
The New Zealand Guard Dataset contains 5,000 examples of questions and statements labeled as either:
related: Inputs… See the full description on the dataset page: https://huggingface.co/datasets/geoffmunn/New_Zealand_Classification.medstack-bedside-eval5-splits
bedside eval5 splits (leak-free)
Deterministic seed=42 train/validation/test split of the governed foundry dataset
(see SPLIT_MANIFEST.json for exact counts + per-domain coverage). The test
split is reserved for the eval-harness and is NEVER seen during training or
early-stopping; training early-stops on validation only. Derived from zeahealth/medstack-bedside-instructions-v1.
new-zealand-dataset-5b
🇳🇿 New Zealand Web Text — 5B-token Sample 🌿
A 5-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-5b.dataset_quechua_espanolmlinkspm-semantika-data
