datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
M3LLM-data-v1.0.0
M3LLM Data
Data for M³LLM training and evaluation on biomedical instruction-following tasks derived from PubMed Central (PMC) articles. This repository is the versioned v1.0.0 data release.
Contents
Collection
Split
Records
Description
PMC-MI supervised instruction corpus
train
224,401
Six instruction formats after partitioning and release filtering
PMC-MI policy-refinement partition
train
10,355
Policy-refinement instances for Stage II after release… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/M3LLM-data-v1.0.0.Japan-successful-bids
🏆 Japanese Successful Bid Texts Dataset
📘 Dataset Summary
This dataset contains a collection of successful Japanese freelance bid (proposal) texts.Each entry represents a message that was actually accepted by a client on freelance job platforms.The dataset is written entirely in natural, polite, and professional Japanese — beginning with “お世話になっております” and ending with “何卒よろしくお願いいたします”.
This data is ideal for:
Fine-tuning text generation models for Japanese business… See the full description on the dataset page: https://huggingface.co/datasets/StellaVision/Japan-successful-bids.thinking_earth_hackathon_bids2025
ThinkingEarth - Harnessing Copernicus Foundation Models to Decode Earth from Space
This repo is part of the ThinkingEarth hackathon, organized at the Big Data from Space 2025 conference in Riga, Latvia. Find our official webpage here.
Hackathon description
The ThinkingEarth Hackathon at BiDS 25, invites AI and Earth observation enthusiasts to explore the power of Copernicus-scale foundation models. Organised by the Horizon Europe project ThinkingEarth, this challenge… See the full description on the dataset page: https://huggingface.co/datasets/franzigrkn/thinking_earth_hackathon_bids2025.bids
