datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indian_Climate_Disaster_data
BharatCRIC: Indian Climate Disaster Data (Heatwave Advisories + Scam Pairs)
Generated by scripts/grade_a_rebuild.py with blueprint grade_a_2026_04_30.
File
Use
instruction_dataset_main.jsonl
Full structured instruction upload (1680 rows)
instruction_dataset_smoke.jsonl
50-row smoke test, 5 languages x 5 formats x 2 rows
instruction_dataset_smoke_v2.jsonl
Same smoke set for the previous filename expected by notes
preference_pairs_scams.jsonl
Mirrored genuine-vs-scam… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/Indian_Climate_Disaster_data.Indian_Climate_Resilience_Instruction_Corpus_
IndianCRIC — Indian Climate Resilience Instruction Corpus
5 languages · 5 formats · genuine ↔ scam pairs
Built for the Adaption Labs Uncharted Data Challenge 2026
Why this dataset exists
The Vulnerable people of Bihar, Uttar Pradesh, and Jharkhand sit at the intersection of high heat vulnerability and low AI coverage.
During extreme weather events, official advisories compete with misinformation — fake helplines, fraudulent relief schemes, and OTP scams disguised… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/Indian_Climate_Resilience_Instruction_Corpus_.ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42
ClimateMBERT Synthetic Qwen3 30B A3B FP8 10K Seed42
Synthetic continuation dataset generated from WxChat/ClimateMBERT_syn train split.
Source dataset: WxChat/ClimateMBERT_syn
Source split: train
Sampling: shuffled with random seed 42, ranks 0..9999
Rows: 10,000
Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8
Inference: vLLM on Clariden GH200 GPUs, tensor parallel size 2, non-eager mode
Max tokens: 4096
Generation config: temperature 0.7, top_p 0.8, top_k 20, min_p 0.0… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42.ClimateMBERT-syn-qwen35-122b-fp8-10k-seed42
ClimateMBERT Synthetic Qwen3.5 FP8 10K Seed42
Synthetic continuation dataset generated from WxChat/ClimateMBERT_syn train split.
Source dataset: WxChat/ClimateMBERT_syn
Source split: train
Sampling: shuffled with random seed 42, ranks 0..9999
Rows: 10,000
Generator: Qwen/Qwen3.5-122B-A10B-FP8
Inference: vLLM on Clariden GH200 GPUs, tensor parallel size 4, non-eager mode
Max tokens: 4096
No-thinking mode: chat_template_kwargs={"enable_thinking": false}
Generation config: temperature… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ClimateMBERT-syn-qwen35-122b-fp8-10k-seed42.
