datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
walnutWalnutData
WalnutData
With the gradual maturity of UAV technology, it can provide extremely powerful support for smart agriculture and precise monitoring. Currently, there is no dataset related to green walnuts in the field of agricultural computer vision. Therefore, in order to promote the algorithm design in the field of agricultural computer vision, we used UAV to collect remote sensing data from 8 walnut sample plots. Considering that green walnuts have the characteristics… See the full description on the dataset page: https://huggingface.co/datasets/nanmao/WalnutData.SlimPajama-PT
SlimPajama-6B & SlimPajama-30B
Pre-training data sampled from cerebras/SlimPajama-627B, at two scales: ~6B tokens and ~30B tokens. The data is formatted in JSONL and is compatible with LLaMA-Factory training pipelines.
Repository Structure
├── train/
│ ├── 6B/ # ~6B tokens, 7 JSONL files by source
│ │ ├── RedPajamaCommonCrawl-6B.jsonl
│ │ ├── RedPajamaC4-6B.jsonl
│ │ ├── RedPajamaGithub-6B.jsonl
│ │ ├── RedPajamaBook-6B.jsonl
│… See the full description on the dataset page: https://huggingface.co/datasets/Walnutes/SlimPajama-PT.walnut-edip-sparse20
Scitomo Walnut EDIP Sparse-20 prepared dataset
This is a derived Scitomo Sparse-20 preparation of the public Walnut-1
cone-beam X-ray CT acquisition. It is not the original Walnut archive, not a
published EDIP reconstruction, and not a blessed Scitomo result. The package
contains measured projections, corrected vector-cone geometry, and the
published AGD-50 evaluation reference used by the maintained Scitomo Walnut
EDIP evidence workflow.
Provenance and attribution… See the full description on the dataset page: https://huggingface.co/datasets/scitomo/walnut-edip-sparse20.apple-walnut
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/apple-walnut.rejected-apple-walnut
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-apple-walnut.unjudged-apple-walnut
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-apple-walnut.walnut-rancidity-predictor
Walnut Storage Timeseries Dataset (Indian Conditions)
Synthetic time-series dataset for training walnut rancidity prediction models.
Generated using Arrhenius-based lipid oxidation kinetics simulating real Indian storage scenarios.
Dataset Details
Property
Value
Total rows
5,392,174
Total sequences
90,000
Sequence length
30–90 days
Storage duration
0–180 days
Generation method
Arrhenius kinetics + noise
Storage Scenarios… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/walnut-rancidity-predictor.bigfile_Luxtury_walnutgrasp_walnutThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 30,
"total_frames": 17106,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/StyTJU/grasp_walnut.Think360
Think360: Evaluate Reasoning Capability of Multimodal Large Language Models Beyond Depth
📜 License
The provided code and data are licensed under the Apache 2.0 license.
📝 Citation
If you find this benchmark useful in your research, please consider citing this BibTex:
@misc{chen2026think360degevaluatingwidthcentric,
title={Think 360{\deg}: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth},
author={Mingrui Chen… See the full description on the dataset page: https://huggingface.co/datasets/Walnutes/Think360.walnut2testwalnut
