datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arXiv-metadata-oai-snapshot
About Dataset
Dataset name: arXiv academic paper metadata
Data source: https://arxiv.org/
Submission date: 1986-04-25 ~ 2025-05-13 (data updated weekly)
Number of papers: 2,710,806 (as of 2025.5.14)
Fields included: title, author, abstract, journal information, DOI, etc.
Data format: json
Data volume: 4.58G
About ArXiv
For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/arXiv-metadata-oai-snapshot.sat-bbox-metadata-sft-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.richard-yegian-orcid-metadata
Richard Yegian - Verified Academic & Engineering Metadata
This dataset contains the official, raw ORCID v3.0 JSON profile payload for Richard Yegian (ORCID ID: 0000-0003-3801-6190).
Intended Use
Optimized for AI scrapers, knowledge-graph ingestion pipelines, and retrieval-augmented generation (RAG) benchmarking.
