datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fmars-dataset
FMARS: Foundation Model Annotations for Remote Sensing Images
FMARS is a large-scale dataset of Very High Resolution (VHR) remote sensing images with annotations generated using Vision Foundation Models.
The dataset focuses on disaster management applications and provides pre-event imagery and annotations for major crisis events worldwide from 2021 to 2023.
Paper: https://arxiv.org/abs/2405.20109
Dataset Features
VHR Imagery: The dataset uses pre-event VHR… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/fmars-dataset.wildfire-risk-static-dataarxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.act_phase2align_Varianz_rechts_links_20260916_124805This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_roll.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/OrderDraconis/act_phase2align_Varianz_rechts_links_20260916_124805.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.Spotify_Songs_with_SoundCloud_linksarxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/rafidirtiza/arxiv-software-repo-links.farcaster-links
Farcaster Public Links Dataset
This dataset contains public follow relationships from the Farcaster social protocol that have not been deleted by their authors. The dataset includes comprehensive metadata for each link, allowing for detailed analysis of the social graph in the Farcaster ecosystem.
Dataset Description
The dataset contains the following fields for each active link:
Fid: The Farcaster ID of the user who created the link
MessageType: Type of message (all are… See the full description on the dataset page: https://huggingface.co/datasets/jc4p/farcaster-links.2ch-24-09-2024-no-linkspovarenok_links
Dataset Card for "povarenok_links"
More Information needed
nodes-datasetslinksunten_classifiedlinksuntensoarm_101_2_20260605_153821platform-urls-sample-roberta-tiny-filtered
Dataset Card for "platform-urls-sample-roberta-tiny-filtered"
More Information needed
sample-linksturkish-podcast-links
Turkish Podcast RSS Links
This dataset contains podcast RSS feed metadata for podcasts classified as Turkish by both lingua-py and fast-langdetect. It includes only feed metadata and URLs (e.g., RSS URLs), not any podcast audio or episode content.
Why and how was it built?
I wanted to create this dataset because I could not find an existing one that met my needs. I then downloaded the Podcast Index dataset, which contains 4.5 million podcasts globally. Although entries… See the full description on the dataset page: https://huggingface.co/datasets/hcsolakoglu/turkish-podcast-links.tweets_with_suspicious_linkspodcast-linksvlm-data-with-images-linkslink-scrap3rwikipedia-ref-links-20260101
Wikipedia Multilingual Dataset
This dataset contains 10 language configurations, each with its own train split. The data for each config is stored in its respective subfolder.
Usage Example
from datasets import load_dataset
# Load the 'eng' config
eng_ds = load_dataset("danghaidang-passau/wikipedia-ref-links-20260101", name="eng")
# Load the 'ger' config
ger_ds = load_dataset("danghaidang-passau/wikipedia-ref-links-20260101", name="ger")
LINKS_SMALLLINKS
