datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.Chinese-LLaVA-Vision-Instructions本数据集是对于LLaVA的翻译,请从LLaVA dataset下载对应的图片。
百度网盘链接: https://pan.baidu.com/s/1-jgINIkW0MxusmJuSif85w?pwd=q62v
arxiv-software-repo-links-datacite-enrichment-format
arXiv Software Repository Links - DataCite Enrichment Format
A collection of metadata enrichments, formatted for DataCite's enrichment API, that add links between arXiv papers (via DOI) and the software repositories they reference or are supplemented by.
Quick Start
from datasets import load_dataset
ds = load_dataset("cometadata/arxiv-software-repo-links-datacite-enrichment-format")
Dataset Description
Each record is a DataCite-style enrichment instruction… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links-datacite-enrichment-format.arxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/rafidirtiza/arxiv-software-repo-links.chavruta-index-linksds-huffpost-news-article-links
huffpost news article dataset
The dataset consists of >1m links to huffpost news articles. They have been crawled for an educational approach to learn how content could be gathered through reading the websites sitemap(s).
The code can be found here Link
huffpost-crawler
educational - how to use sitemap.xml for crawling a website (huffpost.com - could be any)
Every public website has (or should have) a sitemap.xml. The sitemap.xml allows robots like Google to find links… See the full description on the dataset page: https://huggingface.co/datasets/manzked/ds-huffpost-news-article-links.bilibili_computer_operation_video_download_linkslinks_to_pocasts_lecture_and_shows_for_tts
This dataset was made by Charan from our Discord community. Thank you, very much. :)
license: apache-2.0
linksFineRob_OM-CoT_Instructionlinks-between-papers-and-codepwc-github-links-deduplicatedpaimeng_trainlaw-linkspaimengfortune-telling-DPOpaimeng_dpo
