datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linkedin-job-postingslinkedin_job_listingslinkedin-company-profilelinkedin-job-scrape
LinkedIn DS/ML Job Postings
Daily snapshots of data science / machine learning job postings scraped from LinkedIn, one parquet file per scrape run. All splits share one schema (2023 → today); historical splits were migrated in 2026-07 — the original 8-column data is preserved at revision v1-schema.
The same job_id recurs across splits (a posting stays live for days) — that's the longitudinal signal. For a unique-jobs view, dedup on job_id keeping the row with max scrape_dt.… See the full description on the dataset page: https://huggingface.co/datasets/ryang2/linkedin-job-scrape.LinkedInJobPostings
📊 LinkedIn Job Posting Engagement Analysis
Which LinkedIn job posting characteristics predict candidate engagement (views) — and how well can engagement be predicted or classified using only posting-level features?
Personal motivation: As someone in entrepreneurship, understanding which job posting features attract candidates is directly relevant to future hiring decisions.
📹 Presentation Video
<video… See the full description on the dataset page: https://huggingface.co/datasets/MichaelYitzchak/LinkedInJobPostings.neuml-linkedin-202501
NeuML LinkedIn Company Posts
This dataset is 12 months of NeuML's LinkedIn Company Posts as of January 2025. It contains the post text along with engagement metrics.
It was created as follows:
Export the company posts from the analytics page, see this link for instructions.
Run the following code to create a dataset
import pandas as pd
from datasets import load_dataset
df = pd.read_excel("export_data.xls", sheet_name=1, header=1)
df = df.dropna(axis="columns")… See the full description on the dataset page: https://huggingface.co/datasets/NeuML/neuml-linkedin-202501.linkedin-top-content-scraper-sample-data
LinkedIn Top Content & Top Voices Scraper
Scrapes LinkedIn's public Top Content directory to extract curated high-engagement posts and Top Voice influencers across 40+ categories. Get post text, author profiles, follower counts, reaction metrics, and Top Voice badges. No login, no cookies, no account ban risk. $2 per 1,000 posts.
What the actor scrapes
LinkedIn Top Content & Top Voices Scraper Scrape LinkedIn's public Top Content directory — a curated archive of… See the full description on the dataset page: https://huggingface.co/datasets/logiover/linkedin-top-content-scraper-sample-data.jobs-dataset-linkedinlinkedin-posts-score
LinkedIn Corporate Nonsense Score Dataset
Ein automatisch wachsender Datensatz realer LinkedIn-Posts, bewertet nach ihrem Grad an Corporate Nonsense — gesammelt über die LinkedIn Translator App.
Dataset Details
Beschreibung
Nutzer der App geben LinkedIn-Posts ein um sie auf ihren semantischen Kern zu reduzieren. Jeder Post wird dabei von Llama 4 Maverick automatisch anhand von 5 Metriken bewertet. Die Bewertungen und der vollständige Post-Text… See the full description on the dataset page: https://huggingface.co/datasets/aidn/linkedin-posts-score.linkedin_postslinkedin_postslinkedin_postslinkedin
Dataset Card for "linkedin"
More Information needed
linkedin-jobsMerged files from https://www.kaggle.com/datasets/asaniczka/1-3m-linkedin-jobs-and-skills-2024
MIT_LinkedIns
MIT LinkedIn Profiles — Preview
A small set of MIT-affiliated LinkedIn profiles to test enrichment and matching workflows. For the full dataset with broader coverage and updates, see https://www.thedataoutlet.com.
Buy the full dataset: https://www.thedataoutlet.com
What’s inside
One Excel file with 60 rows and 11 columns.
Person names, education, location, and current role.
File list
MIT_LinkedIns.xlsx 60 rows
Quickstart
import pandas aspd
df =… See the full description on the dataset page: https://huggingface.co/datasets/calebheinzman/MIT_LinkedIns.linkedin_job_listingslinkedin-company-profilemintic_linkedin-job-postingslinkedin-industry-listlinkedinjobsEvery day, thousands of companies and individuals turn to LinkedIn in search of talent. This dataset contains a nearly comprehensive record of 124,000+ job postings listed in 2023 and 2024. Each individual posting contains dozens of valuable attributes for both postings and companies, including the title, job description, salary, location, application URL, and work-types (remote, contract, etc), in addition to separate files containing the benefits, skills, and industries associated with each… See the full description on the dataset page: https://huggingface.co/datasets/lof223/linkedinjobs.linkedin-llama2-datasetlinkedin_profiles_syntheticLinkedInPostDatasetultrachat_200k_deduplinkedin-natural-150
LinkedIn Natural 150 — gemma-3-1b-it style tuning
150 curated LinkedIn posts in a simple, conversational, twitter-like style (no cringe drama).
Composition: 44 twitter-gold short (<50w) + 84 core short-medium (50-80w) + 20 medium-long (80-150w) + 2 long gold. Avg 60.7w.
Format per line (JSONL):
{"prompt": "Write a LinkedIn post about: <topic>", "response": "<natural post>"}
Files:
data/final_dataset.jsonl — 150 full examples (use this)
data/train.jsonl — 138 train… See the full description on the dataset page: https://huggingface.co/datasets/harsh-jos/linkedin-natural-150.Qwen2.5-3B-SFT-pairwise-L_RMlinkedin-ad-library-scraper-sample-data
LinkedIn Ads Library Scraper — No Login, No Cookies
Scrapes LinkedIn's official Ad Library to extract advertiser names, ad copy, headlines, creative images, video posters, and detail URLs. Filter by keyword (advertiser or ad text), 250+ countries, date range, and creative type. No login, no cookies, no account ban risk. $1.50 per 1,000 ads.
What the actor scrapes
🎯 LinkedIn Ad Library Scraper — B2B Competitor Ad Spy, No Login Required Extract every public ad… See the full description on the dataset page: https://huggingface.co/datasets/logiover/linkedin-ad-library-scraper-sample-data.jcvd-or-linkedinllama-2-linkedin-dataLinkedin-company-data
