Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
Scraper Code - https://github.com/ArshKA/LinkedIn-Job-Scraper
Every day, thousands of companies and individuals turn to LinkedIn in search of talent. This dataset contains a nearly comprehensive record of 124,000+ job postings listed in 2023 and 2024. Each individual posting contains dozens of valuable attributes for both postings and companies, including the title, job description, salary, location, application URL, and work-types (remote, contract, etc), in addition to separate files containing the benefits, skills, and industries associated with each posting. The majority of jobs are also linked to a company, which are all listed in another csv file containing attributes such as the company description, headquarters location, and number of employees, and follower count.
With so many datapoints, the potential for exploration of this dataset is vast and includes exploring the highest compensated titles, companies, and locations; predicting salaries/benefits through NLP; and examining how industries and companies vary through their internship offerings and benefits. Future updates will permit further exploration into time-based trends, including company growth, prevalence of remote jobs, and demand of individual job titles over time.
Thank you to @zoeyyuzou for scraping an additional 100,000 jobs
job_postings.csv
- job_id: The job ID as defined by LinkedIn (https://www.linkedin.com/jobs/view/ job_id )
- company_id: Identifier for the company associated with the job posting (maps to companies.csv)
- title: Job title.
- description: Job description.
- max_salary: Maximum salary
- med_salary: Median salary
- min_salary: Minimum salary
- pay_period: Pay period for salary (Hourly, Monthly, Yearly)
- formatted_work_type: Type of work (Fulltime, Parttime, Contract)
- location: Job location
- applies: Number of applications that have been submitted
- original_listed_time: Original time the job was listed
- remote_allowed: Whether job permits remote work
- views: Number of times the job posting has been viewed
- job_posting_url: URL to the job posting on a platform
- application_url: URL where applications can be submitted
- application_type: Type of application process (offsite, complex/simple onsite)
- expiry: Expiration date or time for the job listing
- closed_time: Time to close job listing
- formatted_experience_level: Job experience level (entry, associate, executive, etc)
- skills_desc: Description detailing required skills for job
- listed_time: Time when the job was listed
- posting_domain: Domain of the website with application
- sponsored: Whether the job listing is sponsored or promoted.
- work_type: Type of work associated with the job
- currency: Currency in which the salary is provided.
- compensation_type: Type of compensation for the job.
job_details/benefits.csv
- job_id: The job ID
- type: Type of benefit provided (401K, Medical Insurance, etc)
- inferred: Whether the benefit was explicitly tagged or inferred through text by LinkedIn
company_details/companies.csv
- company_id: The company ID as defined by LinkedIn
- name: Company name
- description: Company description
- company_size: Company grouping based on number of employees (0 Smallest - 7 Largest)
- country: Country of company headquarters.
- state: State of company headquarters.
- city: City of company headquarters.
- zip_code: ZIP code of company's headquarters.
- address: Address of company's headquarters
- url: Link to company's LinkedIn page
company_details/employee_counts.csv
- company_id: The company ID
- employee_count: Number of employees at company
- follower_count: Number of company followers on LinkedIn
- time_recorded: Unix time of data collection
If you find this dataset helpful, your upvote would convince me I didn't waste my summer break 😁
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
This dataset contains time-stamped AI & ML job postings scraped from LinkedIn and Indeed over multiple days, covering companies, roles, and locations. It includes:
link: URL to the job postingtitle: Job title (e.g., Data Scientist, ML Engineer)company: Company namelocation: City, state, or countrydate & time: Job posting timestampscrape_date & scrape_time: When the data was collectedDataset Highlights: - ~1,550 unique postings, clean and deduplicated - Ready for EDA, visualization, and ML experiments - Includes scrape metadata for temporal analysis
Potential Use Cases: - Trend analysis of AI/ML hiring over time - Skill extraction and NLP on job titles - Job classification or predictive modeling projects - Company hiring insights and labor market research - Geospatial analysis of AI/ML demand
Included Notebook: EDA_Job_Postings.ipynb
- Exploratory data analysis with top companies, job titles, locations, and word clouds
- Time-series analysis of job postings
License: CC BY 4.0 — free for research, educational, and analysis purposes with attribution.
Note: Data was collected via public job postings; no personal candidate information is included. Users can further enrich the dataset using the job links if legally permissible.
Facebook
TwitterOpen Data Commons Attribution License (ODC-By) v1.0https://www.opendatacommons.org/licenses/by/1.0/
License information was derived automatically
LinkedIn is a popular professional networking platform with millions of job postings across various industries.
This dataset provides a raw dump of data science-related job postings collected from LinkedIn. It includes information about job titles, companies, locations, search parameters, and other relevant details.
The main objective of this dataset is not only to provide insights into the data science job market and the skills required by professionals in this field but also to offer users an opportunity to practice their data cleaning skills.
By working with this dataset, users can gain hands-on experience in cleaning and preprocessing raw data, a critical skill for aspiring data scientists.
If you find this dataset useful or interesting, please upvote it! 😊💝
Photo by Luke Chesser on Unsplash
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by Muhammad Sufyan
Released under CC0: Public Domain
Facebook
TwitterThis dataset contains a curated collection of job listings sourced from LinkedIn, featuring a variety of positions across multiple industries and locations. Each entry includes essential details such as job title, company name, job location, employment type, and base pay range, alongside a comprehensive job summary and required qualifications.
This dataset is ideal for researchers, data scientists, and job seekers looking to analyze job market trends, understand salary expectations, or develop predictive models for career growth. Use this resource to gain insights into the evolving job landscape and make informed career decisions.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This comprehensive dataset contains a curated collection of 100 job postings for Data Analyst positions, sourced from LinkedIn. As the demand for skilled data analysts continues to surge, this dataset serves as a valuable resource for data enthusiasts, aspiring data analysts, and researchers alike.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
This dataset contains information about job postings on LinkedIn. The data is divided into several files, each containing different aspects of the job postings:
job_postings.csv: This file contains detailed information about each job posting, including the job title, description, salary, work type, location, and more.companies.csv: This file contains detailed information about each company that posted a job, including the company name, website, description, size, location, and more.company_industries.csv: This file contains the industries associated with each company.company_specialities.csv: This file contains the specialties associated with each company.employee_counts.csv: This file contains the employee and follower counts for each company.benefits.csv: This file contains the benefits associated with each job.job_industries.csv: This file contains the industries associated with each job.job_skills.csv: This file contains the skills associated with each job.This dataset can be used for various purposes such as: - Analyzing the job market - Analyzing company trends - Analyzing salary trends - Building a job recommendation system - Natural Language Processing (NLP) tasks such as keyword extraction, topic modeling, etc.
This dataset was collected from LinkedIn. Please note that the data may be subject to LinkedIn's terms of use.
This dataset is released under the Open Database License (ODbL).
Facebook
TwitterThis dataset was created by davideev9
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
The data comprises job-related information from LinkedIn job postings scraped over a 2-day period. Key features include company details and job-specific information like title, description, and salary. The dataset provides a comprehensive view for exploring factors influencing job posting characteristics and has been reformatted from its original source to improve its compatibility among various machine learning algorithms.
Facebook
TwitterThis dataset was created by Steve Marcello Liem
Facebook
TwitterAttribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
License information was derived automatically
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Collated 8 different data source. Filtered for only Data and ML jobs, titles and descriptions. Applied text data cleaning and preprocessing, documented here: https://tianyimasf.github.io/ai/data-cleaning/.
LinkedIn-Tech-Job-Data: A compilation of job posts and metadata scraped from various tech categories on LinkedIn
Data Analyst Jobs: This dataset was created by picklesueat and contains more than 2000 job listing for data analyst positions
US Job Postings from 2023-05-05: This dataset is an excerpt of our web scraping activities at Techmap.io and contains a sample of 33k Job Postings from the USA on May 5th 2023.
LinkedIn Job Postings Dataset: This dataset contains information about job postings on LinkedIn.
LinkedIn Job Postings - Machine Learning Data Set: The data comprises job-related information from LinkedIn job postings scraped over a 2-day period.
Linkedin Canada: Data Science Jobs 2024: The "LinkedIn Canada: Data Science Jobs 2024" dataset presents an insightful overview of the data science job market in Canada as sourced from LinkedIn.
Data Scientist - Linkedin Job Postings: This dataset provides valuable insights into data science job postings, including the required skills and software proficiency sought by employers.
LinkedIn Job Postings Dataset: This dataset contains information about job postings on LinkedIn.
Initially used for my project analyzing data job market, including analyzing titles, skills, and company functions. Could be used for other purposes like posting generation.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset was created by Alex Ma
Released under Apache 2.0
Facebook
TwitterAttribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0)https://creativecommons.org/licenses/by-nc-sa/3.0/
License information was derived automatically
This dataset was created by AJ Strauman-Scott
Released under Attribution-NonCommercial-ShareAlike 3.0 IGO (CC BY-NC-SA 3.0 IGO)
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset is a controlled snapshot of the IBM HR Analytics Attrition dataset.
Source: https://www.kaggle.com/datasets/arshkon/linkedin-job-postings
Purpose: This copy is maintained to ensure reproducibility and stability of the analysis, avoiding dependency on external dataset changes.
Notes:
No transformations have been applied. This dataset represents a fixed version used in the project pipeline.
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
**💼 LinkedIn Jobs Sample Dataset – May 2026 ** This dataset contains 28,000 structured job listings collected from LinkedIn using a scalable data pipeline built on Apify.
It is designed as a high-quality sample dataset demonstrating the structure, richness, and scalability of real-world job market data for analytics, machine learning, and recruitment intelligence.
🧾 Dataset Overview
Each row represents a job listing and includes structured attributes such as:
jobTitle – Job role/title companyName – Hiring company location – Job location employmentType – Full-time, Part-time, Contract, etc. experienceLevel – Entry, Mid, Senior (if available) salary / compensation – Salary info (if available) jobDescription – Full job description text skills / keywords – Extracted or listed skills postedDate – Job posting date applicants / insights – Engagement indicators (if available) jobUrl – Direct job listing link jobId / referenceId – Unique identifier
✔️ Clean, structured, and ready for analysis (CSV / Excel / JSON compatible)
⚙️ Data Collection Method
This dataset is generated using custom-built scraping actors:
👉 Powered by Apify 👉 URL: https://apify.com/shahidirfan/fast-linkedin-job-scraper
Pipeline highlights:
Scalable large-volume extraction Structured normalization of job data Duplicate filtering & validation Flexible filters (location, keywords, job type) Batch or real-time delivery 📈 Why This Dataset Is Useful ✅ Real-world job market signals ✅ Rich text data for NLP (job descriptions, skills) ✅ Company + role-level insights ✅ Time-based data for trend analysis
This makes it suitable for serious analytical and commercial use, not just basic datasets.
🚀 Use Cases Job market trend analysis Skill demand analysis (NLP / text mining) Salary benchmarking (where available) Job recommendation systems Recruitment intelligence platforms AI/ML models for career insights 📦 Need Larger or Custom Job Datasets?
This dataset (28,000 rows) is only a sample preview.
I can provide:
Millions of job listings Country or industry-specific datasets Historical + real-time pipelines Scheduled scraping (daily/hourly updates) API-based delivery Custom enrichment (skills extraction, classification, etc.) 🤝 Work With Me
I’m an experienced Apify Actor Builder, specializing in:
Job board & professional platform scraping Large-scale data pipelines Automation & API integrations Custom data solutions
👉 https://apify.com/shahidirfan
If you need bulk datasets or custom scraping solutions, feel free to reach out.
Facebook
Twitterhttp://opendatacommons.org/licenses/dbcl/1.0/http://opendatacommons.org/licenses/dbcl/1.0/
The jobs_linkedin.csv file comprises the scraping results obtained from LinkedIn website. It includes the following columns:
1. title: Signifies the job title associated with each entry.
2. location: Provides information about the job's location.
3. time: Indicates the timestamp when the job post was uploaded.
4. link: Contains a unique identifier (UUID) and a direct link to the respective job post.
5. desc: Contains the comprehensive description of each job opportunity.
For a more detailed exploration of my NLP work, please refer to: - LinkedIn-NLP-Notebook - LinkedIn-NLP&DL-Notebook
Facebook
TwitterThis dataset contains job postings from Linkedin from 2023 with the following features It can be used to analyze the current trends based on job positions, location,company nameetc
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset provides an extensive collection of job listings from LinkedIn, covering the period from August 15th to August 31st, 2023. With nearly 3 million records, this dataset is a valuable resource for HR professionals, data scientists, labor market analysts, and researchers looking to explore job market trends, skills demand, and employment opportunities across various industries.
Key Features:
Job Information: Detailed job titles, descriptions, categories, and unique job identifiers. Company Data: Information on hiring companies, including company names, industries, and locations. Job Requirements: Data on required skills, qualifications, experience levels, and employment types. Job Locations: Comprehensive coverage of job locations across various countries and regions. Application Details: Information on application processes and deadlines (if available). Time of Data Collection: Listings are time-stamped, reflecting the job postings available from August 15th to August 31st, 2023.
Dataset Overview:
Price: $2500.0 Total Records Count: 2,862,984 Domain Name: LinkedIn Date Range: August 15th, 2023 - August 31st, 2023 File Extension: LDJSON (Line-Delimited JSON)
Use Cases:
Job Market Analysis: Analyze job market trends, in-demand skills, and emerging roles across industries. Recruitment Strategies: Optimize recruitment strategies by understanding the competitive landscape and job posting patterns. Skill Demand Analysis: Identify the most sought-after skills and qualifications in various sectors. Labor Market Research: Conduct in-depth research on employment opportunities, job availability, and geographic job distribution. Economic Indicators: Use job listings data as an indicator of economic health and employment trends.
About the Data Collection:
This dataset was collected using sophisticated web scraping techniques to ensure comprehensive coverage and accuracy. For businesses and researchers needing customized data or large-scale web extraction from platforms like LinkedIn, PromptCloud offers bespoke web scraping solutions. These services cater to specific needs, delivering high-quality, structured data tailored to unique research and business objectives. https://www.promptcloud.com/contact/
Disclaimer: This dataset is intended for research and educational purposes. Users are responsible for ensuring that their use of this data complies with LinkedIn’s terms of service and all applicable laws.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset was created by RATNESH SATYARTHI
Released under Apache 2.0
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
Scraper Code - https://github.com/ArshKA/LinkedIn-Job-Scraper
Every day, thousands of companies and individuals turn to LinkedIn in search of talent. This dataset contains a nearly comprehensive record of 124,000+ job postings listed in 2023 and 2024. Each individual posting contains dozens of valuable attributes for both postings and companies, including the title, job description, salary, location, application URL, and work-types (remote, contract, etc), in addition to separate files containing the benefits, skills, and industries associated with each posting. The majority of jobs are also linked to a company, which are all listed in another csv file containing attributes such as the company description, headquarters location, and number of employees, and follower count.
With so many datapoints, the potential for exploration of this dataset is vast and includes exploring the highest compensated titles, companies, and locations; predicting salaries/benefits through NLP; and examining how industries and companies vary through their internship offerings and benefits. Future updates will permit further exploration into time-based trends, including company growth, prevalence of remote jobs, and demand of individual job titles over time.
Thank you to @zoeyyuzou for scraping an additional 100,000 jobs
job_postings.csv
- job_id: The job ID as defined by LinkedIn (https://www.linkedin.com/jobs/view/ job_id )
- company_id: Identifier for the company associated with the job posting (maps to companies.csv)
- title: Job title.
- description: Job description.
- max_salary: Maximum salary
- med_salary: Median salary
- min_salary: Minimum salary
- pay_period: Pay period for salary (Hourly, Monthly, Yearly)
- formatted_work_type: Type of work (Fulltime, Parttime, Contract)
- location: Job location
- applies: Number of applications that have been submitted
- original_listed_time: Original time the job was listed
- remote_allowed: Whether job permits remote work
- views: Number of times the job posting has been viewed
- job_posting_url: URL to the job posting on a platform
- application_url: URL where applications can be submitted
- application_type: Type of application process (offsite, complex/simple onsite)
- expiry: Expiration date or time for the job listing
- closed_time: Time to close job listing
- formatted_experience_level: Job experience level (entry, associate, executive, etc)
- skills_desc: Description detailing required skills for job
- listed_time: Time when the job was listed
- posting_domain: Domain of the website with application
- sponsored: Whether the job listing is sponsored or promoted.
- work_type: Type of work associated with the job
- currency: Currency in which the salary is provided.
- compensation_type: Type of compensation for the job.
job_details/benefits.csv
- job_id: The job ID
- type: Type of benefit provided (401K, Medical Insurance, etc)
- inferred: Whether the benefit was explicitly tagged or inferred through text by LinkedIn
company_details/companies.csv
- company_id: The company ID as defined by LinkedIn
- name: Company name
- description: Company description
- company_size: Company grouping based on number of employees (0 Smallest - 7 Largest)
- country: Country of company headquarters.
- state: State of company headquarters.
- city: City of company headquarters.
- zip_code: ZIP code of company's headquarters.
- address: Address of company's headquarters
- url: Link to company's LinkedIn page
company_details/employee_counts.csv
- company_id: The company ID
- employee_count: Number of employees at company
- follower_count: Number of company followers on LinkedIn
- time_recorded: Unix time of data collection
If you find this dataset helpful, your upvote would convince me I didn't waste my summer break 😁