Saved datasets
Last updated
Download format
Croissant
Croissant is a format for Machine Learning datasets
Learn more about this at mlcommons.org/croissant.
Usage rights
License from data provider
Please review the applicable license to make sure your contemplated use is permitted.
Topic
Provider
Free
Cost to access
Described as free to access or have a license that allows redistribution.
26 datasets found
  1. Social Media & Misinformation Dataset 2024

    • kaggle.com
    zip
    Updated Aug 16, 2025
  2. Indian Political Tweet Engagement Dataset

    • kaggle.com
    zip
    Updated Jan 23, 2026
  3. TruthSocial - 2024 Election Integrity Initiative

    • kaggle.com
    zip
    Updated Nov 1, 2024
  4. Political Social Media Posts

    • kaggle.com
    zip
    Updated Nov 20, 2016
  5. Global Political tweets

    • kaggle.com
    zip
    Updated Aug 23, 2022
  6. US Election 2024 Social Media Sentiment Dataset

    • kaggle.com
    zip
    Updated Sep 15, 2025
  7. Joe Biden's Tweets

    • kaggle.com
    zip
    Updated Dec 19, 2022
  8. Reddit: /r/WTF

    • kaggle.com
    zip
    Updated Dec 18, 2022
  9. Political Inclination Classification Nepali Tweets

    • kaggle.com
    zip
    Updated Nov 23, 2025
  10. donald-trump-truths-dataset

    • kaggle.com
    zip
    Updated Jun 13, 2026
  11. Trump 2024 Campaign Truth Social Truths (Tweets)

    • kaggle.com
    zip
    Updated Dec 15, 2024
  12. Fact-Checking Facebook Politics Pages

    • kaggle.com
    zip
    Updated Jun 5, 2017
  13. Truth Social Reactions to Trump’s Iran War Posts

    • kaggle.com
    zip
    Updated Apr 22, 2026
  14. Word Frequency In Political and Non-Pol. Subreddit

    • kaggle.com
    zip
    Updated Feb 16, 2021
  15. Reddit Data from Before and After Algorithm

    • kaggle.com
    zip
    Updated May 9, 2024
  16. Egypt - Arabic Political 600k Tweets

    • kaggle.com
    zip
    Updated Sep 17, 2019
  17. Narendra Modi Posts After becoming PM

    • kaggle.com
    zip
    Updated Mar 13, 2024
  18. Sound and Audio Data in Montenegro

    • kaggle.com
    zip
    Updated Mar 31, 2025
  19. Sound and Audio Data in Sri Lanka

    • kaggle.com
    zip
    Updated Apr 1, 2025
  20. Sound and Audio Data in Namibia

    • kaggle.com
    zip
    Updated Mar 31, 2025
Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Imaad Mahmood (2025). Social Media & Misinformation Dataset 2024 [Dataset]. https://www.kaggle.com/datasets/imaadmahmood/social-media-and-misinformation-dataset-2024/data
Organization logo

Social Media & Misinformation Dataset 2024

Synthetic dataset of social media posts with engagement, sentiment, misinfo.

Explore at:
zip(4439 bytes)Available download formats
Dataset updated
Aug 16, 2025
Authors
Imaad Mahmood
License

MIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically

Description

📊 Dataset Description: Social Media Content, Engagement & Moderation:

~This dataset contains 40 social media posts collected from multiple platforms (Twitter, Facebook, Instagram, YouTube, TikTok). It provides a detailed view of how different types of content perform, how users engage with them, and how moderation systems respond.

🔑 Key Features:

~**Platform & Content:** Includes post type (Tweet, Story, Video, etc.), unique IDs, and timestamps.

~**User Information:** Follower counts and verification status.

~**Content Metadata:** Text, category, language, country, length, media type, and presence of external links.

~**Engagement Metrics:** Like, share, and comment counts, along with an overall engagement score.

Trust & Safety Signals:

~Misinformation Flag

~Fact-Check Source

~Moderation Action (e.g., Approved, Warning Label, Demonetized, Removed)

NLP & Behavioral Features:

~Sentiment Score (positive/negative tone)

~Toxicity Score (harassment/offensive likelihood)

~Political Leaning (Neutral, Liberal, Conservative, Conspiracy)

~Topic Tags (e.g., climate, vaccine, election, 5G)

~Virality Indicators: Viral score estimating likelihood of content going viral.

📌 Example Use Cases:

~**Fake News & Misinformation Research** – Train ML models to detect misinformation.

~**Content Moderation Systems** – Study how platforms label, remove, or demonetize harmful content.

~**NLP & Sentiment Analysis** – Analyze toxicity, bias, and sentiment across platforms.

~**Trend Analysis** – Compare engagement across topics (climate change, vaccines, elections, 5G).

~**Political Bias Detection** – Explore correlations between political leaning, engagement, and moderation.

📂 Dataset Size:

~40 posts

~25 features

~This dataset is a synthetic but realistic representation of social media activity. It can be useful for machine learning, data analysis, and visualization projects related to misinformation, user engagement, and platform moderation.

Search
Clear search
Close search
Google apps
Main menu