Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
~This dataset contains 40 social media posts collected from multiple platforms (Twitter, Facebook, Instagram, YouTube, TikTok). It provides a detailed view of how different types of content perform, how users engage with them, and how moderation systems respond.
~**Platform & Content:** Includes post type (Tweet, Story, Video, etc.), unique IDs, and timestamps.
~**User Information:** Follower counts and verification status.
~**Content Metadata:** Text, category, language, country, length, media type, and presence of external links.
~**Engagement Metrics:** Like, share, and comment counts, along with an overall engagement score.
~Misinformation Flag
~Fact-Check Source
~Moderation Action (e.g., Approved, Warning Label, Demonetized, Removed)
~Sentiment Score (positive/negative tone)
~Toxicity Score (harassment/offensive likelihood)
~Political Leaning (Neutral, Liberal, Conservative, Conspiracy)
~Topic Tags (e.g., climate, vaccine, election, 5G)
~Virality Indicators: Viral score estimating likelihood of content going viral.
~**Fake News & Misinformation Research** – Train ML models to detect misinformation.
~**Content Moderation Systems** – Study how platforms label, remove, or demonetize harmful content.
~**NLP & Sentiment Analysis** – Analyze toxicity, bias, and sentiment across platforms.
~**Trend Analysis** – Compare engagement across topics (climate change, vaccines, elections, 5G).
~**Political Bias Detection** – Explore correlations between political leaning, engagement, and moderation.
~40 posts
~25 features
~This dataset is a synthetic but realistic representation of social media activity. It can be useful for machine learning, data analysis, and visualization projects related to misinformation, user engagement, and platform moderation.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains 85,154 Twitter posts related to Indian political discourse, collected between October 2022 and March 2023. It includes tweet text, user identifiers, temporal metadata, and engagement metrics such as likes and retweets, enabling analysis of interaction patterns and engagement behavior in high-activity public discussions.
The dataset consists of 9 variables and is fully cleaned, with no missing values, duplicate records, or invalid timestamps. Derived temporal features (Year, Month, Day) are perfectly consistent with the original timestamp, ensuring reliability for time-based analysis.
Multiple forensic checks were performed to evaluate whether engagement metrics reflect real-world social media behavior:
Engagement metrics display realistic long-tail distributions, with a small fraction of highly engaged tweets and minimal zero inflation. The dataset contains over 58,000 distinct users and more than 98% unique tweet content, further supporting data authenticity.
In addition to tweet-level data, the dataset includes a user interaction network represented as directed edges. Each edge denotes an interaction between two users, derived from observable Twitter actions such as replies, mentions, or retweets. This network structure enables graph-based analysis of information flow, influence patterns, and community behavior within political discussions.
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22766094%2Fea12c83af697260cda6a5dfcdd6c544b%2Fgraph.jpg?generation=1769145568146421&alt=media" alt="">
The dataset is suitable for machine learning and analytical tasks such as engagement prediction, content analysis, user behavior modeling, and temporal interaction studies. Political content is treated solely as a high-engagement discussion domain and does not imply ideological inference or endorsement.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset, from Crowdflower's Data For Everyone Library, provides text of 5000 messages from politicians' social media accounts, along with human judgments about the purpose, partisanship, and audience of the messages.
Contributors looked at thousands of social media messages from US Senators and other American politicians to classify their content. Messages were broken down into audience (national or the tweeter’s constituency), bias (neutral/bipartisan, or biased/partisan), and finally tagged as the actual substance of the message itself (options ranged from informational, announcement of a media appearance, an attack on another candidate, etc.)
Data was provided by the Data For Everyone Library on Crowdflower.
Our Data for Everyone library is a collection of our favorite open data jobs that have come through our platform. They're available free of charge for the community, forever.
Here are a couple of questions you can explore with this dataset:
The dataset contains one file, with the following fields:
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
The US Election 2024 Social Media Sentiment Dataset captures 100 authentic, anonymized posts from X (formerly Twitter) collected during November 5-6, 2024, coinciding with the US Presidential Election's critical period. This dataset reflects real-time public opinions, emotions, and discussions surrounding the election, focusing on candidates (Donald Trump, Kamala Harris), voting processes, and media narratives. Sourced via X's official API, the data ensures compliance with platform policies and prioritizes ethical considerations by anonymizing user identities.
Size: 100 unique posts (excluding replies and quoted posts to avoid redundancy).
Attributes:
Posts were collected using X's API with targeted queries (e.g., "#USElection2024", "Trump", "Harris" -filter:replies) and a minimum engagement filter (min_faves:1) to ensure relevance. The dataset was cleaned to remove sensitive information (e.g., full URLs where non-essential) while retaining original text for analysis. The collection focused on the latest posts to capture real-time reactions.
This dataset is a valuable resource for data scientists, political researchers, and students studying social media’s impact on the 2024 US Presidential Election. It provides a snapshot of public discourse, ideal for NLP, social network analysis, and trend detection.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
By Reddit [source]
Immerse yourself in the world of Reddit's Subreddit WTF with this comprehensive dataset! This dataset offers a unique glimpse into the mysterious world of Reddit, giving researchers access to 8 columns packed with essential data such as titles, scores, comments counts and URLs, creation dates and times as well as post bodies and timestamps. With this data in hand, researchers can begin uncovering secrets buried beneath Reddit's user-generated content. Uncover trends that go beyond simple sentiment analysis or explore topics important to the community! The possibilities available for research are abundant; it's time to start uncovering the answers you need from this exciting data set
For more datasets, click here.
- 🚨 Your notebook can be here! 🚨!
Using this dataset, you can uncover previously unseen insights on Reddit's Subreddit WTF. To get the most out of this dataset, you should familiarize yourself with the columns: title, score, url, comms_num (comment numbers), created (creation date and time), body (the posts body) and timestamp. Each column is filled with valuable information that you can use to uncover hidden trends and topics of discussion within Reddit’s Subreddit WTF community.
Start by exploring the data to gain an understanding of how Reddit works. Look for patterns such as hot topics being discussed or particular types of posts being upvoted more frequently than others. Understanding these patterns will help you build a better picture of what’s going on in Reddit’s Subreddit WTF community. Once you understand the data, start mining it for relevant metadata—such as sentiment analysis or keyword searches—that will add deeper insight into its user-generated content. By answering questions such as “what topics are people discussing?” or “what type of posts are generating more buzz?” Researchers have the capability to dive deep into related topics with much accuracy thanks to this comprehensive Dataset!
Depending on your research needs; it might also be worth combining different columns from the dataset like URL + Title/Body = something contextual - to unlock insights that help answer a bigger research question when investigated together multiple pieces of data together .This method might be key in uncovering hidden gems from within this dataset! For example; if studying political impact on users content - then analyzing post bodies combined with their dates & timestamps could reveal some interesting trends about increase/decrease of online activity in relation political events happening around a certain period in times cases).
- Analyzing the Sentiment of Users: By studying the language and body of Reddit posts, researchers can analyze sentiment and uncover any potential biases or patterns in user words. This could be used to infer subtle changes in sentiment or overall sentiment at a given point in time, as well as spot emerging topics or controversies.
- Uncover Popular Topics: Through analyzing the titles, topics discussed on WTF subreddit could be determined, allowing for insight into what content might be popular on Reddit overall. This could also help reveal insight into what type of content is widely accepted by Reddit audiences and'smaller subcultures.'
- Tracking User Engagement: By studying scores and comment counts over time, researcherse can track user engagement with posts over time to see when users are most likely to comment on posts or interact with one another. This could help shed light on user habits and preferences when it comes to engaging with content across different platforms like Reddit or other social media sites
If you use this dataset in your research, please credit the original authors. Data Source
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: WTF.csv | Column name | Description | |:--------------|:--------------------------------------------------------| | title | The title of the post. (String) | | score | The number of upvotes the post has received. (Integer) | | url | The URL of the post. (String) ...
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
To use this dataset on your research paper use the following reference.
@artical{s13102024ijcatr13101005,
Title = "Comparing Political Inclination Classification on Twitter Posts using Naive Bayes, SVM, and XGBoost",
Journal ="International Journal of Computer Applications Technology and Research(IJCATR)",
Volume = "13",
Issue ="10",
Pages ="62 - 65",
Year = "2024",
Authors ="Shashank Shree Neupane, Atish Shakya, Bishan Rokka, Sagar Acharya"}
The details of the article is:
International Journal of Computer Applications Technology and Research Volume 13–Issue 10, 62 – 65, 2024, ISSN:-2319–8656 DOI:10.7753/IJCATR1310.1005
The link to article: https://ijcat.com/archieve/volume13/issue10/ijcatr13101005
The dataset contains the twitter post of nepali political leader who are on political parties. The dataset can be used to know the inclination of people towards a political party with their post on the social media such as X (formerly twitter).
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by Anjay23
Released under CC0: Public Domain
It contains the following files:
Facebook
TwitterDuring the 2016 US presidential election, the phrase “fake news” found its way to the forefront in news articles, tweets, and fiery online debates the world over after misleading and untrue stories proliferated rapidly. BuzzFeed News analyzed over 1,000 stories from hyperpartisan political Facebook pages selected from the right, left, and mainstream media to determine the nature and popularity of false or misleading information they shared.
This dataset supports the original story “Hyperpartisan Facebook Pages Are Publishing False And Misleading Information At An Alarming Rate” published October 20th, 2016. Here are more details on the methodology used for collecting and labeling the dataset (reproduced from the story):
More on Our Methodology and Data Limitations
“Each of our raters was given a rotating selection of pages from each category on different days. In some cases, we found that pages would repost the same link or video within 24 hours, which caused Facebook to assign it the same URL. When this occurred, we did not log or rate the repeat post and instead kept the original date and rating. Each rater was given the same guide for how to review posts:
“*Mostly True*: The post and any related link or image are based on factual information and portray it accurately. This lets them interpret the event/info in their own way, so long as they do not misrepresent events, numbers, quotes, reactions, etc., or make information up. This rating does not allow for unsupported speculation or claims.
“*Mixture of True and False*: Some elements of the information are factually accurate, but some elements or claims are not. This rating should be used when speculation or unfounded claims are mixed with real events, numbers, quotes, etc., or when the headline of the link being shared makes a false claim but the text of the story is largely accurate. It should also only be used when the unsupported or false information is roughly equal to the accurate information in the post or link. Finally, use this rating for news articles that are based on unconfirmed information.
“*Mostly False*: Most or all of the information in the post or in the link being shared is inaccurate. This should also be used when the central claim being made is false.
“*No Factual Content*: This rating is used for posts that are pure opinion, comics, satire, or any other posts that do not make a factual claim. This is also the category to use for posts that are of the “Like this if you think...” variety.
“In gathering the Facebook engagement data, the API did not return results for some posts. It did not return reaction count data for two posts, and two posts also did not return comment count data. There were 70 posts for which the API did not return share count data. We also used CrowdTangle's API to check that we had entered all posts from all nine pages on the assigned days. In some cases, the API returned URLs that were no longer active. We were unable to rate these posts and are unsure if they were subsequently removed by the pages or if the URLs were returned in error.”
This dataset was originally published on GitHub by BuzzFeed News here: https://github.com/BuzzFeedNews/2016-10-facebook-fact-check
Here are some ideas for exploring the hyperpartisan echo chambers on Facebook:
How do left, mainstream, and right categories of Facebook pages differ in the stories they share?
Which types of stories receive the most engagement from their Facebook followers? Are videos or links more effective for engagement?
Can you replicate BuzzFeed’s findings that “the least accurate pages generated some of the highest numbers of shares, reactions, and comments on Facebook”?
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
This dataset contains a curated collection of Donald Trump’s Truth Social posts related to Iran, military escalation, and Middle East conflict rhetoric, along with the associated comment data collected for those posts.
The dataset is structured in two linked tables:
The main purpose of this dataset is to support research and exploratory analysis of:
posts.csv Contains post-level metadata such as post ID, timestamp, cleaned text, engagement counts, media indicators, account fields, and link-card information.
comments.csv Contains comment-level metadata such as comment ID, parent post ID, cleaned text, engagement counts, commenter account fields, media indicators, and link-card information.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
https://techcrunch.com/wp-content/uploads/2015/10/twitter-politics.png" alt="">
Social media is becoming a key medium through which we communicate with each other: it is at the center of the very structures of our daily interactions. Yet this infiltration is not unique to interpersonal relations. Political leaders, governments, and states operate within this social media environment, wherein they continually address crises and institute damage control through platforms such as Twitter.
With the proliferation of the internet into mass masses, social media is emerging as a potential way of communication. It provides a direct channel to politicians for communicating, connecting, and engaging with the public. The power of social media, especially Twitter and Facebook has been proved by its successful application during recent US presidential elections and Arabian countries' revolts. In India too, as the general election is about to knock at the door during early 2014, political parties and leaders are trying to harness the power of social media.
The tweets have the #Politics hashtag. The collection started on 24/7/2021, and will be updated on a daily basis.
The data totally consists of 1 lakh+ records with 13 columns. The description of the features is given below | No |Columns | Descriptions | | -- | -- | -- | | 1 | user_name | The name of the user, as they’ve defined it. | | 2 | user_location | The user-defined location for this account’s profile. | | 3 | user_description | The user-defined UTF-8 string describing their account. | | 4 | user_created | Time and date, when the account was created. | | 5 | user_followers | The number of followers an account currently has. | | 6 | user_friends | The number of friends an account currently has. | | 7 | user_favourites | The number of favorites an account currently has | | 8 | user_verified | When true, indicates that the user has a verified account | | 9 | date | UTC time and date when the Tweet was created | | 10 | text | The actual UTF-8 text of the Tweet | | 11 | hashtags | All the other hashtags posted in the tweet along with #Politics | | 12 | source | Utility used to post the Tweet, Tweets from the Twitter website have a source value - web | | 13 | is_retweet | Indicates whether this Tweet has been Retweeted by the authenticating user. |
You can use this data to dive into the subjects that use this hashtag, look to the geographical distribution, evaluate sentiments, and look at trends.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Techsalerator's News Events Data for Cameroon: A Comprehensive Overview
Techsalerator's News Events Data for Cameroon offers a valuable resource for businesses, researchers, and media organizations. This dataset gathers information on significant news events across Cameroon, sourced from a variety of media outlets, online publications, and social platforms. It provides crucial insights for those interested in tracking trends, analyzing public sentiment, or monitoring industry-specific developments.
Key Data Fields
Event Date: Records the exact date of the news event. This is essential for analysts tracking trends over time or for businesses responding to market changes.
Event Title: A concise headline summarizing the event. This allows users to quickly categorize and evaluate news content based on its relevance.
Source: Identifies the news outlet or platform where the event was reported. This helps users track credible sources and assess the event's reach and influence.
Location: Provides geographic details on where the event occurred within Cameroon. This is particularly useful for regional analysis or localized marketing strategies.
Event Description: A detailed summary of the event, highlighting key developments, participants, and potential impacts. Researchers and businesses use this to understand the context and implications of the event.
Top 5 News Categories in Cameroon
Politics: Coverage of government decisions, political movements, elections, and policy changes affecting the national landscape.
Economy: Focuses on Cameroon’s economic indicators, inflation rates, international trade, and corporate activities impacting the business and finance sectors.
Social Issues: News related to protests, public health, education, and other societal concerns driving public discourse.
Sports: Highlights events in football, basketball, and other popular sports, often capturing significant public interest and engagement.
Technology and Innovation: Reports on technological advancements, startups, and innovations within Cameroon’s growing tech ecosystem.
Top 5 News Sources in Cameroon
Cameroon Tribune: A leading newspaper providing comprehensive coverage of politics, economy, and social issues in Cameroon.
Equinoxe Television: A major news network offering updates on current affairs, politics, and sports through its television and online platforms.
Le Jour: A prominent publication focusing on political news, business developments, and societal issues.
The Guardian Post: Known for its detailed reporting on Cameroonian politics, business, and social topics.
Cameroun Web: A significant online news platform offering real-time updates on breaking news, sports, and entertainment.
Accessing Techsalerator’s News Events Data for Cameroon
To access Techsalerator’s News Events Data for Cameroon, please contact info@techsalerator.com with your specific needs. We will provide a customized quote based on the data fields and records required, with delivery available within 24 hours. Ongoing access options can also be discussed.
Included Data Fields
Techsalerator’s dataset is an invaluable tool for monitoring significant events in Cameroon. It supports informed decision-making, whether for business strategy, market analysis, or academic research, offering a clear view of the country’s news landscape.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Techsalerator’s Location Sentiment Data for Namibia
Techsalerator’s Location Sentiment Data for Namibia provides in-depth insights into the emotions, opinions, and sentiment trends across different regions of the country. This dataset is essential for businesses, researchers, and policymakers looking to understand public perception, consumer behavior, and regional sentiment variations.
For access to the full dataset, contact us at info@techsalerator.com or visit Techsalerator Contact Us.
Techsalerator’s Location Sentiment Data for Namibia delivers structured sentiment analysis derived from social media, news sources, and consumer reviews. This dataset is valuable for market research, social analytics, and economic development strategies.
To obtain Techsalerator’s Location Sentiment Data for Namibia, contact info@techsalerator.com with your specific requirements. Techsalerator provides customized datasets based on requested fields, with delivery available within 24 hours. Ongoing access options can also be discussed.
For comprehensive sentiment analysis and location-based insights in Namibia, Techsalerator’s dataset serves as a valuable resource for businesses, researchers, and policymakers.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Techsalerator's News Events Data for Cambodia: A Comprehensive Overview
Techsalerator's News Events Data for Cambodia provides an essential resource for businesses, researchers, and media organizations. This dataset compiles information on key news events across Cambodia, drawing from a diverse array of media sources, including news outlets, online publications, and social media platforms. It offers valuable insights for those interested in tracking trends, analyzing public sentiment, or monitoring industry-specific developments.
Key Data Fields - Event Date: Records the precise date of the news event, crucial for trend analysis over time or for businesses reacting to market changes. - Event Title: A concise headline describing the event, enabling users to quickly gauge and categorize news content based on relevance. - Source: Identifies the news outlet or platform reporting the event, helping users track credible sources and evaluate the event's reach and influence. - Location: Provides geographic details about where the event occurred within Cambodia, valuable for regional analysis or targeted marketing. - Event Description: Offers a detailed summary of the event, including key developments, participants, and potential impact, aiding in understanding the context and implications.
Top 5 News Categories in Cambodia - Politics: Covers major news on government decisions, political movements, elections, and policy changes affecting the national landscape. - Economy: Focuses on Cambodia’s economic indicators, trade activities, inflation rates, and corporate news impacting business and finance sectors. - Social Issues: Highlights news on public health, education, social protests, and other societal concerns driving public discourse. - Sports: Features events in popular sports, such as football and martial arts, attracting considerable attention and engagement. - Technology and Innovation: Reports on tech advancements, startups, and innovations in Cambodia’s evolving tech sector.
Top 5 News Sources in Cambodia - The Phnom Penh Post: One of Cambodia's leading English-language newspapers, offering comprehensive coverage of politics, economy, and social issues. - Cambodia Daily: A well-regarded source for news related to national affairs, business, and cultural events. - Fresh News: A prominent online news platform providing real-time updates on breaking news, sports, and entertainment. - Koh Santepheap Daily: A major Khmer-language newspaper known for its extensive reporting on current affairs and local issues. - VOD (Voice of Democracy): An independent news outlet focusing on in-depth coverage of politics, social issues, and investigative journalism.
Accessing Techsalerator’s News Events Data for Cambodia To access Techsalerator’s News Events Data for Cambodia, please contact info@techsalerator.com with your specific needs. We will provide a customized quote based on the data fields and records you require, with delivery available within 24 hours. Ongoing access options can also be discussed.
Included Data Fields - Event Date - Event Title - Source - Location - Event Description - Event Category (Politics, Economy, Sports, etc.) - Participants (if applicable) - Event Impact (Social, Economic, etc.)
Techsalerator’s dataset is a valuable tool for tracking significant events in Cambodia. It supports informed decision-making, whether for business strategy, market analysis, or academic research, offering a comprehensive view of the country’s news landscape.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Techsalerator's News Events Data for Indonesia: A Comprehensive Overview
Techsalerator's News Events Data for Indonesia offers a robust resource for businesses, researchers, and media organizations. This dataset aggregates information on significant news events throughout Indonesia, drawing from diverse media sources including news outlets, online publications, and social platforms. It provides crucial insights for those aiming to track trends, analyze public sentiment, or monitor industry-specific developments.
Key Data Fields
Event Date: Records the exact date of the news event. This is essential for analysts monitoring trends over time or businesses responding to market changes.
Event Title: A concise headline describing the event. This enables users to quickly categorize and evaluate news content based on relevance to their interests.
Source: Identifies the news outlet or platform where the event was reported. This helps users track reliable sources and gauge the event's reach and influence.
Location: Provides geographic details about where the event occurred within Indonesia. This is particularly useful for regional analysis or localized marketing efforts.
Event Description: A detailed summary of the event, including key developments, participants, and potential impacts. Researchers and businesses use this to understand the context and ramifications of the event.
Top 5 News Categories in Indonesia
Politics: Coverage on government decisions, political movements, elections, and policy changes affecting the national landscape.
Economy: Focuses on Indonesia’s economic indicators, inflation rates, international trade, and corporate activities influencing business and finance sectors.
Social Issues: News events addressing protests, public health, education, and other societal concerns that drive public discourse.
Sports: Highlights events in football, badminton, and other popular sports, often attracting widespread attention and engagement across the country.
Technology and Innovation: Reports on technological advancements, startups, and innovations within Indonesia’s growing tech ecosystem, featuring prominent companies and emerging trends.
Top 5 News Sources in Indonesia
Kompas: A leading Indonesian newspaper offering comprehensive coverage of politics, economy, and social issues.
Detik: A major online news platform providing real-time updates on breaking news, sports, and entertainment.
Tempo: A well-regarded source for in-depth reporting on national politics, social issues, and investigative journalism.
The Jakarta Post: An English-language newspaper delivering news related to business, politics, and cultural events across Indonesia.
CNN Indonesia: A prominent news network broadcasting updates on current affairs, sports, and live events throughout the country.
Accessing Techsalerator’s News Events Data for Indonesia
To access Techsalerator’s News Events Data for Indonesia, please contact info@techsalerator.com with your specific needs. We will provide a customized quote based on the data fields and records you require, with delivery available within 24 hours. Ongoing access options can also be discussed.
Included Data Fields
Techsalerator’s dataset is an invaluable tool for keeping track of significant events in Indonesia. It supports informed decision-making, whether for business strategy, market analysis, or academic research, providing a clear picture of the country’s news landscape.
Not seeing a result you expected?
Learn how you can add new datasets to our index.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
~This dataset contains 40 social media posts collected from multiple platforms (Twitter, Facebook, Instagram, YouTube, TikTok). It provides a detailed view of how different types of content perform, how users engage with them, and how moderation systems respond.
~**Platform & Content:** Includes post type (Tweet, Story, Video, etc.), unique IDs, and timestamps.
~**User Information:** Follower counts and verification status.
~**Content Metadata:** Text, category, language, country, length, media type, and presence of external links.
~**Engagement Metrics:** Like, share, and comment counts, along with an overall engagement score.
~Misinformation Flag
~Fact-Check Source
~Moderation Action (e.g., Approved, Warning Label, Demonetized, Removed)
~Sentiment Score (positive/negative tone)
~Toxicity Score (harassment/offensive likelihood)
~Political Leaning (Neutral, Liberal, Conservative, Conspiracy)
~Topic Tags (e.g., climate, vaccine, election, 5G)
~Virality Indicators: Viral score estimating likelihood of content going viral.
~**Fake News & Misinformation Research** – Train ML models to detect misinformation.
~**Content Moderation Systems** – Study how platforms label, remove, or demonetize harmful content.
~**NLP & Sentiment Analysis** – Analyze toxicity, bias, and sentiment across platforms.
~**Trend Analysis** – Compare engagement across topics (climate change, vaccines, elections, 5G).
~**Political Bias Detection** – Explore correlations between political leaning, engagement, and moderation.
~40 posts
~25 features
~This dataset is a synthetic but realistic representation of social media activity. It can be useful for machine learning, data analysis, and visualization projects related to misinformation, user engagement, and platform moderation.