Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
https://raw.githubusercontent.com/Masterx-AI/Project_Twitter_Sentiment_Analysis_/main/twitt.jpg" alt="">
Twitter is an online Social Media Platform where people share their their though as tweets. It is observed that some people misuse it to tweet hateful content. Twitter is trying to tackle this problem and we shall help it by creating a strong NLP based-classifier model to distinguish the negative tweets & block such tweets. Can you build a strong classifier model to predict the same?
Each row contains the text of a tweet and a sentiment label. In the training set you are provided with a word or phrase drawn from the tweet (selected_text) that encapsulates the provided sentiment.
Make sure, when parsing the CSV, to remove the beginning / ending quotes from the text field, to ensure that you don't include them in your training.
You're attempting to predict the word or phrase from the tweet that exemplifies the provided sentiment. The word or phrase should include all characters within that span (i.e. including commas, spaces, etc.)
The dataset is download from Kaggle Competetions:
https://www.kaggle.com/c/tweet-sentiment-extraction/data?select=train.csv
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains a collection of tweets from the Indonesian community, expressing their opinions on the government's implementation of PPKM (Enforcement of Community Activity Restrictions). The dataset consists of approximately 20,000 tweets gathered within the time range from April 1, 2020, to April 1, 2022.
The selected time range for data collection is based on when Indonesia started implementing PPKM extensively and when the government revoked the policy. Within this dataset, diverse opinions, comments, and reactions from the public regarding the PPKM policy during that period can be found.
This dataset provides an opportunity to analyze the sentiment and public views regarding the PPKM policy, as well as observe changes in opinions over time. It offers valuable insights into understanding the perceptions and reactions of the community towards government policies related to PPKM.
Label: 0 (Positive), 1 (Neutral), 2 (Negative)
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Dataset description Users assessed tweets related to various brands and products, providing evaluations on whether the sentiment conveyed was positive, negative, or neutral. Additionally, if the tweet conveyed any sentiment, contributors identified the specific brand or product targeted by that emotion.
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11965067%2Fa48606bfcaf80acebbb6edff7895484a%2Fdownload.png?generation=1704673111671747&alt=media" alt="">
Train Dataset : 8589 rows x 3 columns
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11965067%2Fe998ba81ca461699a787ff7305486b24%2FTrainDS.JPG?generation=1704672608361793&alt=media" alt="">
Test Dataset : 504 rows x 1 columns
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11965067%2F07df18965e91f84df123270aabb641e1%2Ftest.JPG?generation=1704679582009718&alt=media" alt="">
Facebook
Twitterhttps://brightdata.com/licensehttps://brightdata.com/license
Our Twitter Sentiment Analysis Dataset provides a comprehensive collection of tweets, enabling businesses, researchers, and analysts to assess public sentiment, track trends, and monitor brand perception in real time. This dataset includes detailed metadata for each tweet, allowing for in-depth analysis of user engagement, sentiment trends, and social media impact.
Key Features:
Tweet Content & Metadata: Includes tweet text, hashtags, mentions, media attachments, and engagement metrics such as likes, retweets, and replies.
Sentiment Classification: Analyze sentiment polarity (positive, negative, neutral) to gauge public opinion on brands, events, and trending topics.
Author & User Insights: Access user details such as username, profile information, follower count, and account verification status.
Hashtag & Topic Tracking: Identify trending hashtags and keywords to monitor conversations and sentiment shifts over time.
Engagement Metrics: Measure tweet performance based on likes, shares, and comments to evaluate audience interaction.
Historical & Real-Time Data: Choose from historical datasets for trend analysis or real-time data for up-to-date sentiment tracking.
Use Cases:
Brand Monitoring & Reputation Management: Track public sentiment around brands, products, and services to manage reputation and customer perception.
Market Research & Consumer Insights: Analyze consumer opinions on industry trends, competitor performance, and emerging market opportunities.
Political & Social Sentiment Analysis: Evaluate public opinion on political events, social movements, and global issues.
AI & Machine Learning Applications: Train sentiment analysis models for natural language processing (NLP) and predictive analytics.
Advertising & Campaign Performance: Measure the effectiveness of marketing campaigns by analyzing audience engagement and sentiment.
Our dataset is available in multiple formats (JSON, CSV, Excel) and can be delivered via API, cloud storage (AWS, Google Cloud, Azure), or direct download.
Gain valuable insights into social media sentiment and enhance your decision-making with high-quality, structured Twitter data.
Facebook
TwitterEleutherAI/twitter-sentiment dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterDataset contains airline-related tweets that were labeled with positive, negative, and neutral sentiment.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains over 26 million English-language tweets related to Bitcoin (BTC), collected between 2013 and 2023. The data was sourced from Kaggle and includes posts from a wide range of users, from everyday investors to high-profile figures. Each tweet includes metadata such as timestamp, user information, and text content. The dataset has been thoroughly cleaned to remove spam, non-English content, bot activity, and duplicated entries. It serves as the primary input for sentiment analysis and subsequent price prediction models in this study.
Facebook
TwitterDataset Card for cardiffnlp/tweet_sentiment_multilingual
Dataset Summary
Tweet Sentiment Multilingual consists of sentiment analysis dataset on Twitter in 8 different lagnuages.
arabic english french german hindi italian portuguese spanish
Supported Tasks and Leaderboards
text_classification: The dataset can be trained using a SentenceClassification model from HuggingFace transformers.
Dataset Structure
Data Instances
An instance from… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset was created by TuhinAI_Labs
Released under Apache 2.0
Facebook
TwitterSentiment140 consists of Twitter messages with emoticons, which are used as noisy labels for sentiment classification. For more detailed information please refer to the paper.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The dataset has three sentiments namely, negative(-1), neutral(0), and positive(+1). It contains two fields for the tweet and label.
HUSSEIN, SHERIF (2021), “Twitter Sentiments Dataset”, Mendeley Data, V1, doi: 10.17632/z9zw7nt5h2.1
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
FinTwitBERT: Synthetic Financial Tweets Dataset
Description
This dataset contains a collection of synthetically generated tweets related to financial markets, including discussions on stocks and cryptocurrencies. The tweets were generated using the NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO model, employing 10-shot random examples from the TimKoornstra/financial-tweets-sentiment dataset. Each entry in this dataset provides insights into financial discussions and is… See the full description on the dataset page: https://huggingface.co/datasets/TimKoornstra/synthetic-financial-tweets-sentiment.
Facebook
Twitter📢 Introducing the Twitter Sentiment Analysis Dataset 🐦📊
Unlock the power of sentiment analysis with our comprehensive dataset from Twitter! 🌟📈 Analyzing entity-level sentiments, this dataset allows you to judge the sentiment of messages about specific entities. 📋✨
With three distinct classes—Positive, Negative, and Neutral—you can delve into the sentiments expressed in tweets. We consider messages that are not relevant to the entity as Neutral, ensuring a comprehensive analysis. 🔄🔍
Unleash the potential of this dataset for sentiment analysis tasks. Gain valuable insights into public opinions, brand reputation, and customer sentiments in real-time. 📈💬
Join researchers, data scientists, and language enthusiasts as you explore the vast world of tweets. Develop and train sentiment analysis models to accurately classify the sentiments associated with various entities mentioned in the messages. 📚🔬
Engage in conversations and share your findings within the community. Discuss the nuances of sentiment analysis, uncover trends, and refine your techniques together. 🗣️💭
Note: The Twitter Sentiment Analysis Dataset is designed for research and analysis purposes only. The dataset categorizes sentiments into Positive, Negative, and Neutral classes. Let's embark on this exciting journey of sentiment analysis! 😊🐦✨
{ 0: Negative; 1: Positive; 2: Neutral }
Facebook
Twitterhttps://brightdata.com/licensehttps://brightdata.com/license
Utilize our Tweets dataset for a range of applications to enhance business strategies and market insights. Analyzing this dataset offers a comprehensive view of social media dynamics, empowering organizations to optimize their communication and marketing strategies. Access the full dataset or select specific data points tailored to your needs. Popular use cases include sentiment analysis to gauge public opinion and brand perception, competitor analysis by examining engagement and sentiment around rival brands, and crisis management through real-time tracking of tweet sentiment and influential voices during critical events.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Wider spatiotemporal English COVID-19 Tweets
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset is made up of unique annotated English-Malay code-switching, pure English, and pure Malay tweets using raw_tweets_012019_to_062020.csv on Kaggle (Carlson, 2020). The raw tweets file is the collected users’ tweets about a Malaysian brand called, ‘The dUCk Group’ which is founded by Vivy Yusof focuses on selling scarves, bags, cosmetics, stationaries, and Home & Living products. When preparing this dataset, the duplicated, invalid and unusable data rows are removed. The tweets are then annotated with the language category “ENG” for pure English tweets, “BM” for pure Malay tweets, and “ENG-BM” for the code-switching tweets. Besides, the tweets are annotated with sentiment value 0 for neutral, 1 for positive, and -1 for negative.
The sub-folders contain in this dataset are as follows:
1) Full Training Dataset: This sub-folder contains a full set of annotated pure English, pure Malay, and English-Malay code-switching tweets regarding ‘The dUCk Group’ brand, which can be used to train machine learning models. The tweets are kept in both CSV and XML format files namely 'full_training_dataset.csv' and 'full_training_dataset.xml'.
2) Full Testing Dataset: This sub-folder contains a full set of annotated pure English, pure Malay, and English-Malay code-switching tweets regarding ‘The dUCk Group’ brand, which can be used to test the performance of learning models. The tweets are kept in both CSV and XML format files namely 'full_testing_dataset.csv' and 'full_testing_dataset.xml'.
3) Code-Switching Training Dataset: This sub-folder comprises only annotated English-Malay code-switching tweets regarding ‘The dUCk Group’ brand for training the learning models. The tweets are kept in both CSV and XML format files namely 'eng_malay_training_dataset.csv' and 'eng_malay_training_dataset.xml'.
4) Code-Switching Testing Dataset: This sub-folder comprises only annotated English-Malay code-switching tweets regarding ‘The dUCk Group’ brand, which can be used to evaluate the performance of the learning models. The tweets are kept in both CSV and XML format files namely 'eng_malay_testing_dataset.csv' and 'eng_malay_testing_dataset.xml.
*Note: 'Language' column represents the language category of the tweet belongs to 'TweetText' column represents the whole tweet 'TweetSentiment' column represents the sentiment value of the tweet (0, 1, and -1)
Facebook
TwitterThis dataset was created by Dr. Zohair Ahmed
Facebook
TwitterRusya-Ukrayna Savaşı Twitter Duygu Analizi Veri Seti
Bu veri seti, TÜBİTAK projemiz kapsamında Rusya-Ukrayna savaşıyla ilgili Twitter/X paylaşımlarının duygu analizi için hazırlanmıştır. Veriler açık kaynak veri setlerinden toplanmış, filtrelenmiş ve model eğitimi için düzenlenmiştir. Paylaşılan dosyada kullanıcı adı, kullanıcı id'si, profil bağlantısı gibi alanlar bulunmamaktadır.
Dosyalar
train.csv: Temizlenmiş metin ve duygu etiketi içeren ana veri dosyası.… See the full description on the dataset page: https://huggingface.co/datasets/Batuhanbey/rusya-ukrayna-twitter-sentiment-dataset.
Facebook
TwitterThe extraction of data from the Twitter site was the initial step, without which no analysis is possible. Using the ‘Advanced Twitter search’ option, appropriate hashtags and dates were used to check if the Tweets from the desired dates are available. To scrape the Tweets and fetch the historical data, a Twitter framework was created in ‘Octoparse’ software. The final output of tweets was downloaded in Excel format. Nearly 234 tweets were obtained using the hashtags ‘#BipinRawat’ ‘#Karma’ ‘#IAFChoppercrash’ and ‘#IndianAirForce.’ Tweets in regional language; news and tweets of different contexts but similar hashtags; updates from online news channels; and retweets, or replies were filtered. To annotate the reviews manually the guidelines were framed following the design proposed by Mohammad (2016) in their manual, ‘A Practical Guide to Sentiment Annotation: Challenges and Solutions’. The questionnaire was also prepared based on the same.
Facebook
TwitterThis is a data set of 482,251 public tweets and retweets (Twitter IDs) posted by the #edchat online community of educators who discuss current trends in teaching with technology. The data set was collected via Twitter's Streaming API between Feb 1, 2018 and Apr 4, 2018, and was used as part of the research on developing a learning analytics dashboard for teaching and learning with Twitter. Following Twitter's terms of service, the data set only includes unique identifiers of relevant tweets. To collect the actual tweets that are part of this data set, you will need to use one of the available third party tools such as Hydrator or Twarc ("hydrate" function). As part of this release, we are also attaching an enriched version of this data set that contains sentiment and opinion analysis labels that were produced by analyzing each tweet with the help of the NLTK SentimentAnalyzer Python package. *This work was supported in part by eCampusOntario and The Social Sciences and Humanities Research Council of Canada.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
https://raw.githubusercontent.com/Masterx-AI/Project_Twitter_Sentiment_Analysis_/main/twitt.jpg" alt="">
Twitter is an online Social Media Platform where people share their their though as tweets. It is observed that some people misuse it to tweet hateful content. Twitter is trying to tackle this problem and we shall help it by creating a strong NLP based-classifier model to distinguish the negative tweets & block such tweets. Can you build a strong classifier model to predict the same?
Each row contains the text of a tweet and a sentiment label. In the training set you are provided with a word or phrase drawn from the tweet (selected_text) that encapsulates the provided sentiment.
Make sure, when parsing the CSV, to remove the beginning / ending quotes from the text field, to ensure that you don't include them in your training.
You're attempting to predict the word or phrase from the tweet that exemplifies the provided sentiment. The word or phrase should include all characters within that span (i.e. including commas, spaces, etc.)
The dataset is download from Kaggle Competetions:
https://www.kaggle.com/c/tweet-sentiment-extraction/data?select=train.csv