100+ datasets found
  1. Twitter Tweets Sentiment Dataset

    • kaggle.com
    zip
    Updated Apr 8, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    M Yasser H (2022). Twitter Tweets Sentiment Dataset [Dataset]. https://www.kaggle.com/datasets/yasserh/twitter-tweets-sentiment-dataset
    Explore at:
    zip(1289519 bytes)Available download formats
    Dataset updated
    Apr 8, 2022
    Authors
    M Yasser H
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    https://raw.githubusercontent.com/Masterx-AI/Project_Twitter_Sentiment_Analysis_/main/twitt.jpg" alt="">

    Description:

    Twitter is an online Social Media Platform where people share their their though as tweets. It is observed that some people misuse it to tweet hateful content. Twitter is trying to tackle this problem and we shall help it by creating a strong NLP based-classifier model to distinguish the negative tweets & block such tweets. Can you build a strong classifier model to predict the same?

    Each row contains the text of a tweet and a sentiment label. In the training set you are provided with a word or phrase drawn from the tweet (selected_text) that encapsulates the provided sentiment.

    Make sure, when parsing the CSV, to remove the beginning / ending quotes from the text field, to ensure that you don't include them in your training.

    You're attempting to predict the word or phrase from the tweet that exemplifies the provided sentiment. The word or phrase should include all characters within that span (i.e. including commas, spaces, etc.)

    Columns:

    1. textID - unique ID for each piece of text
    2. text - the text of the tweet
    3. sentiment - the general sentiment of the tweet

    Acknowledgement:

    The dataset is download from Kaggle Competetions:
    https://www.kaggle.com/c/tweet-sentiment-extraction/data?select=train.csv

    Objective:

    • Understand the Dataset & cleanup (if required).
    • Build classification models to predict the twitter sentiments.
    • Compare the evaluation metrics of vaious classification algorithms.
  2. Indonesian Twitter Sentiment Analysis Dataset-PPKM

    • kaggle.com
    zip
    Updated Jul 31, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Angga Widiarta (2023). Indonesian Twitter Sentiment Analysis Dataset-PPKM [Dataset]. https://www.kaggle.com/datasets/anggapurnama/twitter-dataset-ppkm
    Explore at:
    zip(3757803 bytes)Available download formats
    Dataset updated
    Jul 31, 2023
    Authors
    Angga Widiarta
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    This dataset contains a collection of tweets from the Indonesian community, expressing their opinions on the government's implementation of PPKM (Enforcement of Community Activity Restrictions). The dataset consists of approximately 20,000 tweets gathered within the time range from April 1, 2020, to April 1, 2022.

    The selected time range for data collection is based on when Indonesia started implementing PPKM extensively and when the government revoked the policy. Within this dataset, diverse opinions, comments, and reactions from the public regarding the PPKM policy during that period can be found.

    This dataset provides an opportunity to analyze the sentiment and public views regarding the PPKM policy, as well as observe changes in opinions over time. It offers valuable insights into understanding the perceptions and reactions of the community towards government policies related to PPKM.

    Label: 0 (Positive), 1 (Neutral), 2 (Negative)

  3. Brand Sentiment Analysis Dataset (Twitter)

    • kaggle.com
    zip
    Updated Jan 7, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Tushar Paul (2024). Brand Sentiment Analysis Dataset (Twitter) [Dataset]. https://www.kaggle.com/datasets/tusharpaul2001/brand-sentiment-analysis-dataset
    Explore at:
    zip(375745 bytes)Available download formats
    Dataset updated
    Jan 7, 2024
    Authors
    Tushar Paul
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Dataset description Users assessed tweets related to various brands and products, providing evaluations on whether the sentiment conveyed was positive, negative, or neutral. Additionally, if the tweet conveyed any sentiment, contributors identified the specific brand or product targeted by that emotion.

    https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11965067%2Fa48606bfcaf80acebbb6edff7895484a%2Fdownload.png?generation=1704673111671747&alt=media" alt="">

    Train Dataset : 8589 rows x 3 columns https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11965067%2Fe998ba81ca461699a787ff7305486b24%2FTrainDS.JPG?generation=1704672608361793&alt=media" alt="">

    Test Dataset : 504 rows x 1 columns https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11965067%2F07df18965e91f84df123270aabb641e1%2Ftest.JPG?generation=1704679582009718&alt=media" alt="">

  4. Twitter Sentiment Analysis Datasets

    • brightdata.com
    .json, .csv, .xlsx
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Bright Data, Twitter Sentiment Analysis Datasets [Dataset]. https://brightdata.com/products/datasets/twitter/sentiment-analysis
    Explore at:
    .json, .csv, .xlsxAvailable download formats
    Dataset authored and provided by
    Bright Datahttps://brightdata.com/
    License

    https://brightdata.com/licensehttps://brightdata.com/license

    Area covered
    Worldwide
    Description

    Our Twitter Sentiment Analysis Dataset provides a comprehensive collection of tweets, enabling businesses, researchers, and analysts to assess public sentiment, track trends, and monitor brand perception in real time. This dataset includes detailed metadata for each tweet, allowing for in-depth analysis of user engagement, sentiment trends, and social media impact.

    Key Features:
    
      Tweet Content & Metadata: Includes tweet text, hashtags, mentions, media attachments, and engagement metrics such as likes, retweets, and replies.
      Sentiment Classification: Analyze sentiment polarity (positive, negative, neutral) to gauge public opinion on brands, events, and trending topics.
      Author & User Insights: Access user details such as username, profile information, follower count, and account verification status.
      Hashtag & Topic Tracking: Identify trending hashtags and keywords to monitor conversations and sentiment shifts over time.
      Engagement Metrics: Measure tweet performance based on likes, shares, and comments to evaluate audience interaction.
      Historical & Real-Time Data: Choose from historical datasets for trend analysis or real-time data for up-to-date sentiment tracking.
    
    
    Use Cases:
    
      Brand Monitoring & Reputation Management: Track public sentiment around brands, products, and services to manage reputation and customer perception.
      Market Research & Consumer Insights: Analyze consumer opinions on industry trends, competitor performance, and emerging market opportunities.
      Political & Social Sentiment Analysis: Evaluate public opinion on political events, social movements, and global issues.
      AI & Machine Learning Applications: Train sentiment analysis models for natural language processing (NLP) and predictive analytics.
      Advertising & Campaign Performance: Measure the effectiveness of marketing campaigns by analyzing audience engagement and sentiment.
    
    
    
      Our dataset is available in multiple formats (JSON, CSV, Excel) and can be delivered via API, cloud storage (AWS, Google Cloud, Azure), or direct download. 
      Gain valuable insights into social media sentiment and enhance your decision-making with high-quality, structured Twitter data.
    
  5. twitter-sentiment

    • huggingface.co
    Updated Jan 15, 2008
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    EleutherAI (2008). twitter-sentiment [Dataset]. https://huggingface.co/datasets/EleutherAI/twitter-sentiment
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jan 15, 2008
    Dataset authored and provided by
    EleutherAIhttps://eleuther.ai/
    Description

    EleutherAI/twitter-sentiment dataset hosted on Hugging Face and contributed by the HF Datasets community

  6. m

    Twitter US Airline Sentiment

    • metatext.io
    • kaggle.com
    csv
    Updated Aug 10, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Figure Eight (2026). Twitter US Airline Sentiment [Dataset]. https://metatext.io/datasets/twitter-us-airline-sentiment
    Explore at:
    csvAvailable download formats
    Dataset updated
    Aug 10, 2026
    Dataset authored and provided by
    Figure Eight
    Variables measured
    Classification, Sentiment Analysis
    Description

    Dataset contains airline-related tweets that were labeled with positive, negative, and neutral sentiment.

  7. Bitcoin Twitter Sentiment Dataset (2013–2023)

    • kaggle.com
    zip
    Updated May 18, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Andrea Penas Martinez (2025). Bitcoin Twitter Sentiment Dataset (2013–2023) [Dataset]. https://www.kaggle.com/datasets/andreapenasmartinez/bitcoin-twitter-sentiment-dataset-20132023
    Explore at:
    zip(15923455155 bytes)Available download formats
    Dataset updated
    May 18, 2025
    Authors
    Andrea Penas Martinez
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    This dataset contains over 26 million English-language tweets related to Bitcoin (BTC), collected between 2013 and 2023. The data was sourced from Kaggle and includes posts from a wide range of users, from everyday investors to high-profile figures. Each tweet includes metadata such as timestamp, user information, and text content. The dataset has been thoroughly cleaned to remove spam, non-English content, bot activity, and duplicated entries. It serves as the primary input for sentiment analysis and subsequent price prediction models in this study.

  8. h

    tweet_sentiment_multilingual

    • huggingface.co
    • opendatalab.com
    Updated Dec 25, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Cardiff NLP (2022). tweet_sentiment_multilingual [Dataset]. https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Dec 25, 2022
    Dataset authored and provided by
    Cardiff NLP
    Description

    Dataset Card for cardiffnlp/tweet_sentiment_multilingual

      Dataset Summary
    

    Tweet Sentiment Multilingual consists of sentiment analysis dataset on Twitter in 8 different lagnuages.

    arabic english french german hindi italian portuguese spanish

      Supported Tasks and Leaderboards
    

    text_classification: The dataset can be trained using a SentenceClassification model from HuggingFace transformers.

      Dataset Structure
    
    
    
    
    
      Data Instances
    

    An instance from… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_sentiment_multilingual.

  9. Twitter sentiment Data Set CSV file

    • kaggle.com
    zip
    Updated Jun 3, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    TuhinAI_Labs (2025). Twitter sentiment Data Set CSV file [Dataset]. https://www.kaggle.com/datasets/tuhinkundu2025/twitter-sentiment-data-set-csv-file
    Explore at:
    zip(2017433 bytes)Available download formats
    Dataset updated
    Jun 3, 2025
    Authors
    TuhinAI_Labs
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    Dataset

    This dataset was created by TuhinAI_Labs

    Released under Apache 2.0

    Contents

  10. h

    sentiment140

    • huggingface.co
    • opendatalab.com
    • +1more
    Updated Apr 23, 2010
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Stanford NLP (2010). sentiment140 [Dataset]. https://huggingface.co/datasets/stanfordnlp/sentiment140
    Explore at:
    Dataset updated
    Apr 23, 2010
    Dataset authored and provided by
    Stanford NLP
    Description

    Sentiment140 consists of Twitter messages with emoticons, which are used as noisy labels for sentiment classification. For more detailed information please refer to the paper.

  11. Twitter Sentiment Dataset

    • kaggle.com
    zip
    Updated May 14, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Saurabh Shahane (2021). Twitter Sentiment Dataset [Dataset]. https://www.kaggle.com/saurabhshahane/twitter-sentiment-dataset
    Explore at:
    zip(7966522 bytes)Available download formats
    Dataset updated
    May 14, 2021
    Authors
    Saurabh Shahane
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Context

    The dataset has three sentiments namely, negative(-1), neutral(0), and positive(+1). It contains two fields for the tweet and label.

    Acknowledgements

    HUSSEIN, SHERIF (2021), “Twitter Sentiments Dataset”, Mendeley Data, V1, doi: 10.17632/z9zw7nt5h2.1

  12. h

    synthetic-financial-tweets-sentiment

    • huggingface.co
    Updated Feb 5, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Tim Koornstra (2024). synthetic-financial-tweets-sentiment [Dataset]. https://huggingface.co/datasets/TimKoornstra/synthetic-financial-tweets-sentiment
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Feb 5, 2024
    Authors
    Tim Koornstra
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    FinTwitBERT: Synthetic Financial Tweets Dataset

      Description
    

    This dataset contains a collection of synthetically generated tweets related to financial markets, including discussions on stocks and cryptocurrencies. The tweets were generated using the NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO model, employing 10-shot random examples from the TimKoornstra/financial-tweets-sentiment dataset. Each entry in this dataset provides insights into financial discussions and is… See the full description on the dataset page: https://huggingface.co/datasets/TimKoornstra/synthetic-financial-tweets-sentiment.

  13. Twitter Sentiment Dataset 3 million labelled rows

    • kaggle.com
    zip
    Updated Jul 1, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Prakhar Awasthi (2023). Twitter Sentiment Dataset 3 million labelled rows [Dataset]. https://www.kaggle.com/datasets/prkhrawsthi/twitter-sentiment-dataset-3-million-labelled-rows/discussion
    Explore at:
    zip(108866071 bytes)Available download formats
    Dataset updated
    Jul 1, 2023
    Authors
    Prakhar Awasthi
    Description

    📢 Introducing the Twitter Sentiment Analysis Dataset 🐦📊

    Unlock the power of sentiment analysis with our comprehensive dataset from Twitter! 🌟📈 Analyzing entity-level sentiments, this dataset allows you to judge the sentiment of messages about specific entities. 📋✨

    With three distinct classes—Positive, Negative, and Neutral—you can delve into the sentiments expressed in tweets. We consider messages that are not relevant to the entity as Neutral, ensuring a comprehensive analysis. 🔄🔍

    Unleash the potential of this dataset for sentiment analysis tasks. Gain valuable insights into public opinions, brand reputation, and customer sentiments in real-time. 📈💬

    Join researchers, data scientists, and language enthusiasts as you explore the vast world of tweets. Develop and train sentiment analysis models to accurately classify the sentiments associated with various entities mentioned in the messages. 📚🔬

    Engage in conversations and share your findings within the community. Discuss the nuances of sentiment analysis, uncover trends, and refine your techniques together. 🗣️💭

    Note: The Twitter Sentiment Analysis Dataset is designed for research and analysis purposes only. The dataset categorizes sentiments into Positive, Negative, and Neutral classes. Let's embark on this exciting journey of sentiment analysis! 😊🐦✨

    { 0: Negative; 1: Positive; 2: Neutral }

  14. Tweets Dataset

    • brightdata.com
    .json, .csv, .xlsx
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Bright Data, Tweets Dataset [Dataset]. https://brightdata.com/products/datasets/twitter/tweets
    Explore at:
    .json, .csv, .xlsxAvailable download formats
    Dataset authored and provided by
    Bright Datahttps://brightdata.com/
    License

    https://brightdata.com/licensehttps://brightdata.com/license

    Area covered
    Worldwide
    Description

    Utilize our Tweets dataset for a range of applications to enhance business strategies and market insights. Analyzing this dataset offers a comprehensive view of social media dynamics, empowering organizations to optimize their communication and marketing strategies. Access the full dataset or select specific data points tailored to your needs. Popular use cases include sentiment analysis to gauge public opinion and brand perception, competitor analysis by examining engagement and sentiment around rival brands, and crisis management through real-time tracking of tweet sentiment and influential voices during critical events.

  15. m

    Extended Covid Twitter Datasets

    • data.mendeley.com
    Updated May 24, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Md Abrar Jahin (2023). Extended Covid Twitter Datasets [Dataset]. http://doi.org/10.17632/2ynwykrfgf.1
    Explore at:
    Dataset updated
    May 24, 2023
    Authors
    Md Abrar Jahin
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Wider spatiotemporal English COVID-19 Tweets

  16. m

    Annotated The dUCk Tweets Dataset

    • data.mendeley.com
    Updated Aug 13, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Hui Ming Tham (2021). Annotated The dUCk Tweets Dataset [Dataset]. http://doi.org/10.17632/876tc4dkts.2
    Explore at:
    Dataset updated
    Aug 13, 2021
    Authors
    Hui Ming Tham
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset is made up of unique annotated English-Malay code-switching, pure English, and pure Malay tweets using raw_tweets_012019_to_062020.csv on Kaggle (Carlson, 2020). The raw tweets file is the collected users’ tweets about a Malaysian brand called, ‘The dUCk Group’ which is founded by Vivy Yusof focuses on selling scarves, bags, cosmetics, stationaries, and Home & Living products. When preparing this dataset, the duplicated, invalid and unusable data rows are removed. The tweets are then annotated with the language category “ENG” for pure English tweets, “BM” for pure Malay tweets, and “ENG-BM” for the code-switching tweets. Besides, the tweets are annotated with sentiment value 0 for neutral, 1 for positive, and -1 for negative.

    The sub-folders contain in this dataset are as follows:

    1) Full Training Dataset: This sub-folder contains a full set of annotated pure English, pure Malay, and English-Malay code-switching tweets regarding ‘The dUCk Group’ brand, which can be used to train machine learning models. The tweets are kept in both CSV and XML format files namely 'full_training_dataset.csv' and 'full_training_dataset.xml'.

    2) Full Testing Dataset: This sub-folder contains a full set of annotated pure English, pure Malay, and English-Malay code-switching tweets regarding ‘The dUCk Group’ brand, which can be used to test the performance of learning models. The tweets are kept in both CSV and XML format files namely 'full_testing_dataset.csv' and 'full_testing_dataset.xml'.

    3) Code-Switching Training Dataset: This sub-folder comprises only annotated English-Malay code-switching tweets regarding ‘The dUCk Group’ brand for training the learning models. The tweets are kept in both CSV and XML format files namely 'eng_malay_training_dataset.csv' and 'eng_malay_training_dataset.xml'.

    4) Code-Switching Testing Dataset: This sub-folder comprises only annotated English-Malay code-switching tweets regarding ‘The dUCk Group’ brand, which can be used to evaluate the performance of the learning models. The tweets are kept in both CSV and XML format files namely 'eng_malay_testing_dataset.csv' and 'eng_malay_testing_dataset.xml.

    *Note: 'Language' column represents the language category of the tweet belongs to 'TweetText' column represents the whole tweet 'TweetSentiment' column represents the sentiment value of the tweet (0, 1, and -1)

  17. Twitter Sentiment Analysis Dataset

    • kaggle.com
    zip
    Updated Feb 13, 2021
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dr. Zohair Ahmed (2021). Twitter Sentiment Analysis Dataset [Dataset]. https://www.kaggle.com/datasets/zohairahmed007/twitter-sentiment-analysis-dataset
    Explore at:
    zip(38737743 bytes)Available download formats
    Dataset updated
    Feb 13, 2021
    Authors
    Dr. Zohair Ahmed
    Description

    Dataset

    This dataset was created by Dr. Zohair Ahmed

    Contents

  18. h

    rusya-ukrayna-twitter-sentiment-dataset

    • huggingface.co
    Updated Jun 25, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Köse (2026). rusya-ukrayna-twitter-sentiment-dataset [Dataset]. https://huggingface.co/datasets/Batuhanbey/rusya-ukrayna-twitter-sentiment-dataset
    Explore at:
    Dataset updated
    Jun 25, 2026
    Authors
    Köse
    Area covered
    Ukraine, Russia
    Description

    Rusya-Ukrayna Savaşı Twitter Duygu Analizi Veri Seti

    Bu veri seti, TÜBİTAK projemiz kapsamında Rusya-Ukrayna savaşıyla ilgili Twitter/X paylaşımlarının duygu analizi için hazırlanmıştır. Veriler açık kaynak veri setlerinden toplanmış, filtrelenmiş ve model eğitimi için düzenlenmiştir. Paylaşılan dosyada kullanıcı adı, kullanıcı id'si, profil bağlantısı gibi alanlar bulunmamaktadır.

      Dosyalar
    

    train.csv: Temizlenmiş metin ve duygu etiketi içeren ana veri dosyası.… See the full description on the dataset page: https://huggingface.co/datasets/Batuhanbey/rusya-ukrayna-twitter-sentiment-dataset.

  19. M

    Rawat - Sentiment Analysis of Tweets

    • datasetcatalog.nlm.nih.gov
    • data.mendeley.com
    Updated Dec 15, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Tamilselvan, Vedhavarshini (2023). Rawat - Sentiment Analysis of Tweets [Dataset]. https://datasetcatalog.nlm.nih.gov/dataset?q=0001268817
    Explore at:
    Dataset updated
    Dec 15, 2023
    Authors
    Tamilselvan, Vedhavarshini
    Description

    The extraction of data from the Twitter site was the initial step, without which no analysis is possible. Using the ‘Advanced Twitter search’ option, appropriate hashtags and dates were used to check if the Tweets from the desired dates are available. To scrape the Tweets and fetch the historical data, a Twitter framework was created in ‘Octoparse’ software. The final output of tweets was downloaded in Excel format. Nearly 234 tweets were obtained using the hashtags ‘#BipinRawat’ ‘#Karma’ ‘#IAFChoppercrash’ and ‘#IndianAirForce.’ Tweets in regional language; news and tweets of different contexts but similar hashtags; updates from online news channels; and retweets, or replies were filtered. To annotate the reviews manually the guidelines were framed following the design proposed by Mohammad (2016) in their manual, ‘A Practical Guide to Sentiment Annotation: Challenges and Solutions’. The questionnaire was also prepared based on the same.

  20. d

    EdChat Public Tweets

    • search.dataone.org
    Updated Dec 28, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Gruzd, Anatoliy; Conroy, Nadia (2023). EdChat Public Tweets [Dataset]. http://doi.org/10.5683/SP/JQ0NLC
    Explore at:
    Dataset updated
    Dec 28, 2023
    Dataset provided by
    Borealis
    Authors
    Gruzd, Anatoliy; Conroy, Nadia
    Description

    This is a data set of 482,251 public tweets and retweets (Twitter IDs) posted by the #edchat online community of educators who discuss current trends in teaching with technology. The data set was collected via Twitter's Streaming API between Feb 1, 2018 and Apr 4, 2018, and was used as part of the research on developing a learning analytics dashboard for teaching and learning with Twitter. Following Twitter's terms of service, the data set only includes unique identifiers of relevant tweets. To collect the actual tweets that are part of this data set, you will need to use one of the available third party tools such as Hydrator or Twarc ("hydrate" function). As part of this release, we are also attaching an enriched version of this data set that contains sentiment and opinion analysis labels that were produced by analyzing each tweet with the help of the NLTK SentimentAnalyzer Python package. *This work was supported in part by eCampusOntario and The Social Sciences and Humanities Research Council of Canada.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
M Yasser H (2022). Twitter Tweets Sentiment Dataset [Dataset]. https://www.kaggle.com/datasets/yasserh/twitter-tweets-sentiment-dataset
Organization logo

Twitter Tweets Sentiment Dataset

Twitter Tweets Sentiment Analysis for Natural Language Processing

Explore at:
43 scholarly articles cite this dataset (View in Google Scholar)
zip(1289519 bytes)Available download formats
Dataset updated
Apr 8, 2022
Authors
M Yasser H
License

https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

Description

https://raw.githubusercontent.com/Masterx-AI/Project_Twitter_Sentiment_Analysis_/main/twitt.jpg" alt="">

Description:

Twitter is an online Social Media Platform where people share their their though as tweets. It is observed that some people misuse it to tweet hateful content. Twitter is trying to tackle this problem and we shall help it by creating a strong NLP based-classifier model to distinguish the negative tweets & block such tweets. Can you build a strong classifier model to predict the same?

Each row contains the text of a tweet and a sentiment label. In the training set you are provided with a word or phrase drawn from the tweet (selected_text) that encapsulates the provided sentiment.

Make sure, when parsing the CSV, to remove the beginning / ending quotes from the text field, to ensure that you don't include them in your training.

You're attempting to predict the word or phrase from the tweet that exemplifies the provided sentiment. The word or phrase should include all characters within that span (i.e. including commas, spaces, etc.)

Columns:

  1. textID - unique ID for each piece of text
  2. text - the text of the tweet
  3. sentiment - the general sentiment of the tweet

Acknowledgement:

The dataset is download from Kaggle Competetions:
https://www.kaggle.com/c/tweet-sentiment-extraction/data?select=train.csv

Objective:

  • Understand the Dataset & cleanup (if required).
  • Build classification models to predict the twitter sentiments.
  • Compare the evaluation metrics of vaious classification algorithms.
Search
Clear search
Close search
Google apps
Main menu