48 datasets found
  1. Customer Support on Twitter

    • berd-platform.de
    csv
    Updated Jul 31, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Stuart Axelbrooke; Stuart Axelbrooke (2025). Customer Support on Twitter [Dataset]. http://doi.org/10.34740/kaggle/dsv/8841
    Explore at:
    csvAvailable download formats
    Dataset updated
    Jul 31, 2025
    Dataset provided by
    Kagglehttp://kaggle.com/
    Authors
    Stuart Axelbrooke; Stuart Axelbrooke
    Time period covered
    Mar 12, 2017
    Description

    The Customer Support on Twitter dataset is a large, modern corpus of tweets and replies to aid innovation in natural language understanding and conversational models, and for study of modern customer support practices and impact. The dataset includes replies of companies like Apple, Amazon, Uber, Delta, Spotify and others.

  2. Twitter Customer Service Interaction Summarization

    • kaggle.com
    zip
    Updated May 14, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rock-Lagoon (2026). Twitter Customer Service Interaction Summarization [Dataset]. https://www.kaggle.com/datasets/rocklagoon/twitter-customer-service-interaction-summarization
    Explore at:
    zip(253517940 bytes)Available download formats
    Dataset updated
    May 14, 2026
    Authors
    Rock-Lagoon
    Description

    Dataset

    This dataset was created by Rock-Lagoon

    Contents

  3. Customer Support on Twitter

    • kaggle.com
    zip
    Updated Oct 17, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Amin Aslami (2024). Customer Support on Twitter [Dataset]. https://www.kaggle.com/datasets/aminaslam/customer-support-on-twitter
    Explore at:
    zip(78948 bytes)Available download formats
    Dataset updated
    Oct 17, 2024
    Authors
    Amin Aslami
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    Dataset

    This dataset was created by Amin Aslami

    Released under Apache 2.0

    Contents

  4. Customer Support Twitter Data

    • kaggle.com
    zip
    Updated Aug 29, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Muhammad Asif (2025). Customer Support Twitter Data [Dataset]. https://www.kaggle.com/datasets/muhammadasif786/customer-support-twitter-data
    Explore at:
    zip(176765850 bytes)Available download formats
    Dataset updated
    Aug 29, 2025
    Authors
    Muhammad Asif
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    Dataset

    This dataset was created by Muhammad Asif

    Released under Apache 2.0

    Contents

  5. Twitter customer support twitter llm finetune

    • kaggle.com
    zip
    Updated Sep 1, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Muhammad Asif (2025). Twitter customer support twitter llm finetune [Dataset]. https://www.kaggle.com/datasets/muhammadasif786/twitter-customer-support-twitter-llm-finetune/suggestions
    Explore at:
    zip(176765850 bytes)Available download formats
    Dataset updated
    Sep 1, 2025
    Authors
    Muhammad Asif
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    Dataset

    This dataset was created by Muhammad Asif

    Released under Apache 2.0

    Contents

  6. Customer Support Tweets (945M rows)

    • kaggle.com
    zip
    Updated Oct 31, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Galal Qassas (2025). Customer Support Tweets (945M rows) [Dataset]. https://www.kaggle.com/datasets/galalqassas/customer-support-tweets-945m-rows
    Explore at:
    zip(74154613 bytes)Available download formats
    Dataset updated
    Oct 31, 2025
    Authors
    Galal Qassas
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Dataset

    This dataset was created by Galal Qassas

    Released under MIT

    Contents

  7. customer care tweets KSA

    • kaggle.com
    zip
    Updated May 20, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Mansour (2022). customer care tweets KSA [Dataset]. https://www.kaggle.com/mansourhussain/customer-care-tweets-ksa
    Explore at:
    zip(642212 bytes)Available download formats
    Dataset updated
    May 20, 2022
    Authors
    Mansour
    Area covered
    Saudi Arabia
    Description

    - this data contains 10000 tweets for a telecom company's customer care account on Twitter.

    - this data need to use in Sentiment Analysis in Arabic.

  8. Twitter Airline Sentiment Dataset

    • kaggle.com
    zip
    Updated Nov 14, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Chandana Ramakrishna (2025). Twitter Airline Sentiment Dataset [Dataset]. https://www.kaggle.com/datasets/chandana890/twitter-airline-sentiment-dataset
    Explore at:
    zip(1134990 bytes)Available download formats
    Dataset updated
    Nov 14, 2025
    Authors
    Chandana Ramakrishna
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    Overview

    This dataset contains tweets related to major US airlines and is widely used for NLP and sentiment analysis tasks. Each record includes the tweet text, timestamp, airline name, and sentiment label (positive, negative, neutral). This uploaded version is prepared to support advanced text processing, machine learning, and anomaly detection experiments.

    What's Included

    • Tweets.csv – Full collection of airline-related tweets
    • Text content suitable for NLP tasks
    • Timestamp information (useful for time-based analysis)
    • Sentiment labels for classification and evaluation
    • Cleaned text field for direct use in ML pipelines

    Purpose of This Dataset

    This dataset is used in a machine learning workflow focused on: - sentiment analysis
    - embedding generation (transformers)
    - dimensionality reduction (PCA, UMAP)
    - clustering and visualization
    - unsupervised anomaly detection using Isolation Forest

    It is especially suited for exploring changes in public sentiment, event detection, and contextual analysis in social media data.

    Key Use Cases

    • Building and testing NLP models
    • Semantic similarity and embedding-based analysis
    • Sentiment classification
    • Detecting anomalous posts or time periods
    • Visualizing tweet clusters using UMAP
    • Studying customer feedback patterns in the airline industry

    Source

    Originally derived from the Twitter US Airline Sentiment dataset on Kaggle.
    This uploaded version is intended for educational, analytical, and research purposes.

    Notes

    If you're using this dataset in a notebook, ensure you update your file path accordingly: ```python df = pd.read_csv("/kaggle/input/twitter-airline-sentiment-dataset/Tweets.csv")

  9. Twitter Customer Reviews of Popular Smart Phone

    • kaggle.com
    zip
    Updated Jun 8, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Shibbir Ahmed Arif (2024). Twitter Customer Reviews of Popular Smart Phone [Dataset]. https://www.kaggle.com/datasets/shibbir282/twitter-customer-reviews-of-popular-smart-phone
    Explore at:
    zip(1236373 bytes)Available download formats
    Dataset updated
    Jun 8, 2024
    Authors
    Shibbir Ahmed Arif
    Description

    Context

    This dataset is a part of our research work titled "Opinion Mining of Customer Reviews Using Supervised Learning Algorithms". If you use this dataset then please cite our work. You can find the article in https://ieeexplore.ieee.org/document/9733435

    Content

    Nowadays, a lot of people express their opinions on various topics using social networking sites. Twitter has become a famous social networking site where people can express their opinions to the point and so it has become a great source for opinion mining. In this research, the goal was to train and build a model that can automatically and accurately categorize the opinion of customer tweet reviews about popular cell phone brands. We have used python TextBlob library for getting the polarity values of all the tweet reviews of the dataset. We have also used Support Vector Machine (SVM), Naïve Bayes, Logistic Regression, Decision Tree and Random Forest algorithms along with Bag of Words and TF-IDF vectorizers separately to train and build the model. We have investigated the opinions using five classes which are Strongly Positive, Positive, Neutral, Negative and Strongly Negative.

    When referencing this dataset please cite the below paper

    Bibtex @inproceedings{arif2021opinion, title={Opinion Mining of Customer Reviews Using Supervised Learning Algorithms}, author={Arif, Shibbir Ahmed and Hossain, Taslima Binte}, booktitle={2021 5th International Conference on Electrical Information and Communication Technology (EICT)}, pages={1--6}, year={2021}, organization={IEEE} }

  10. Tweets and User Engagement

    • kaggle.com
    zip
    Updated Dec 6, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    The Devastator (2023). Tweets and User Engagement [Dataset]. https://www.kaggle.com/datasets/thedevastator/tweets-and-user-engagement
    Explore at:
    zip(9121838 bytes)Available download formats
    Dataset updated
    Dec 6, 2023
    Authors
    The Devastator
    Description

    Tweets and User Engagement

    Twitter Data: Tweet Characteristics and Engagement Metrics

    By Krystal Jensen [source]

    About this dataset

    The dataset Twitter Data: Tweets and User Interactions provides comprehensive information about tweets and user interactions on the popular social media platform Twitter. The dataset includes various attributes that shed light on the characteristics and engagement metrics of tweets, allowing for in-depth analysis of user behavior and content performance.

    One of the key variables in this dataset is the Klout score, which represents the influence and reputation of the Twitter users who posted the tweets. This numeric metric helps assess the impact a user has on their audience and provides insights into their social media presence.

    Another essential attribute is the text content of each tweet. By examining this textual data, analysts can uncover valuable information about trending topics, opinions, sentiments, conversations, or news shared by users. It serves as a primary source for understanding what people share publicly on Twitter.

    The dataset Twitter+data+in+sheets.csv serves as a reliable resource for conducting research or performing analytics that require detailed information about Twitter activity. It covers aspects such as tweet characteristics (including length and language), engagement metrics (such as retweets and favorites), sentiment analysis (revealing positive or negative emotions expressed), as well as individual user details.

    By utilizing this extensive dataset, researchers can gain valuable insights into patterns of online communication within Twitter's vast network. They can identify influential individuals with high Klout scores who have substantial reach among their followers or communities. Additionally, they can analyze various aspects related to tweet content such as sentiment analysis to understand public opinion trends or measure engagement levels through counts like retweets and favorites.

    Overall, this dataset serves as an invaluable resource for anyone interested in comprehensively analyzing tweets' characteristics, exploring how users interact with them across different dimensions like popularity or sentiment analysis groups—or examining correlations between Klout scores with other factors influencing engagement levels like time posted

    How to use the dataset

    Welcome to the Twitter Data: Tweets and User Interactions dataset! This dataset provides valuable insights into tweet characteristics and user engagement on Twitter. Here is a useful guide on how to make the most out of this dataset:

    • Understanding the Columns: There are two main columns in this dataset:

      • Klout Score (Numeric): The Klout score indicates the influence of the user who posted the tweet. A higher Klout score suggests greater influence and reach.
      • Text Content of Tweet (Text): This column contains the actual text content of each tweet.
    • Analyzing Tweet Characteristics: The text content column will help you understand various aspects of tweets, such as language, sentiment, trending topics, or specific keywords used by users. You can perform text analysis techniques like word frequency analysis or sentiment analysis to gain insights into tweet characteristics.

    • Examining User Engagement: The Klout score provides a measure of user influence on Twitter. By analyzing this column, you can identify highly influential users who generate higher engagement rates with their tweets. You can further explore interactions (likes, retweets, replies) between these influential users and other Twitter users mentioned in their tweets.

    • Identifying Trends and Patterns: With this dataset's rich information about tweet content and user engagement, you can identify popular trends or patterns among highly engaged tweets or influential users over different time periods.

    Remember that dates are not included in this guide since they were not provided in the original request for creating it.

    Please note that it is essential to responsibly use this data for any analysis or research purposes while adhering to ethical considerations related to privacy rights and data usage policies set by both Kaggle platform rules as well as any relevant privacy regulations.

    Best regards, [Your Name]

    Research Ideas

    • Analyzing the relationship between Klout score and the content of tweets: This dataset can be used to investigate whether there is a correlation between a user's Klout score (a measure of their social media influence) and the characteristics of their tweets. By examining factors such as tweet length, sentiment, and engagement metrics, researchers can gain...
  11. Sentiment with 1.6 million tweets with locations

    • kaggle.com
    zip
    Updated Mar 12, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    vivek chary (2023). Sentiment with 1.6 million tweets with locations [Dataset]. https://www.kaggle.com/datasets/vivekchary/sentiment-with-16-million-tweets-with-locations
    Explore at:
    zip(86959692 bytes)Available download formats
    Dataset updated
    Mar 12, 2023
    Authors
    vivek chary
    Description

    The "Sentiment with 16 million tweets with locations" dataset is a collection of tweets with their respective geographical location information and sentiment labels. The dataset includes 16 million tweets from various locations around the world, spanning a period of several years. The sentiment labels for each tweet are binary, indicating whether the sentiment expressed in the tweet is positive or negative.

    This dataset can be used for sentiment analysis and natural language processing tasks, such as training machine learning models to classify the sentiment of text data. Researchers and developers can use this dataset to analyze trends in sentiment across different locations and time periods, as well as to develop new algorithms and models for sentiment analysis.

    Please note that this dataset is intended for research purposes only and should not be used for any commercial or legal applications. The dataset may also contain offensive or inappropriate language, and users should exercise caution when working with this data

    Context In addition to the technical details of the "Sentiment with 16 million tweets with locations" dataset, some context that may be relevant to include in the About Dataset section could be:

    • The dataset was compiled and made publicly available by Vivek Chary, a data scientist and machine learning engineer.
    • The tweets were collected using the Twitter API, and the dataset was last updated in 2017.
    • The dataset includes tweets in various languages, although the majority are in English.
    • Sentiment analysis is a common application of natural language processing, and has a wide range of potential use cases, such as in market research, social media monitoring, and customer service.
    • Sentiment analysis can be challenging due to the complexity and ambiguity of language, as well as the variability of individual expression and context.

    • Large datasets like this one are important for developing accurate and robust sentiment analysis models, as they provide a diverse and representative sample of real-world text data.

    Content It contains the following 7 fields:

    1. Sentiment Target: The polarity of the tweet, indicated by a numeric value of 0 (negative), 2 (neutral), or 4 (positive).

    2. Tweet ID: The unique identifier of the tweet.

    3. Date: The date and time the tweet was posted in Coordinated Universal Time (UTC) format.

    4. Query Flag: The keyword or phrase used to filter the tweets. If no query was used, the value is NO_QUERY.

    5. User: The username of the Twitter account that posted the tweet.

    6. Text: The actual text content of the tweet.

    7. Location: The location of the tweet

  12. Saudi Customer Care Tweets

    • kaggle.com
    zip
    Updated Mar 13, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Abdullah Alsharif (2024). Saudi Customer Care Tweets [Dataset]. https://www.kaggle.com/alshreefabdullh/saudi-customer-care-tweets
    Explore at:
    zip(10030314 bytes)Available download formats
    Dataset updated
    Mar 13, 2024
    Authors
    Abdullah Alsharif
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Area covered
    Saudi Arabia
    Description

    This data was collected from several customer care accounts as inquiries of the customers.

    "fullText": This variable contains the full-text content of the tweet. "lang": This variable indicates the language in which the tweet is written. "viewsCount": This variable represents the count of views or impressions the tweet has received. "bookmarkCount": This variable represents the count of times the tweet has been bookmarked by users. "favoriteCount": This variable represents the count of times the tweet has been favorited by users. "replyCount": This variable represents the count of replies the tweet has received. "retweetCount": This variable represents the count of times the tweet has been retweeted by users. "quoteCount": This variable represents the count of times the tweet has been quoted by users.

  13. Twitter dataset of Facebook/Meta

    • kaggle.com
    zip
    Updated Apr 10, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Haber (2022). Twitter dataset of Facebook/Meta [Dataset]. https://www.kaggle.com/datasets/haber0322/twitter-dataset-of-facebook
    Explore at:
    zip(1113136 bytes)Available download formats
    Dataset updated
    Apr 10, 2022
    Authors
    Haber
    License

    Attribution 3.0 (CC BY 3.0)https://creativecommons.org/licenses/by/3.0/
    License information was derived automatically

    Description

    Public dataset that everyone can use Creating a dataframe from the tweets list above. E-mail supervision In order to keep a regular discussion going, it is useful to use e-mail. There are many distance discussions that take place by e-mail or fax, backed up with some visits for full-blown supervisions. If the supervisor is in another country, then e-mail contact is essential, as the face-to-face supervisory contacts will be condensed into the periods when you can both be in the same country. Make e-mail contacts lucid, short and precise, with some friendly tone to establish a personal touch. Try not to get involved in excessively chatty discussions but concentrate on asking questions, seeking information and reporting on findings for comment. E-mail is quite an insistent medium. If you make contact too frequently, the supervisor will feel harassed. If you make contact too infrequently, the supervisor will feel guilty (and so will you), wondering what you are up to. Regular brief contact with some very full discussions on work in progress at regular intervals will maintain a sense of a working relationship over time and space.

  14. (🌇Sunset) 🇺🇦 Ukraine Conflict Twitter Dataset

    • kaggle.com
    zip
    Updated Apr 2, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    BwandoWando (2024). (🌇Sunset) 🇺🇦 Ukraine Conflict Twitter Dataset [Dataset]. https://www.kaggle.com/datasets/bwandowando/ukraine-russian-crisis-twitter-dataset-1-2-m-rows
    Explore at:
    zip(18174367560 bytes)Available download formats
    Dataset updated
    Apr 2, 2024
    Authors
    BwandoWando
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Area covered
    Ukraine
    Description

    IMPORTANT (02-Apr-2024)

    Kaggle has fixed the issue with gzip files and Version 510 should now reflect properly working files

    IMPORTANT (28-Mar-2024)

    Please use the version 508 of the dataset, as 509 is broken. See link below of the dataset that is properly working https://www.kaggle.com/datasets/bwandowando/ukraine-russian-crisis-twitter-dataset-1-2-m-rows/versions/508

    Context

    The context and history of the current ongoing conflict can be found https://en.wikipedia.org/wiki/2022_Russian_invasion_of_Ukraine.

    Announcement

    [Jun 16] (🌇Sunset) Twitter has finally pulled the plug on all of my remaining TWITTER API accounts as part of their efforts for developers to migrate to the new API. The last tweets that I pulled was dated last Jun 14, and no more data from Jun 15 onwards. It was fun til it lasted and I hope that this dataset was able and will continue to help a lot. I'll just leave the dataset here for future download and reference. Thank you all!

    [Apr 19] Two additional developer accounts have been permanently suspended, expect a lower throughtput in the next few weeks. I will pull data til they ban my last account.

    [Apr 08] I woke up this morning and saw that Twitter has banned/ permanently suspended 4 of my developer accounts, I have around a few more but it is just a matter of time till all my accounts will most likely get banned as well. This was a fun project that I maintained for as long as I can. I will pull data til my last account gets banned.

    [Feb 26] I've started to pull in RETWEETS again, so I am expecting a significant amount of throughput in tweets again on top of the dedicated processes that I have that gets NONRETWEETS. If you don't want RETWEETS, just filter them out.

    [Feb 24] It's been a year since I started getting tweets of this conflict and had no idea that a year later this is still ongoing. Almost everyone assumed that Ukraine will crumble in a matter of days, but it is not the case. To those who have been using my dataset, i hope that I am helping all of you in one way or another. Ill do my best to maintain updating this dataset as long as I can.

    [Feb 02] I seem to be getting less tweets as my crawlers are getting throttled, i used to get 2500 tweets per 15 mins but around 2-3 of my crawlers are getting throttling limit errors. There may be some kind of update that Twitter has done about rate limits or something similar. Will try to find ways to increase the throughput again.

    [Jan 02] For all new datasets, it will now be prefixed by a year, so for Jan 01, 2023, it will be 20230101_XXXX.

    [Dec 28] For those looking for a cleaned version of my dataset, with the retweets removed from before Aug 08, here is a dataset by @@vbmokin https://www.kaggle.com/datasets/vbmokin/russian-invasion-ukraine-without-retweets

    [Nov 19] I noticed that one of my developer accounts, which ISNT TWEETING ANYTHING and just pulling data out of twitter has been permanently banned by Twitter.com, thus the decrease of unique tweets. I will try to come up with a solution to increase my throughput and signup for a new developer account.

    [Oct 19] I just noticed that this dataset is finally "GOLD", after roughly seven months since I first uploaded my gzipped csv files.

    [Oct 11] Sudden spike in number of tweets revolving around most recent development(s) about the Kerch Bridge explosion and the response from Russia.

    [Aug 19- IMPORTANT] I raised the missing dataset issue to Kaggle team and they confirmed it was a bug brought by a ReactJs upgrade, the conversation and details can be seen here https://www.kaggle.com/discussions/product-feedback/345915 . It has been fixed already and I've reuploaded all the gzipped files that were lost PLUS the new files that were generated AFTER the issue was identified.

    [Aug 17] Seems the latest version of my dataset lost around 100+ files, good thing this dataset is versioned so one can just go back to the previous version(s) and download them. Version 188 HAS ALL THE LOST FILES, I wont be reuploading all datasets as it will be tedious and I've deleted them already in my local and I only store the latest 2-3 days.

    [Aug 10] 3/5 of my Python processes errored out and resulted to around 10-12 hours of NO data gathering for those processes thus the sharp decrease of tweets for Aug 09 dataset. I've applied an exception/ error checking to prevent this from happening.

    [Aug 09] Significant drop in tweets extracted, but I am now getting ORIGINAL/ NON-RETWEETS.

    [Aug 08] I've noticed that I had a spike of Tweets extracted, but they are literally thousands of retweets of a single original tweet. I also noticed that my crawlers seem to deviate because of this tactic being used by some Twitter users where they flood Twitter w...

  15. Sentiment140 dataset (1,600,000 tweets)

    • kaggle.com
    zip
    Updated Jan 10, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Miloud Belarebia (2021). Sentiment140 dataset (1,600,000 tweets) [Dataset]. https://www.kaggle.com/datasets/milobele/sentiment140-dataset-1600000-tweets
    Explore at:
    zip(84886122 bytes)Available download formats
    Dataset updated
    Jan 10, 2021
    Authors
    Miloud Belarebia
    Description

    Context

    Sentiment140 dataset contains 1,600,000 tweets extracted from Twitter by using the Twitter API.

    Content

    The tweets have been categorized into three classes: - 0:negative - 2:neutral - 4:positive The information contained in the dataset: - The polarity of the tweet - id of the tweet - date of the tweet - query - User that tweeted - The content of the tweet. Dataset size: 305.13 MB

    Loading the dataset using TensorFlow import codecs import csv import os import tensorflow.compat.v2 as tf import tensorflow_datasets.public_api as tfds class Sentiment140(tfds.core.GeneratorBasedBuilder): VERSION = tfds.core.Version("1.0.0") def _info(self): return tfds.core.DatasetInfo( builder=self, features=tfds.features.FeaturesDict({ "polarity": tf.int32, "date": tfds.features.Text(), "query": tfds.features.Text(), "user": tfds.features.Text(), "text": tfds.features.Text(), }), supervised_keys=("text", "polarity"), homepage=_HOMEPAGE_URL, ) def _split_generators(self, dl_manager): dl_paths = dl_manager.download_and_extract(_DOWNLOAD_URL) return [ tfds.core.SplitGenerator( name=tfds.Split.TRAIN, gen_kwargs={ "path": os.path.join(dl_paths, "training.1600000.processed.noemoticon.csv") }), tfds.core.SplitGenerator( name=tfds.Split.TEST, gen_kwargs={ "path": os.path.join(dl_paths, "testdata.manual.2009.06.14.csv") }), ]

    Acknowledgements

    We wouldn't be here without the help of others. If you owe any attributions or thanks, include them here along with any citations of past research.

    Inspiration

    Your data will be in front of the world's largest data science community. What questions do you want to see answered?

  16. Drug-related Tweets Dataset

    • kaggle.com
    zip
    Updated Sep 17, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Techno Care (2025). Drug-related Tweets Dataset [Dataset]. https://www.kaggle.com/datasets/technocare/drug-related-tweets-dataset
    Explore at:
    zip(9522011 bytes)Available download formats
    Dataset updated
    Sep 17, 2025
    Authors
    Techno Care
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    This dataset contains drug-related text entries structured to resemble tweets. It was generated from the drugsComTest_raw.csv dataset, which originally included patient reviews of medications. Source: Extracted from patient-submitted reviews on Drugs.com. Format: CSV file with two columns: drugName – the name of the drug mentioned. tweet – the review text reformatted to simulate a tweet-like message.

    Purpose: To support Natural Language Processing (NLP) tasks such as sentiment analysis, drug-effect classification, and social media mining. To act as a proxy dataset for training or testing models on drug-related discussions, where actual Twitter data collection is restricted or unavailable.

    Limitations: Not real Twitter data, but synthetic tweets generated from formal drug reviews. May differ in tone and structure compared to actual tweets.

  17. Cryptocurrency Tweets

    • kaggle.com
    zip
    Updated Sep 10, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Waqas Aman (2024). Cryptocurrency Tweets [Dataset]. https://www.kaggle.com/datasets/infsceps/cryptocurrency-tweets
    Explore at:
    zip(9083119 bytes)Available download formats
    Dataset updated
    Sep 10, 2024
    Authors
    Waqas Aman
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Cryptocurrency Discussions on X (formerly Twitter):

    Dataset Overview: This dataset consists of random tweets related to cryptocurrency discussions collected from X (formerly Twitter) over a 12-month period, from June 2022 to May 2023. The dataset captures user-generated content, opinions, and sentiment around cryptocurrency, providing valuable insights into social media trends, market sentiment, and public discourse on various cryptocurrencies.

    Features: - Tweets: User-generated tweets on cryptocurrency discussions. - Time Period: Data collected from June 2022 to May 2023. - Content Focus: Conversations, mentions, and sentiment around different cryptocurrencies. - Preprocessing: Duplicate tweets removed, user information removed for privacy.

    Use Cases: - Sentiment Analysis: Assess public sentiment and trends in the cryptocurrency space. - Social Media Analysis: Examine patterns in cryptocurrency-related discussions on social media platforms. - Market Influence Studies: Analyze how public opinion on cryptocurrency affects market behavior. - Natural Language Processing (NLP): Train models for topic modeling, text classification, or opinion mining in the context of cryptocurrency.

    **Potential Impact: **This dataset is a valuable resource for researchers, data scientists, and cryptocurrency enthusiasts aiming to explore the intersection of social media and financial markets. It can support studies in sentiment analysis, public discourse, and the relationship between social media activity and cryptocurrency market dynamics.

  18. Twitter Data on #IndonesiaHumanRightsSOS

    • kaggle.com
    zip
    Updated Dec 4, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Francis (2024). Twitter Data on #IndonesiaHumanRightsSOS [Dataset]. https://www.kaggle.com/datasets/noeyislearning/twitter-data-on-indonesiahumanrightssos
    Explore at:
    zip(14269743 bytes)Available download formats
    Dataset updated
    Dec 4, 2024
    Authors
    Francis
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset provides a comprehensive overview of tweets related to the hashtag #IndonesiaHumanRightsSOS from December 18, 2020, 10:59 AM, to December 19, 2020, 23:18 PM. The data is sourced from Twitter using the twint application and offers detailed insights into social media engagement, user activity, and discussions on human rights. The dataset is structured to include key metrics such as user ID, username, Twitter name, tweets, mentions, URLs, photos, replies count, retweets count, likes count, hashtags, cashtags, and more, providing a robust foundation for analyzing social media trends and user behavior.

    Key Features

    • User Identification: The dataset includes unique identifiers for each user, such as user ID, username, and Twitter name, facilitating easy identification and tracking of user activity.
    • Temporal Precision: Data is categorized by date and time, offering insights into the timing and frequency of tweets.
    • Content Analysis: Information is presented by tweets, mentions, URLs, photos, and hashtags, allowing for detailed analysis of content and engagement patterns.
    • Engagement Metrics: The dataset includes metrics such as replies count, retweets count, and likes count, providing insights into user interaction and engagement levels.
    • Geolocation: Data includes timezone information, enabling analysis of regional engagement patterns.

    Potential Uses

    • Social Media Analysis: Assist in understanding social media trends and user behavior related to human rights discussions.
    • Sentiment Analysis: Support sentiment analysis by providing detailed data on user sentiments and reactions to human rights issues.
    • Policy Development: Inform policymakers in developing and adjusting human rights policies based on social media trends and user feedback.
    • Strategic Planning: Provide insights into social media engagement patterns, informing strategic planning for advocacy and awareness campaigns.
    • Research and Academic Studies: Enable research and academic studies on social media behavior and human rights issues.
  19. Customer Support on Twitter

    • kaggle.com
    zip
    Updated Dec 3, 2017
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Thought Vector (2017). Customer Support on Twitter [Dataset]. https://www.kaggle.com/dsv/8841
    Explore at:
    zip(176772673 bytes)Available download formats
    Dataset updated
    Dec 3, 2017
    Dataset authored and provided by
    Thought Vector
    License

    Attribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
    License information was derived automatically

    Description

    The Customer Support on Twitter dataset is a large, modern corpus of tweets and replies to aid innovation in natural language understanding and conversational models, and for study of modern customer support practices and impact.

    https://i.imgur.com/nTv3Iuu.png" alt="Example Analysis - Inbound Volume for the Top 20 Brands">

    Context

    Natural language remains the densest encoding of human experience we have, and innovation in NLP has accelerated to power understanding of that data, but the datasets driving this innovation don't match the real language in use today. The Customer Support on Twitter dataset offers a large corpus of modern English (mostly) conversations between consumers and customer support agents on Twitter, and has three important advantages over other conversational text datasets:

    • Focused - Consumers contact customer support to have a specific problem solved, and the manifold of problems to be discussed is relatively small, especially compared to unconstrained conversational datasets like the reddit Corpus.
    • Natural - Consumers in this dataset come from a much broader segment than those in the Ubuntu Dialogue Corpus and have much more natural and recent use of typed text than the Cornell Movie Dialogs Corpus.
    • Succinct - Twitter's brevity causes more natural responses from support agents (rather than scripted), and to-the-point descriptions of problems and solutions. Also, its convenient in allowing for a relatively low message limit size for recurrent nets.

    Inspiration

    The size and breadth of this dataset inspires many interesting questions:

    • Can we predict company responses? Given the bounded set of subjects handled by each company, the answer seems like yes!
    • Do requests get stale? How quickly do the best companies respond, compared to the worst?
    • Can we learn high quality dense embeddings or representations of similarity for topical clustering?
    • How does tone affect the customer support conversation? Does saying sorry help?
    • Can we help companies identify new problems, or ones most affecting their customers?

    Acknowledgements

    Dataset built with PointScrape.

    Content

    The dataset is a CSV, where each row is a tweet. The different columns are described below. Every conversation included has at least one request from a consumer and at least one response from a company. Which user IDs are company user IDs can be calculated using the inbound field.

    tweet_id

    A unique, anonymized ID for the Tweet. Referenced by response_tweet_id and in_response_to_tweet_id.

    author_id

    A unique, anonymized user ID. @s in the dataset have been replaced with their associated anonymized user ID.

    inbound

    Whether the tweet is "inbound" to a company doing customer support on Twitter. This feature is useful when re-organizing data for training conversational models.

    created_at

    Date and time when the tweet was sent.

    text

    Tweet content. Sensitive information like phone numbers and email addresses are replaced with mask values like _email_.

    response_tweet_id

    IDs of tweets that are responses to this tweet, comma-separated.

    in_response_to_tweet_id

    ID of the tweet this tweet is in response to, if any.

    Contributing

    Know of other brands the dataset should include? Found something that needs to be fixed? Start a discussion, or email me directly at $FIRSTNAME@$LASTNAME.com!

    Acknowledgements

    A huge thank you to my friends who helped bootstrap the list of companies that do customer support on Twitter! There are many rocks that would have been left un-turned were it not for your suggestions!

    Relevant Resources

    Licensing

    For commercial applications and use of full dataset, please contact stuart@thoughtvector.io.

  20. Product Tweets Dataset

    • kaggle.com
    zip
    Updated May 29, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dhwanil Shah (2022). Product Tweets Dataset [Dataset]. https://www.kaggle.com/datasets/dshah1612/product-tweets-dataset/discussion
    Explore at:
    zip(375561 bytes)Available download formats
    Dataset updated
    May 29, 2022
    Authors
    Dhwanil Shah
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    This dataset contains tweets posted for various services and products along with the emotion contained in the tweet. It contains three columns, the tweet text, the product/service, and the emotion contained in the tweet. It can be used to train various ML models for analyzing the sentiments in the tweets.

    If you find the data useful do give an upvote and let me know in the discussions about any improvements.

    Cheers!!

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Stuart Axelbrooke; Stuart Axelbrooke (2025). Customer Support on Twitter [Dataset]. http://doi.org/10.34740/kaggle/dsv/8841
Organization logo

Customer Support on Twitter

Explore at:
7 scholarly articles cite this dataset (View in Google Scholar)
csvAvailable download formats
Dataset updated
Jul 31, 2025
Dataset provided by
Kagglehttp://kaggle.com/
Authors
Stuart Axelbrooke; Stuart Axelbrooke
Time period covered
Mar 12, 2017
Description

The Customer Support on Twitter dataset is a large, modern corpus of tweets and replies to aid innovation in natural language understanding and conversational models, and for study of modern customer support practices and impact. The dataset includes replies of companies like Apple, Amazon, Uber, Delta, Spotify and others.

Search
Clear search
Close search
Google apps
Main menu