Saved datasets
Last updated
Download format
Croissant
Croissant is a format for Machine Learning datasets
Learn more about this at mlcommons.org/croissant.
Usage rights
License from data provider
Please review the applicable license to make sure your contemplated use is permitted.
Topic
Provider
Free
Cost to access
Described as free to access or have a license that allows redistribution.
43 datasets found
  1. Customer Support on Twitter

    • kaggle.com
    zip
    Updated Dec 3, 2017
  2. Twitter Tweets Sentiment Dataset

    • kaggle.com
    zip
    Updated Apr 8, 2022
  3. Twitter Customer Service Interaction Summarization

    • kaggle.com
    zip
    Updated May 14, 2026
  4. Twitter customer support twitter llm finetune

    • kaggle.com
    zip
    Updated Sep 1, 2025
  5. Support data for Chatbots

    • kaggle.com
    zip
    Updated Feb 26, 2025
  6. customer care tweets KSA

    • kaggle.com
    zip
    Updated May 20, 2022
  7. Customer Support Tickets Dataset

    • kaggle.com
    zip
    Updated Jul 8, 2026
  8. Customer Support Tweets (945M rows)

    • kaggle.com
    zip
    Updated Oct 31, 2025
  9. Twitter New Dataset 2024 March Data

    • kaggle.com
    zip
    Updated Mar 11, 2024
  10. US Airline Twitter Sentiment Analysis Dataset

    • kaggle.com
    zip
    Updated Feb 8, 2026
  11. Cryptocurrency Tweets

    • kaggle.com
    zip
    Updated Sep 10, 2024
  12. Apple Twitter Sentiment (CrowdFlower)

    • kaggle.com
    zip
    Updated Oct 6, 2021
  13. Hinduphobic COVID-19 X (Twitter) Dataset

    • kaggle.com
    zip
    Updated Feb 13, 2025
  14. Twitter Friends

    • kaggle.com
    zip
    Updated Sep 2, 2016
  15. Monkeypox misinformation: Twitter dataset

    • kaggle.com
    zip
    Updated Aug 31, 2022
  16. Hate Speech and Offensive Language Detection

    • kaggle.com
    zip
    Updated Dec 2, 2023
  17. Sentiment140 dataset with 1.6 million tweets

    • kaggle.com
    zip
    Updated Sep 13, 2017
  18. English Tweets Mentioning Bitcoin (2021-2022)

    • kaggle.com
    zip
    Updated Nov 21, 2022
  19. Product Tweets Dataset

    • kaggle.com
    zip
    Updated May 29, 2022
  20. Greekgodx Tweets: Analyzing Conversation

    • kaggle.com
    zip
    Updated Dec 27, 2022
Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Thought Vector (2017). Customer Support on Twitter [Dataset]. https://www.kaggle.com/dsv/8841
Organization logo

Customer Support on Twitter

Over 3 million tweets and replies from the biggest brands on Twitter

Explore at:
3 scholarly articles cite this dataset (View in Google Scholar)
zip(176772673 bytes)Available download formats
Dataset updated
Dec 3, 2017
Dataset authored and provided by
Thought Vector
License

Attribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically

Description

The Customer Support on Twitter dataset is a large, modern corpus of tweets and replies to aid innovation in natural language understanding and conversational models, and for study of modern customer support practices and impact.

https://i.imgur.com/nTv3Iuu.png" alt="Example Analysis - Inbound Volume for the Top 20 Brands">

Context

Natural language remains the densest encoding of human experience we have, and innovation in NLP has accelerated to power understanding of that data, but the datasets driving this innovation don't match the real language in use today. The Customer Support on Twitter dataset offers a large corpus of modern English (mostly) conversations between consumers and customer support agents on Twitter, and has three important advantages over other conversational text datasets:

  • Focused - Consumers contact customer support to have a specific problem solved, and the manifold of problems to be discussed is relatively small, especially compared to unconstrained conversational datasets like the reddit Corpus.
  • Natural - Consumers in this dataset come from a much broader segment than those in the Ubuntu Dialogue Corpus and have much more natural and recent use of typed text than the Cornell Movie Dialogs Corpus.
  • Succinct - Twitter's brevity causes more natural responses from support agents (rather than scripted), and to-the-point descriptions of problems and solutions. Also, its convenient in allowing for a relatively low message limit size for recurrent nets.

Inspiration

The size and breadth of this dataset inspires many interesting questions:

  • Can we predict company responses? Given the bounded set of subjects handled by each company, the answer seems like yes!
  • Do requests get stale? How quickly do the best companies respond, compared to the worst?
  • Can we learn high quality dense embeddings or representations of similarity for topical clustering?
  • How does tone affect the customer support conversation? Does saying sorry help?
  • Can we help companies identify new problems, or ones most affecting their customers?

Acknowledgements

Dataset built with PointScrape.

Content

The dataset is a CSV, where each row is a tweet. The different columns are described below. Every conversation included has at least one request from a consumer and at least one response from a company. Which user IDs are company user IDs can be calculated using the inbound field.

tweet_id

A unique, anonymized ID for the Tweet. Referenced by response_tweet_id and in_response_to_tweet_id.

author_id

A unique, anonymized user ID. @s in the dataset have been replaced with their associated anonymized user ID.

inbound

Whether the tweet is "inbound" to a company doing customer support on Twitter. This feature is useful when re-organizing data for training conversational models.

created_at

Date and time when the tweet was sent.

text

Tweet content. Sensitive information like phone numbers and email addresses are replaced with mask values like _email_.

response_tweet_id

IDs of tweets that are responses to this tweet, comma-separated.

in_response_to_tweet_id

ID of the tweet this tweet is in response to, if any.

Contributing

Know of other brands the dataset should include? Found something that needs to be fixed? Start a discussion, or email me directly at $FIRSTNAME@$LASTNAME.com!

Acknowledgements

A huge thank you to my friends who helped bootstrap the list of companies that do customer support on Twitter! There are many rocks that would have been left un-turned were it not for your suggestions!

Relevant Resources

Licensing

For commercial applications and use of full dataset, please contact stuart@thoughtvector.io.

Search
Clear search
Close search
Google apps
Main menu