Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
TikTok is one of the hottest social media platforms out there, and it's only getting bigger. If you're looking to get in on the action, this dataset is for you!
This dataset contains a collection of videos from TikTok, including information on the user who posted the video, the number of likes, shares, and comments the video received, as well as the video's length and description. With this data, you can see what types of videos are popular on TikTok and start planning your own viral content!
- The dataset contains a collection of videos from the social media platform TikTok.
- The videos include information on the user who posted the video, the number of likes, shares, and comments the video received, as well as the video's length and description.
- The dataset also contains information on popular TikTok authors, including their unique ID, nickname, avatar thumbnail, signature, and whether or not their account is verified or private.
- Additionally, the dataset includes a list of trending videos on TikTok, as well as the number of likes, shares, comments, and plays each video has received
- Identifying popular TikTok authors to target for scraping videos and liked videos
- Finding trending videos on TikTok for further analysis
- Generating a list of videos from the TikTok app that are tagged with the #funny hashtag
License
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: tiktok_collected_liked_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | user_name | The name of the user who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of shares the video has received. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_collected_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | user_name | The name of the user who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of shares the video has received. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_funny_hashtag_videos.csv | Column name | Description | |:--------------------------|:-----------------------------------------------------------| | author_nickname | The author's nickname. (String) | | author_avatarThumb | The author's avatar thumbnail. (String) | | author_signature | The author's signature. (String) | | author_verification | Whether or not the author's account is verified. (Boolean) | | author_privateAccount | Whether or not the author's account is private. (Boolean) | | author_followingCount | The number of people the author is following. (Integer) | | author_followerCount | The number of people following the author. (Integer) | | author_heartCount | The number of hearts the author has. (Integer) | | author_diggCount | The number of diggs the author has. (Integer) | | music_title | The title of the music. (String) | | music_playUrl | The play url of the music. (String) | | music_coverThumb | The cover thumbnail of the music. (String) | | music_authorName | The author name of the music. (String) | | music_originality | The originality of the music. (String) | | music_duration | The duration of the music. (String) |
File: trending_authors.csv | Column name | Description ...
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains information about TikTok videos, including user interactions and video details. It includes features such as video ID, username, video title, likes, comments, shares, views, and more. This dataset is useful for analyzing video performance and user engagement on TikTok.
Columns:
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
We are probably all familiar with TikTok. People tend to spend hours each day scrolling through the millions of videos which are uploaded every single day. Not to mention the uploaders who are giving anything to get as many likes and followers as possible. But what makes one TikTok video a true hit or a miss? I give you an opportunity to figure this out ;)
I scraped the first 1000 trending videos on TikTok, using an unofficial TikTok web-scraper. Note to mention I had to provide my user information to scrape the trending information, so trending might be a personalized page. But that doesn't change the fact that certain people and videos got a certain amount of likes and comments.
I transformed the data into usable csv files and attached the actual videos as well.
Videos.zip This file contains the actual 1000 trending TikTok videos. Each filename corresponds to the id key in the trending.json file.
trending.json The raw scraped dataset. I figured splitting up the dataset resulted in messy errors. For example: a user might have one avatar while posting a video and another while posting the next video. This resulted in multiple users with the same name, id etc. except for the avatar. So I decided to post the raw data and I will show you how to translate this multi-level JSON structure to a single DataFrame in my first Notebook.
Many thanks to Andrew Nord the creator of the tiktok-scraper, and his contributers.
So what does make a TikTok video a true hit? Is it the moment when a video is uploaded? Or perhaps the amount of followers is an important factor? Maybe the hashtags or even the music being used?
So... are you the one who unlocks the mystery?
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by Marcus Ong
Released under CC0: Public Domain
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by Marcus Ong
Released under CC0: Public Domain
Facebook
TwitterThe dataset was originally obtained from TikTok's trending API by a GitHub user named Ivan Tran. It contains metadata on engagement with user-created videos and user profile data. The original create time is in Unix timecode format and is extracted directly from the video id number. TikTok's API has become much more difficult to access recently, so more current data is harder to obtain. The hashtags column contains lists.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset records various features of top trending videos on TikTok and Youtube Shorts in the summer of 2022. Features include video (theme, type, style, length), and music(genre, release year, and part of the music used).
For use of data examples, please refer to the dashboards I made with Tableau here: TikTok Top Trending Video dashboard: https://public.tableau.com/app/profile/caroline.zhu6047/viz/TopTrendingVideoDashboard_16691429927590/Overview
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset captures the pulse of viral social media trends across TikTok, Instagram, Twitter, and YouTube. It provides insights into the most popular hashtags, content types, and user engagement levels, offering a comprehensive view of how trends unfold across platforms. With regional data and influencer-driven content, this dataset is perfect for:
Dive in to explore what makes content go viral, the behaviors that drive engagement, and how trends evolve on a global scale! 🌍
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Short-form content dominates every social platform — YouTube Shorts, Instagram Reels, and TikTok. Creators and analysts are constantly searching for insights to understand what makes a video go viral: the hook, the niche, the music, or the first-hour views.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains information on TikTok users' reports of videos and comments that include user claims. These reports flag content for moderator review, generating a significant volume of user reports that need timely attention.
TikTok is developing a predictive model to determine whether a video contains a claim or offers an opinion. A successful prediction model will help reduce the backlog of user reports and enable more efficient prioritization.
This dataset is intended for exploratory data analysis (EDA), statistical analysis, and predictive modeling. It has been created for pedagogical purposes and aims to facilitate learning and research in data analysis and machine learning
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
How do you measure the success of a video on social media? Is it the number of likes? The number of shares? The number of comments?
This dataset contains information on videos posted to the social media platform TikTok. The data includes the video ID, description, creation time, length, number of likes, shares, and comments, as well as a link to the video.
With this data, you can explore what factors make a video popular on TikTok and learn more about user preferences on this rapidly growing social media platform
This dataset can be used to study user preferences in social media. The data includes the number of likes, shares, comments, and plays for each video, as well as the video's description, length, and link
- Identifying trends in social media
- Analyzing user preferences in social media
- Predicting future trends in social media
Dataset by TikTok
License
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: omnibuslaw_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | createTime | The date and time the video was posted. (DateTime) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_liked_videos.csv | Column name | Description | |:---------------|:----------------------------------------------------------| | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) | | user_name | The username of the person who posted the video. (String) |
File: trending.csv | Column name | Description | |:---------------|:----------------------------------------------------------| | user_name | The username of the person who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: washingtonpost_videos.csv | Column name | Description | |:---------------|:----------------------------------------------------------| | user_name | The username of the person who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
As a recent user of TikTok, I was interested in working on a dataset that helps me understand how the spam/cancel culture works its ways among famous creators. I decided to focus on one of my favorite creators, David Dobrik and pick his top videos and collect the comments. This data is rich in the actual commenter's profile and metadata which adds an additional layer to detecting "true" fans from spammers.
As you go through the different columns, it's easy to understand the nature of the data starting with the actual comment and the ID of the video where it was posted, the number of likes per each comment, the country of origin, and most importantly the profile of the poster and whether they're verified or not.
To access the videos, you can just plug in the video URL after's David's username e.g. https://www.tiktok.com/@daviddobrik/video/6877635569963273478
I used RapidAPI's TikTok API (to add a link soon)
What questions do you want to see answered? - Who are David Dobrik's true fans vs. spammers? - New research on spam detection algorithms now applied on TikTok (I haven't found any online) - Understand the demographics of some of these viral videos and how each user and creator fit into the picture
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
https://s3-prod.adage.com/s3fs-public/20230807_celeb_run_agencies_3x2.jpg" alt="Celebs">
The dataset you provided appears to focus on TikTok celebrities and contains the following columns:
Celebrity: The name or handle of the TikTok celebrity. Followers: The number of followers the celebrity has, often represented in millions or billions. Following: The number of accounts the celebrity follows, which may be represented as thousands (K) or just a number. Likes: The total number of likes the celebrity’s videos have received, often represented in millions or billions. T.Videos: The total number of videos posted by the celebrity. Video Duration: The typical duration of their videos, which ranges from a few seconds (e.g., 10 - 15 seconds) to over a minute. Average Views: The average number of views their videos receive, often in millions. Net Worth: The estimated net worth of the celebrity, often represented in millions or billions of dollars. Most Viewed Video: The number of views for their most popular video, usually in millions or billions. Most Liked Video: The number of likes for their most popular video, represented in millions or billions. Video Category: The types or categories of videos the celebrity posts, such as comedy, dance, acting, challenges, etc.
Facebook
TwitterAs of January 2022, The United States was the country with the largest TikTok audience by far, with approximately 131 million users engaging with the popular social video platform. Indonesia followed, with around 92 million TikTok users. Brazil came in third, with 74 million users using TikTok to watch short-videos.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
A global dataset capturing short-form video performance across YouTube Shorts and TikTok in 2025.
It includes over 50,000 video records, available in both raw and machine learning–ready formats.
Designed for reproducible EDA, dashboarding, and baseline ML modeling on social media engagement dynamics.
| File | Description | Shape |
|---|---|---|
youtube_shorts_tiktok_trends_2025.csv | Raw video-level data with full feature set | ~48k × ~58 |
youtube_shorts_tiktok_trends_2025_ml.csv | ML-ready, cleaned and engineered version | ~50k × 32 |
monthly_trends_2025.csv | Monthly aggregates (Jan–Aug 2025) | ~480 × 8 |
country_platform_summary_2025.csv | Country × platform summary statistics | ~60 × 14 |
top_hashtags_2025.csv | Ranked list of top trending hashtags | ~82 × 18 |
top_creators_impact_2025.csv | Creator-level impact and influence metrics | ~1,000 × 20 |
DATA_DICTIONARY.csv | Column names and definitions | ~58 × 2 |
All files are UTF-8 encoded, cleaned, and schema-aligned for direct analysis.
video_id, platform, country, category, creator_tierviews, likes, comments, shares, saves, completionsengagement_rate = (likes + comments + shares) / views, plus save_rate, share_rate, comment_ratetrend_label or predict engagement_rate and views trend_label is a snapshot trend proxy; baseline models typically reach 25–35% accuracy without temporal features. publish_date_approx is derived and coarse — for trend direction only. If you find this dataset helpful, supporting it with an upvote helps others discover it too ✨
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Techsalerator’s YouTube & Video Data for Indonesia
Techsalerator’s YouTube & Video Data for Indonesia aggregates comprehensive insights from video platforms to provide a detailed view of the country’s digital video ecosystem. This dataset captures video metadata, creator activity, audience engagement, and content performance signals across Indonesia, helping organizations analyze content trends, viewer behavior, and digital media consumption patterns.
This dataset is designed to support media platforms, advertisers, researchers, NGOs, and policy analysts seeking insights into video engagement, creator growth, and content distribution across Indonesia.
For access to the full dataset, contact us at info@techsalerator.com or visit Techsalerator Contact Us.
Video Metadata (Title, Category, Language, Upload Date)
Identifies core attributes of videos including topic, format, and publication timing.
Channel & Creator Information
Captures data on content creators, subscriber counts, upload frequency, and channel growth.
Engagement Metrics (Views, Likes, Comments, Shares)
Measures audience interaction and content performance across videos.
Audience Demographics & Geography
Provides insights into viewer location, age distribution, and engagement behavior.
Watch Time & Retention Metrics
Tracks average viewing duration, retention rates, and drop-off points.
One of the World’s Largest Mobile Video Audiences
Indonesia has massive mobile-first video consumption driven by affordable smartphones and data access.
High Demand for Bahasa Indonesia and Regional Language Content
Content in Bahasa Indonesia dominates, alongside Javanese and other local languages.
Rapid Growth of Creator Economy and Influencer Culture
Indonesian creators are among the most active in Southeast Asia.
Strong Engagement with Entertainment, Music, and Lifestyle Content
Music, comedy, vlogs, and entertainment content perform exceptionally well.
Heavy Reliance on Social Media for Video Discovery
Platforms like YouTube, TikTok, Instagram, and WhatsApp drive content distribution.
Digital Advertising & Audience Targeting
Enables large-scale, language-based, and behavior-based targeting.
Influencer Marketing & Brand Campaigns
Supports discovery of high-performing creators across diverse niches.
Media & Entertainment Industry Insights
Helps broadcasters, studios, and OTT platforms understand demand patterns.
Consumer Behavior & Cultural Research
Provides insights into regional content preferences and digital habits.
Education, Government & Public Awareness Campaigns
Supports large-scale video-based communication initiatives across sectors.
To obtain Techsalerator’s YouTube & Video Data for Indonesia, contact us at info@techsalerator.com with your specific data requirements. Custom datasets, historical records, and real-time analytics are available, with delivery within 24 hours and flexible access agreements upon request.
For actionable insights into video performance, audience behavior, and digital media trends in Indonesia, Techsalerator’s YouTube & Video Data empowers advertisers, researchers, media companies, NGOs, and policymakers with reliable, structured, and scalable intelligence.
📩 Email: info@techsalerator.com
Facebook
TwitterTikTok's platform is mostly fueled by viral videos of users doing outlandish, scary, or funny things. On the platform, these trend and meme videos typically come with a hashtag that includes the word challenge. But what is a TikTok challenge and how do you find or create them? Here's everything you need to know.
This TikTok book challenge was made by @haleyisfearless, . It asks you to show, your prettiest book,your tiniest book a book you highly suggest a book you're currently reading and one of your favorite books . In the most basic sense, these challenges originate from viral TikTok challenge isn't complete without its defining hashtag in the video's description
These TikTok challenges are the perfect way to ease into what can be an intimidating social media platform and help you find your fellow book lovers.
This dataset is generated entirely from TikTok , so we want to thank @haleyisfearless for building such this challange video
the goal of this project is to make Python script which takes a video as input and returns all texts visible on the video. the videos are titlok videos so texts can appear everywhere on screen, with different background, font size etc..
Facebook
TwitterI always notice how TikTok videos make viewers laugh - this led to me wonder: What exactly about TikTok made people happy? Is it the video length? Or is it the music / sounds? Or is it the content? With these questions in mind, I scraped all the videos from Top 5 influencers from 8 selected countries: Australia, Indonesia, Japan, Norway, Russia, Singapore, South Korea, US, and UK.
For every video (row), the information included are the variables : user_name, user_id, video_id, video_desc, video_time, video_length, video_link, n_likes, n_shares, n_comments, n_plays, video_timestamp, country, year.
user_name: user name of the user who posted the video
user_id: the id of the user recorded
video_id: the id of the video posted
video_desc: the description or caption of the video, written by the user
video_time: the time of posting of the video in UTC format
video_length: the length of the video in seconds
video_link: the url link to the video
n_likes: the number of likes received by the video
n_shares: the number of shares received by the video
n_plays: the number of plays recorded by the video
video_timestamp: the date of posting of the video, converted from UTC
country: the country of the user who posted the video
year: the year of posting the video
Huge thank you to the unofficial TikTok Api and its creator (@davidteather on Github) for letting this webscraping process become a lot more easy!
What sort of content in TikTok makes people happy?
Facebook
Twitterhttp://opendatacommons.org/licenses/dbcl/1.0/http://opendatacommons.org/licenses/dbcl/1.0/
This dataset is an extension of the TikHarm dataset, created to enhance multimodal harmful content detection on TikTok. It was developed as part of the MTikGuard system, a real-time moderation pipeline designed to protect young audiences from unsafe TikTok videos.
🔹 Purpose
The dataset supplements TikHarm with 775 additional annotated videos, collected from TikTok trending and targeted hashtag queries. These videos were selected to address class imbalance and content diversity gaps in the original dataset, improving model generalization for real-world deployment.
🔹 Content
Each video is labeled into one of four categories: - Safe - Adult Content - Harmful Content (e.g., dangerous challenges, graphic violence) - Suicide / Self-harm
🔹 Data Collection & Annotation
Collection: Automated crawling using Selenium and TikTok Content Scraper, coordinated via Apache Airflow and Apache Kafka.
Annotation: Conducted via a custom web-based tool, following detailed guidelines to ensure consistency and reliability. Multiple annotators reviewed each video, with disagreements resolved via majority voting.
Class balance: Oversampling of underrepresented categories (e.g., Suicide, Harmful Content) during collection.
🔹 Applications
Training and evaluating multimodal classification models for harmful content detection.
Benchmarking real-time content moderation pipelines.
Research on multimodal fusion strategies and multi-label classification.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Please upvote if you like this dataset
TikTok, known in China as Douyin (Chinese: 抖音; pinyin: Dǒuyīn), is a short-form video hosting service owned by Chinese company ByteDance. It hosts a variety of short-form user videos, from genres like pranks, stunts, tricks, jokes, dance, and entertainment with durations from 15 seconds to ten minutes. TikTok is an international version of Douyin, which was originally released in the Chinese market in September 2016. TikTok was launched in 2017 for iOS and Android in most markets outside of mainland China; however, it became available worldwide only after merging with another Chinese social media service, Musical.ly, on 2 August 2018.
TikTok and Douyin have almost the same user interface but no access to each other's content. Their servers are each based in the market where the respective app is available. The two products are similar, but features are not identical. Douyin includes an in-video search feature that can search by people's faces for more videos of them and other features such as buying, booking hotels and making geo-tagged reviews. Since its launch in 2016, TikTok and Douyin rapidly gained popularity in virtually all parts of the world. TikTok surpassed 2 billion mobile downloads worldwide in October 2020.
In this dataset you will find the details about top 1000 tiktokers all over the world.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
TikTok is one of the hottest social media platforms out there, and it's only getting bigger. If you're looking to get in on the action, this dataset is for you!
This dataset contains a collection of videos from TikTok, including information on the user who posted the video, the number of likes, shares, and comments the video received, as well as the video's length and description. With this data, you can see what types of videos are popular on TikTok and start planning your own viral content!
- The dataset contains a collection of videos from the social media platform TikTok.
- The videos include information on the user who posted the video, the number of likes, shares, and comments the video received, as well as the video's length and description.
- The dataset also contains information on popular TikTok authors, including their unique ID, nickname, avatar thumbnail, signature, and whether or not their account is verified or private.
- Additionally, the dataset includes a list of trending videos on TikTok, as well as the number of likes, shares, comments, and plays each video has received
- Identifying popular TikTok authors to target for scraping videos and liked videos
- Finding trending videos on TikTok for further analysis
- Generating a list of videos from the TikTok app that are tagged with the #funny hashtag
License
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: tiktok_collected_liked_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | user_name | The name of the user who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of shares the video has received. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_collected_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | user_name | The name of the user who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of shares the video has received. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_funny_hashtag_videos.csv | Column name | Description | |:--------------------------|:-----------------------------------------------------------| | author_nickname | The author's nickname. (String) | | author_avatarThumb | The author's avatar thumbnail. (String) | | author_signature | The author's signature. (String) | | author_verification | Whether or not the author's account is verified. (Boolean) | | author_privateAccount | Whether or not the author's account is private. (Boolean) | | author_followingCount | The number of people the author is following. (Integer) | | author_followerCount | The number of people following the author. (Integer) | | author_heartCount | The number of hearts the author has. (Integer) | | author_diggCount | The number of diggs the author has. (Integer) | | music_title | The title of the music. (String) | | music_playUrl | The play url of the music. (String) | | music_coverThumb | The cover thumbnail of the music. (String) | | music_authorName | The author name of the music. (String) | | music_originality | The originality of the music. (String) | | music_duration | The duration of the music. (String) |
File: trending_authors.csv | Column name | Description ...