Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains information about TikTok videos, including user interactions and video details. It includes features such as video ID, username, video title, likes, comments, shares, views, and more. This dataset is useful for analyzing video performance and user engagement on TikTok.
Columns:
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by Marcus Ong
Released under CC0: Public Domain
Facebook
TwitterThis dataset contains comprehensive information about TikTok posts, originally fetched from RapidAPI. It provides valuable insights into various aspects of TikTok content, including details about the videos, their creators, and audience engagement metrics.
Here's a breakdown of the columns included in this dataset:
video_id: A unique identifier for each TikTok video. author: The username or handle of the TikTok account that posted the video. description: The textual description or caption provided by the creator for the video. (Note: This column contains some missing values.) likes: The number of likes the video has received. comments: The number of comments on the video. shares: The number of times the video has been shared. plays: The total number of plays or views the video has accumulated. (Note: This column contains some missing values.) hashtags: A list of hashtags used in the video's description, which helps categorize content and improve discoverability. (Note: This column contains some missing values.) music: Information about the background music or sound used in the video. create_time: The timestamp indicating when the video was created or published. (Note: This column contains some missing values.) video_url: The direct URL to the TikTok video. fetch_time: The timestamp when the data for the video was fetched from the API. (Note: This column has a high number of missing values.) views: Another metric for the number of views. (Note: This column has a high number of missing values and appears to overlap with plays.) posted_time: The time the video was posted. (Note: This column has a high number of missing values and appears to overlap with create_time.) Potential Uses of This Dataset:
Content Analysis: Analyze popular TikTok content by examining descriptions, hashtags, and engagement metrics. Trend Identification: Identify trending topics, music, and creators on TikTok. Audience Engagement Studies: Understand how different types of content generate likes, comments, shares, and plays. Creator Analysis: Study the posting habits and performance of various TikTok creators. Social Media Research: Conduct research on the dynamics of content dissemination and user interaction on short-form video platforms. Notes on Data Quality:
The description, plays, hashtags, and create_time columns have some missing values, which may require handling (e.g., imputation or removal) depending on your analysis. The fetch_time, views, and posted_time columns are largely empty, suggesting they may not be reliable for comprehensive analysis. It is recommended to primarily rely on create_time for timestamps and plays for engagement metrics. This dataset can be a valuable resource for anyone looking to explore the vast and dynamic world of TikTok content and user engagement.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by Marcus Ong
Released under CC0: Public Domain
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset records various features of top trending videos on TikTok and Youtube Shorts in the summer of 2022. Features include video (theme, type, style, length), and music(genre, release year, and part of the music used).
For use of data examples, please refer to the dashboards I made with Tableau here: TikTok Top Trending Video dashboard: https://public.tableau.com/app/profile/caroline.zhu6047/viz/TopTrendingVideoDashboard_16691429927590/Overview
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
his dataset contains a large collection of TikTok video metadata fetched using the TikTok Scraper API. It includes videos from multiple regions (e.g., US, India,) and categories (e.g., fyp, dance, comedy, food, travel, etc.). Each video entry
provides detailed information such as:
Video ID: Unique identifier for the video. Region: The region where the video is popular. Category: The keyword/category used to fetch the video (e.g., dance, comedy). Title: The title of the video. Duration: The length of the video in seconds. Play URL: Direct link to the video. Watermarked URL: Link to the watermarked version of the video. Cover Image: URL of the video's cover image. Music URL: Link to the music used in the video. Timestamp: The date and time when the data was fetched.
How This Dataset Can Be Helpful
Trend Analysis: Analyze trending videos across different regions and categories. Identify patterns in video popularity based on region, duration, or category.
Machine Learning: Train models to predict video popularity based on features like duration, region, and category. Build recommendation systems for TikTok videos.
Content Moderation: Use the dataset to analyze video content for moderation purposes.
Sentiment Analysis: Perform sentiment analysis on video titles to understand user preferences.
Cross-Region Insights: Compare video trends across different regions to understand cultural differences.
How to Use This Dataset Filter by Region: Analyze videos from a specific region (e.g., US or India).
Filter by Category: Focus on videos from a specific category (e.g., dance or comedy).
Trend Analysis: Identify trending videos based on timestamp and region.
Machine Learning: Use the dataset to train models for video popularity prediction or recommendation systems.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
TikTok is one of the hottest social media platforms out there, and it's only getting bigger. If you're looking to get in on the action, this dataset is for you!
This dataset contains a collection of videos from TikTok, including information on the user who posted the video, the number of likes, shares, and comments the video received, as well as the video's length and description. With this data, you can see what types of videos are popular on TikTok and start planning your own viral content!
- The dataset contains a collection of videos from the social media platform TikTok.
- The videos include information on the user who posted the video, the number of likes, shares, and comments the video received, as well as the video's length and description.
- The dataset also contains information on popular TikTok authors, including their unique ID, nickname, avatar thumbnail, signature, and whether or not their account is verified or private.
- Additionally, the dataset includes a list of trending videos on TikTok, as well as the number of likes, shares, comments, and plays each video has received
- Identifying popular TikTok authors to target for scraping videos and liked videos
- Finding trending videos on TikTok for further analysis
- Generating a list of videos from the TikTok app that are tagged with the #funny hashtag
License
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: tiktok_collected_liked_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | user_name | The name of the user who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of shares the video has received. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_collected_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | user_name | The name of the user who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of shares the video has received. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_funny_hashtag_videos.csv | Column name | Description | |:--------------------------|:-----------------------------------------------------------| | author_nickname | The author's nickname. (String) | | author_avatarThumb | The author's avatar thumbnail. (String) | | author_signature | The author's signature. (String) | | author_verification | Whether or not the author's account is verified. (Boolean) | | author_privateAccount | Whether or not the author's account is private. (Boolean) | | author_followingCount | The number of people the author is following. (Integer) | | author_followerCount | The number of people following the author. (Integer) | | author_heartCount | The number of hearts the author has. (Integer) | | author_diggCount | The number of diggs the author has. (Integer) | | music_title | The title of the music. (String) | | music_playUrl | The play url of the music. (String) | | music_coverThumb | The cover thumbnail of the music. (String) | | music_authorName | The author name of the music. (String) | | music_originality | The originality of the music. (String) | | music_duration | The duration of the music. (String) |
File: trending_authors.csv | Column name | Description ...
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
How do you measure the success of a video on social media? Is it the number of likes? The number of shares? The number of comments?
This dataset contains information on videos posted to the social media platform TikTok. The data includes the video ID, description, creation time, length, number of likes, shares, and comments, as well as a link to the video.
With this data, you can explore what factors make a video popular on TikTok and learn more about user preferences on this rapidly growing social media platform
This dataset can be used to study user preferences in social media. The data includes the number of likes, shares, comments, and plays for each video, as well as the video's description, length, and link
- Identifying trends in social media
- Analyzing user preferences in social media
- Predicting future trends in social media
Dataset by TikTok
License
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: omnibuslaw_videos.csv | Column name | Description | |:---------------|:---------------------------------------------------------| | createTime | The date and time the video was posted. (DateTime) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: tiktok_liked_videos.csv | Column name | Description | |:---------------|:----------------------------------------------------------| | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) | | user_name | The username of the person who posted the video. (String) |
File: trending.csv | Column name | Description | |:---------------|:----------------------------------------------------------| | user_name | The username of the person who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
File: washingtonpost_videos.csv | Column name | Description | |:---------------|:----------------------------------------------------------| | user_name | The username of the person who posted the video. (String) | | n_likes | The number of likes the video has received. (Integer) | | n_shares | The number of times the video has been shared. (Integer) | | n_comments | The number of comments the video has received. (Integer) | | n_plays | The number of times the video has been played. (Integer) |
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
https://s3-prod.adage.com/s3fs-public/20230807_celeb_run_agencies_3x2.jpg" alt="Celebs">
The dataset you provided appears to focus on TikTok celebrities and contains the following columns:
Celebrity: The name or handle of the TikTok celebrity. Followers: The number of followers the celebrity has, often represented in millions or billions. Following: The number of accounts the celebrity follows, which may be represented as thousands (K) or just a number. Likes: The total number of likes the celebrity’s videos have received, often represented in millions or billions. T.Videos: The total number of videos posted by the celebrity. Video Duration: The typical duration of their videos, which ranges from a few seconds (e.g., 10 - 15 seconds) to over a minute. Average Views: The average number of views their videos receive, often in millions. Net Worth: The estimated net worth of the celebrity, often represented in millions or billions of dollars. Most Viewed Video: The number of views for their most popular video, usually in millions or billions. Most Liked Video: The number of likes for their most popular video, represented in millions or billions. Video Category: The types or categories of videos the celebrity posts, such as comedy, dance, acting, challenges, etc.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
This dataset, titled TikTok Viral Trends 2025, provides a curated snapshot of 50 trending TikTok videos from September 2025, capturing the platform's dynamic content landscape. Sourced from real-time web analyses and social media insights (e.g., X posts, trend reports from reputable sources like Ramdam, NapoleonCat, and Tokchart), it focuses on viral videos across diverse categories such as Entertainment, Music, Comedy, Lifestyle, Beauty, Sustainability, and Technology. The dataset is designed for data scientists, researchers, and enthusiasts interested in analyzing social media trends, predicting virality, or exploring multimodal machine learning applications (e.g., NLP, time-series, or clustering). It stands out from existing Kaggle datasets by offering fresh, 2025-specific data with rich metadata, including engagement metrics, hashtags, and sound/trend associations.
tiktok_data.csv).post:72, web:65).The dataset contains the following 12 columns:
- video_id: Unique identifier for each video or trend (integer or hashtag-based).
- author: Creator username or group (anonymized as "Unknown" where not specified).
- description: Brief summary of the video content or trend, derived from source context.
- upload_date: Approximate or exact posting date (YYYY-MM-DD).
- views: Reported view count (e.g., millions, billions for hashtag aggregates; "N/A" if unavailable).
- likes: Reported like count (e.g., thousands, millions; "N/A" if unavailable).
- shares: Share count (often "N/A" due to limited public data).
- comments: Comment count (often "N/A" due to limited public data).
- hashtags: Key hashtags associated with the video or trend (e.g., #Kpop, #Viral).
- category: Inferred content category (e.g., Entertainment, Music, Comedy, Lifestyle, Sustainability, Tech).
- sound_or_trend: Associated audio track or challenge name driving the trend (e.g., "Soda Pop dance", "JUMP").
- source: Citation of data origin (e.g., post:72 for X post ID, web:65 for web source ID).
#Perfume reaching 39.3B views.This dataset is ideal for a variety of machine learning and data analysis tasks on Kaggle, including but not limited to:
- Virality Prediction: Use views, likes, and hashtags to train regression or classification models (e.g., XGBoost, neural networks) to predict video success.
- Trend Analysis: Apply clustering (e.g., K-means) or topic modeling (e.g., LDA) to identify emerging content themes or regional differences.
- NLP Applications: Analyze descriptions and hashtags with BERT or word embeddings to study sentiment, cultural trends, or influencer impact.
- Time-Series Forecasting: Leverage upload_date and engagement metrics for temporal analysis of trend lifecycles.
- Recommendation Systems: Build content recommendation models based on category, sound, or hashtag similarities.
- Social Media Ethics: Explore AI-driven trends (e.g., deepfake Identity Swaps) for studies on misinformation or content authenticity.
#Ominous). Exact metrics may vary slightly due to real-time fluctuations.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
As a recent user of TikTok, I was interested in working on a dataset that helps me understand how the spam/cancel culture works its ways among famous creators. I decided to focus on one of my favorite creators, David Dobrik and pick his top videos and collect the comments. This data is rich in the actual commenter's profile and metadata which adds an additional layer to detecting "true" fans from spammers.
As you go through the different columns, it's easy to understand the nature of the data starting with the actual comment and the ID of the video where it was posted, the number of likes per each comment, the country of origin, and most importantly the profile of the poster and whether they're verified or not.
To access the videos, you can just plug in the video URL after's David's username e.g. https://www.tiktok.com/@daviddobrik/video/6877635569963273478
I used RapidAPI's TikTok API (to add a link soon)
What questions do you want to see answered? - Who are David Dobrik's true fans vs. spammers? - New research on spam detection algorithms now applied on TikTok (I haven't found any online) - Understand the demographics of some of these viral videos and how each user and creator fit into the picture
Facebook
TwitterThe dataset was originally obtained from TikTok's trending API by a GitHub user named Ivan Tran. It contains metadata on engagement with user-created videos and user profile data. The original create time is in Unix timecode format and is extracted directly from the video id number. TikTok's API has become much more difficult to access recently, so more current data is harder to obtain. The hashtags column contains lists.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
A global dataset capturing short-form video performance across YouTube Shorts and TikTok in 2025.
It includes over 50,000 video records, available in both raw and machine learning–ready formats.
Designed for reproducible EDA, dashboarding, and baseline ML modeling on social media engagement dynamics.
| File | Description | Shape |
|---|---|---|
youtube_shorts_tiktok_trends_2025.csv | Raw video-level data with full feature set | ~48k × ~58 |
youtube_shorts_tiktok_trends_2025_ml.csv | ML-ready, cleaned and engineered version | ~50k × 32 |
monthly_trends_2025.csv | Monthly aggregates (Jan–Aug 2025) | ~480 × 8 |
country_platform_summary_2025.csv | Country × platform summary statistics | ~60 × 14 |
top_hashtags_2025.csv | Ranked list of top trending hashtags | ~82 × 18 |
top_creators_impact_2025.csv | Creator-level impact and influence metrics | ~1,000 × 20 |
DATA_DICTIONARY.csv | Column names and definitions | ~58 × 2 |
All files are UTF-8 encoded, cleaned, and schema-aligned for direct analysis.
video_id, platform, country, category, creator_tierviews, likes, comments, shares, saves, completionsengagement_rate = (likes + comments + shares) / views, plus save_rate, share_rate, comment_ratetrend_label or predict engagement_rate and views trend_label is a snapshot trend proxy; baseline models typically reach 25–35% accuracy without temporal features. publish_date_approx is derived and coarse — for trend direction only. If you find this dataset helpful, supporting it with an upvote helps others discover it too ✨
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
The TikHarm dataset is a curated collection of TikTok videos designed to train models for classifying harmful content. The dataset is in the format of UCF101, and it is specifically focused on content accessible to children, with the aim of distinguishing between different types of potentially harmful material.
Data was gathered from TikTok, targeting videos that are accessible to children to ensure the dataset reflects the type of content they are likely to encounter.
Collected videos were manually labeled into four predefined categories: - Harmful Content: Videos that depict violence, dangerous actions that children might imitate, or other harmful behavior. - Adult Content: Videos containing sexual content or other material deemed inappropriate for children. - Safe: Videos that are appropriate and safe for children to view: popular cartoon, etc. - Suicide: Videos that depict, suggest, or discuss suicidal behavior or ideation.
| Subset | Samples | Min Duration (s) | Max Duration (s) | Avg Duration (s) | Total Duration (h) |
|---|---|---|---|---|---|
| Train | 2762 | 3.88 | 600 | 38.71 | 29.71 |
| Dev | 396 | 1.95 | 600 | 38.77 | 8.51 |
| Test | 790 | 5.04 | 600 | 38.57 | 4.24 |
| Class | Samples | Min Duration (s) | Max Duration (s) | Avg Duration (s) | Total Duration (h) |
|---|---|---|---|---|---|
| Safe | 997 | 5.04 | 568.8 | 65.36 | 18.1 |
| Adult | 977 | 1.95 | 600 | 36.25 | 9.84 |
| Harmful | 990 | 4.8 | 600 | 35.92 | 9.88 |
| Suicide | 984 | 3.88 | 181.23 | 16.96 | 4.63 |
These tables present the duration statistics for each subset and class within the TikHarm dataset.
This comprehensive dataset is invaluable for developing robust video classification models to automatically detect and categorize harmful content on social media platforms.
Facebook
TwitterTikTok's platform is mostly fueled by viral videos of users doing outlandish, scary, or funny things. On the platform, these trend and meme videos typically come with a hashtag that includes the word challenge. But what is a TikTok challenge and how do you find or create them? Here's everything you need to know.
This TikTok book challenge was made by @haleyisfearless, . It asks you to show, your prettiest book,your tiniest book a book you highly suggest a book you're currently reading and one of your favorite books . In the most basic sense, these challenges originate from viral TikTok challenge isn't complete without its defining hashtag in the video's description
These TikTok challenges are the perfect way to ease into what can be an intimidating social media platform and help you find your fellow book lovers.
This dataset is generated entirely from TikTok , so we want to thank @haleyisfearless for building such this challange video
the goal of this project is to make Python script which takes a video as input and returns all texts visible on the video. the videos are titlok videos so texts can appear everywhere on screen, with different background, font size etc..
Facebook
TwitterI always notice how TikTok videos make viewers laugh - this led to me wonder: What exactly about TikTok made people happy? Is it the video length? Or is it the music / sounds? Or is it the content? With these questions in mind, I scraped all the videos from Top 5 influencers from 8 selected countries: Australia, Indonesia, Japan, Norway, Russia, Singapore, South Korea, US, and UK.
For every video (row), the information included are the variables : user_name, user_id, video_id, video_desc, video_time, video_length, video_link, n_likes, n_shares, n_comments, n_plays, video_timestamp, country, year.
user_name: user name of the user who posted the video
user_id: the id of the user recorded
video_id: the id of the video posted
video_desc: the description or caption of the video, written by the user
video_time: the time of posting of the video in UTC format
video_length: the length of the video in seconds
video_link: the url link to the video
n_likes: the number of likes received by the video
n_shares: the number of shares received by the video
n_plays: the number of plays recorded by the video
video_timestamp: the date of posting of the video, converted from UTC
country: the country of the user who posted the video
year: the year of posting the video
Huge thank you to the unofficial TikTok Api and its creator (@davidteather on Github) for letting this webscraping process become a lot more easy!
What sort of content in TikTok makes people happy?
Facebook
TwitterAs of January 2022, The United States was the country with the largest TikTok audience by far, with approximately 131 million users engaging with the popular social video platform. Indonesia followed, with around 92 million TikTok users. Brazil came in third, with 74 million users using TikTok to watch short-videos.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Please upvote if you like this dataset
TikTok, known in China as Douyin (Chinese: 抖音; pinyin: Dǒuyīn), is a short-form video hosting service owned by Chinese company ByteDance. It hosts a variety of short-form user videos, from genres like pranks, stunts, tricks, jokes, dance, and entertainment with durations from 15 seconds to ten minutes. TikTok is an international version of Douyin, which was originally released in the Chinese market in September 2016. TikTok was launched in 2017 for iOS and Android in most markets outside of mainland China; however, it became available worldwide only after merging with another Chinese social media service, Musical.ly, on 2 August 2018.
TikTok and Douyin have almost the same user interface but no access to each other's content. Their servers are each based in the market where the respective app is available. The two products are similar, but features are not identical. Douyin includes an in-video search feature that can search by people's faces for more videos of them and other features such as buying, booking hotels and making geo-tagged reviews. Since its launch in 2016, TikTok and Douyin rapidly gained popularity in virtually all parts of the world. TikTok surpassed 2 billion mobile downloads worldwide in October 2020.
In this dataset you will find the details about top 1000 tiktokers all over the world.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset is web-scraped from popular short video platforms like YouTube Shorts, TikTok, and Instagram Reels. It captures user interaction data, including views, likes, comments, shares, and watch duration, along with multimodal features from video content like text (titles, descriptions), image (visual characteristics), and audio (sound properties). The data has been processed and flattened into a structured CSV format with 17,654 Rows.
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
A multimodal dataset for training and distilling models that detect online financial scams in short-form video. Each example is a YouTube or TikTok video reduced to the signals a detector needs: sampled frames, a speech transcript, on-screen OCR text, and the title / description, paired with a scam / legitimate label and a policy-grounded rationale.
This is Version 1: the raw, as-collected release. Labels and rationales come from a mix of automated ("agentic") verification, companion-JSON sources, and partial human review, so their quality is uneven (see Known limitations). A cleaned, re-verified version is planned.
Three global scam types (crypto, giftcard, giveaway) plus three Filipino-specific archetypes (ewallet, p2e, task).
| Category | Scam | Legitimate | Total |
|---|---|---|---|
| crypto | 165 | 165 | 330 |
| ewallet | 150 | 150 | 300 |
| giftcard | 165 | 165 | 330 |
| giveaway | 150 | 150 | 300 |
| p2e (play-to-earn) | 190 | 190 | 380 |
| task | 180 | 180 | 360 |
| Total | 1,000 | 1,000 | 2,000 |
Every scam label is grounded in seven policy-based criteria (the OptiScam C1-C7 scheme, adapted from the YouTube-policy criteria in Kulsum et al.). The criteria a video matches are recorded in the companion JSON, in the response "Criteria Hits" line and the human_review annotation.
On top of the global scam types, the dataset adds three Filipino-specific archetypes:
Sources: publicly posted short-form videos scraped from YouTube and TikTok (browser automation with Playwright and Selenium), across the six scam archetypes above plus matched legitimate videos on the same topics, with a Philippine / Filipino focus.
Per-video pipeline:
1. Frame sampling - up to about 60 keyframes per video are extracted and saved as PNG under images/.
2. Speech to text - audio is transcribed with OpenAI Whisper and stored as the whisper turn (raw WAV is shipped for the giveaway category).
3. On-screen text - timestamped OCR is captured as ocr_temporal_text and the ocr turn.
4. Metadata - the video title and description are recorded as the human turn.
5. Labeling - teacher LLM labelers (Google Gemma and Google Gemini) assign a scam or legitimate verdict, a rationale, and the matched C1-C7 criteria, stored as the response turn and the human_review annotation.
Verification and balancing: labels come from a mix of automated / agentic verification, companion-JSON label sources, and partial human review; each video keeps a provenance rank (0 agentic-verified, 1 companion-json, 2 none) in its category .transfer_manifest.json. The released set is balanced to 1,000 scam and 1,000 legitimate. Only public content is used and no personal data is deliberately collected.
Each video lives in category / {scam | legitimate} / video_folder / and contains: a video_folder.json (metadata + label), an images/ folder of frames named video_folder_N.png, and, for giveaway in v1, a video_folder.wav audio file. The video_folder name is either youtube_ID or tiktok_ID. Each category also includes a .transfer_manifest.json recording label provenance.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains information about TikTok videos, including user interactions and video details. It includes features such as video ID, username, video title, likes, comments, shares, views, and more. This dataset is useful for analyzing video performance and user engagement on TikTok.
Columns: