Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Find the top TikTok accounts.
What's inside is more than just rows and columns. Make it easy for others to get started by describing how you acquired the data and what time period it represents, too.
Data source: https://hypeauditor.com/top-tiktok/
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset captures the pulse of viral social media trends across TikTok, Instagram, Twitter, and YouTube. It provides insights into the most popular hashtags, content types, and user engagement levels, offering a comprehensive view of how trends unfold across platforms. With regional data and influencer-driven content, this dataset is perfect for:
Dive in to explore what makes content go viral, the behaviors that drive engagement, and how trends evolve on a global scale! 🌍
Facebook
Twitterhttps://brightdata.com/licensehttps://brightdata.com/license
Use our TikTok profiles dataset to extract business and non-business information from complete public profiles and filter by account name, followers, create date, or engagement score. You may purchase the entire dataset or a customized subset depending on your needs. Popular use cases include sentiment analysis, brand monitoring, influencer marketing, and more. The TikTok dataset includes all major data points: timestamp, account name, nickname, bio,average engagement score, creation date, is_verified,l ikes, followers, external link in bio, and more. Get your TikTok dataset today!
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Aggregate statistics over 6,133,853 established TikTok creators: the follower-tier pyramid (median ~12,000 followers), discovery reach per search keyword, verification rate (1.4%), and public business-contact rate by tier — for creator discovery and influencer-marketing research.
Facebook
TwitterThe TikTok Creator Profiles Dataset provides access to millions of publicly available TikTok creator profiles across industries, audience sizes, and regions worldwide.
Designed for marketers, agencies, researchers, and analytics teams, the dataset supports influencer discovery, market research, competitive analysis, audience intelligence, and creator economy insights.
Each profile may include publicly available information such as username, display name, bio, profile URL, follower count, following count, total likes, video count, verified status, category, country, language, external links, and contact details where available.
The dataset covers creators across major categories including lifestyle, beauty, fashion, gaming, fitness, technology, entertainment, travel, food, and more.
Data is available in CSV, JSON, and API formats, with regular updates to ensure fresh and reliable coverage of the global TikTok creator ecosystem.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset tracks influencer marketing campaigns across major social media platforms, providing a robust foundation for analyzing campaign effectiveness, engagement, reach, and sales outcomes. Each record represents a unique campaign and includes details such as the campaign’s platform (Instagram, YouTube, TikTok, Twitter), influencer category (e.g., Fashion, Tech, Fitness), campaign type (Product Launch, Brand Awareness, Giveaway, etc.), start and end dates, total user engagements, estimated reach, product sales, and campaign duration. The dataset structure supports diverse analyses, including ROI calculation, campaign benchmarking, and influencer performance comparison.
Columns:
- campaign_id: Unique identifier for each campaign
- platform: Social media platform where the campaign ran
- influencer_category: Niche or industry focus of the influencer
- campaign_type: Objective or style of the campaign
- start_date, end_date: Campaign time frame
- engagements: Total user interactions (likes, comments, shares, etc.)
- estimated_reach: Estimated number of unique users exposed to the campaign
- product_sales: Number of products sold as a result of the campaign
- campaign_duration_days: Duration of the campaign in days
import pandas as pd
df = pd.read_csv('influencer_marketing_roi_dataset.csv', parse_dates=['start_date', 'end_date'])
print(df.head())
print(df.info())
# Overview of campaign types and platforms
print(df['campaign_type'].value_counts())
print(df['platform'].value_counts())
# Summary statistics
print(df[['engagements', 'estimated_reach', 'product_sales']].describe())
# Average engagements and sales by platform
platform_stats = df.groupby('platform')[['engagements', 'product_sales']].mean()
print(platform_stats)
# Top influencer categories by product sales
top_categories = df.groupby('influencer_category')['product_sales'].sum().sort_values(ascending=False)
print(top_categories)
# Assume a fixed campaign cost for demonstration
df['campaign_cost'] = 500 + df['estimated_reach'] * 0.01 # Example formula
# Calculate ROI: (Revenue - Cost) / Cost
# Assume each product sold yields $40 revenue
df['revenue'] = df['product_sales'] * 40
df['roi'] = (df['revenue'] - df['campaign_cost']) / df['campaign_cost']
# View campaigns with highest ROI
top_roi = df.sort_values('roi', ascending=False).head(10)
print(top_roi[['campaign_id', 'platform', 'roi']])
import matplotlib.pyplot as plt
import seaborn as sns
# Engagements vs. Product Sales scatter plot
plt.figure(figsize=(8,6))
sns.scatterplot(data=df, x='engagements', y='product_sales', hue='platform', alpha=0.6)
plt.title('Engagements vs. Product Sales by Platform')
plt.xlabel('Engagements')
plt.ylabel('Product Sales')
plt.legend()
plt.show()
# Average ROI by Influencer Category
category_roi = df.groupby('influencer_category')['roi'].mean().sort_values()
category_roi.plot(kind='barh', color='teal')
plt.title('Average ROI by Influencer Category')
plt.xlabel('Average ROI')
plt.show()
# Campaigns over time
df['month'] = df['start_date'].dt.to_period('M')
monthly_sales = df.groupby('month')['product_sales'].sum()
monthly_sales.plot(figsize=(10,4), marker='o', title='Monthly Product Sales from Influencer Campaigns')
plt.ylabel('Product Sales')
plt.show()
Facebook
TwitterStructured dataset of 30M+ creator profiles across YouTube, Instagram and TikTok, with up to 245 data points per creator including real engagement rate (computed against active followers), audience demographics (country, age, gender), sponsorship history with cost-per-video estimates, content-niche classification, verified contact emails and up to 4+ years of follower and engagement time-series. Refreshed almost daily.
Facebook
TwitterAbout This Dataset This dataset is a sample from WebAutomation's TikTok Profiles Dataset, containing publicly available information from TikTok creator and business profiles. The sample includes approximately 100M TikTok profiles and provides a diverse snapshot of creators, influencers, brands, and organizations across multiple industries, languages, and regions. It is designed for researchers, marketers, data scientists, creator economy platforms, and businesses seeking insights into TikTok's rapidly growing ecosystem.
What's Included
Each row represents a single TikTok profile and may include the following fields: username — TikTok username tiktok_id — TikTok account identifier sec_uid — TikTok secure unique identifier display_name — Public display name bio — Profile biography external_url — Website or external link listed in the profile avatar_url — Profile image URL language — Primary detected profile language followers_count — Number of followers following_count — Number of accounts followed likes_count — Total profile likes received video_count — Number of published videos friends_count — Number of mutual connections (where available) digg_count — Number of liked videos is_verified — Verification status is_private — Indicates whether the account is private is_organization — Organization account indicator is_commerce_user — Commerce-enabled account indicator is_tt_seller — TikTok Shop seller indicator has_active_live_room — Indicates active livestream availability open_favorite — Public favorites visibility setting comment_setting — Comment permissions setting duet_setting — Duet permissions setting stitch_setting — Stitch permissions setting download_setting — Video download permissions setting created_at — Account creation date (where available) nickname_modified_at — Most recent nickname change timestamp username_modified_at — Most recent username change timestamp scraped_at — Date and time the data was collected
Additional fields may be included depending on the data collection period and TikTok profile attributes.
Data Collection
Data is collected from publicly accessible TikTok profile pages using WebAutomation's proprietary web data infrastructure. No private APIs, account credentials, or authentication-restricted information are used. All collected information is publicly available at the time of collection.
Scrape date of this sample: 6/7/2026
Use Cases Influencer discovery and creator segmentation Creator economy research Social media analytics Audience and engagement analysis Brand partnership identification Market intelligence and trend analysis Machine learning and recommendation systems Academic and industry research
Data Quality The sample dataset has undergone basic quality checks, including: Removal of duplicate profiles Validation of unique account identifiers Verification of scrape timestamps Standardization of profile-level metrics Filtering of incomplete or invalid records
Full Dataset
This is a sample only.
WebAutomation maintains significantly larger TikTok datasets covering creator profiles, engagement metrics, account attributes, business accounts, and public profile metadata at scale.
👉**Explore more datasets:** https://webautomation.io
For custom data delivery, API access, enrichment projects, or enterprise licensing, contact: victor@webautomation.io
About WebAutomation
WebAutomation.io is a web data and scraping company specializing in large-scale public datasets across social media platforms, marketplaces, review sites, and business directories. We provide structured datasets and data collection solutions for market research firms, analytics platforms, AI companies, enterprise teams, and data-driven organizations worldwide.
Facebook
TwitterThis dataset contains metadata from online news coverage of the arrest and prosecution of female TikTok influencers in Egypt, a case labelled by the media as "the TikTok Girls" case. The data were collected from five major Egyptian and Arabic-language news websites between 2020 and 2024. It includes article-level metadata (such as titles, publication dates, authors, and URLs). The dataset was compiled to examine how these prosecutions and the women involved have been represented and debated in digital news media during this period.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Dataset The Role of Digital User Experience in TikTok Influencer Content on Generation Z Purchase Intention in Indonesia
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
The dataset consists of primary data collected through an online questionnaire distributed to followers of a selected TikTok influencer in Indonesia. Respondents were selected using purposive sampling based on predefined criteria. The data were measured using a Likert scale and include variables of influencer credibility, purchase intention, and purchase decision. The dataset was analyzed using SEM-PLS to examine the relationships between variables.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
This social media content dataset is simulate realistic influencer posts across multiple popular platforms, reflecting diverse content types, sponsorship details, audience demographics, and engagement metrics. The dataset contains over 52,000 rows representing individual content posts generated over the past two years. It includes a balanced distribution of sponsored and non-sponsored content, with detailed disclosure information to support transparency studies and analyses. The variety of platforms, languages, content categories, and audience demographics makes this dataset ideal for exploring influencer marketing dynamics, content performance analytics, disclosure practices, and audience segmentation in social media research.
Dataset Features
id: Unique identifier for each content post (starting from 1).
platform: The social media platform where the content was posted. Values: YouTube, TikTok, Instagram, Bilibili, RedNote.
content_id: Unique ID for each content piece (e.g., content_0, content_1, …).
creator_id: Unique identifier for the content creator, cycling through 5000 distinct creators.
creator_name: Username of the content creator.
content_url: URL pointing to the content.
content_type: Format of the content. Values: video, image, text, mixed.
content_category: The main theme or niche of the content. Values: beauty, lifestyle, tech.
post_date: Timestamp of the post, randomly distributed over the past two years.
language: Language of the content, with probabilities favoring English. Values: English, Chinese, Spanish, Hindi, Japanese.
content_length: Length of the content in seconds (for video) or word count (for text), varying by content type.
content_description: Textual description or caption of the content.
hashtags: A comma-separated string of hashtags used in the post (0 to 5 tags).
views: Number of views (simulated via a Poisson distribution).
likes: Number of likes received.
shares: Number of shares.
comments_count: Count of comments on the post.
comments_text: Aggregated text of comments (0 to 5 comments concatenated).
follower_count: Number of followers the creator had at the time of posting.
is_sponsored: Boolean indicating whether the post is sponsored.
disclosure_type: Disclosure type regarding sponsorship for sponsored posts. Values: explicit, implicit, none (non-sponsored always 'none').
sponsor_name: Name of the sponsoring company if sponsored, else 'Not sponsors'.
sponsor_category: Sponsorship industry category. Values: cosmetics, electronics, fashion, food, gaming, travel or 'Not sponsors'.
disclosure_location: Where sponsorship disclosure appears in the post. Values: video, caption, hashtags, none (non-sponsored always 'none').
audience_age_distribution: Predominant age group of the audience. Values: 13-18, 19-25, 26-35, 36-50, 50+.
audience_gender_distribution: Predominant gender of the audience. Values: male, female, non-binary, unknown.
audience_location: Primary geographic location of the audience. Values: USA, China, India, Japan, Brazil, Germany, UK, Russia.
Facebook
Twitterhttps://brightdata.com/licensehttps://brightdata.com/license
Gain a competitive edge with our comprehensive Advertising Dataset, designed for marketers, analysts, and businesses to track ad performance, analyze competitor strategies, and optimize campaign effectiveness. Dataset Features Sponsored Posts & Ads: Access structured data on paid advertisements, including post content, engagement metrics, and platform details. Competitor Advertising Insights: Extract data on competitor campaigns, influencer partnerships, and promotional strategies. Audience Engagement Metrics: Analyze likes, shares, comments, and impressions to measure ad effectiveness. Multi-Platform Coverage: Track ads across LinkedIn, Instagram, Facebook, TikTok, Twitter (X), Pinterest, and more. Historical & Data: Retrieve historical ad performance data or access regularly updated records for insights. Customizable Subsets for Specific Needs Our Advertising Dataset is fully customizable, allowing you to filter data based on platform, ad type, engagement levels, or specific brands. Whether you need broad coverage for market research or focused data for ad optimization, we tailor the dataset to your needs. Popular Use Cases Targeted Advertising & Audience Segmentation: Refine ad targeting by analyzing competitor content, audience demographics, and engagement trends. Campaign Performance Analysis: Measure ad effectiveness by tracking engagement metrics, reach, and conversion rates. Competitive Intelligence: Monitor competitor ad strategies, influencer collaborations, and promotional trends. Market Research & Trend Forecasting: Identify emerging advertising trends, high-performing content types, and consumer preferences. AI & Predictive Analytics: Use structured ad data to train AI models for automated ad optimization, sentiment analysis, and performance forecasting. Whether you're optimizing ad campaigns, analyzing competitor strategies, or refining audience targeting, our Advertising Dataset provides the structured data you need. Get started today and customize your dataset to fit your business objectives.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Please upvote if you like this dataset
TikTok, known in China as Douyin (Chinese: 抖音; pinyin: Dǒuyīn), is a short-form video hosting service owned by Chinese company ByteDance. It hosts a variety of short-form user videos, from genres like pranks, stunts, tricks, jokes, dance, and entertainment with durations from 15 seconds to ten minutes. TikTok is an international version of Douyin, which was originally released in the Chinese market in September 2016. TikTok was launched in 2017 for iOS and Android in most markets outside of mainland China; however, it became available worldwide only after merging with another Chinese social media service, Musical.ly, on 2 August 2018.
TikTok and Douyin have almost the same user interface but no access to each other's content. Their servers are each based in the market where the respective app is available. The two products are similar, but features are not identical. Douyin includes an in-video search feature that can search by people's faces for more videos of them and other features such as buying, booking hotels and making geo-tagged reviews. Since its launch in 2016, TikTok and Douyin rapidly gained popularity in virtually all parts of the world. TikTok surpassed 2 billion mobile downloads worldwide in October 2020.
In this dataset you will find the details about top 1000 tiktokers all over the world.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Influencer marketing statistics dataset for 2026: 11 cited data points across 5 named sources (Statista, Sprout Social, Aspire, CreatorIQ, Influencer Marketing Hub) plus a FORKOFF first-party network figure of 5B+ views processed. Covers global market size and growth, marketer adoption, budget intent, AI adoption in influencer operations, TikTok Shop and social commerce, ROI and program validation under economic pressure, and the influencer functions most commonly outsourced to agencies.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
🚀TikTok Shop Viral Video Sales, Attribution & CAC Analytics Dataset
📊 100,000 Transactions • 180+ Features • Creator Economics • Multi-Touch Attribution • CAC Optimization Traditional e-commerce analytics models are failing. In modern social commerce ecosystems like TikTok Shop, virality, creator authority, and micro-retention curves dictate revenue—not standard search traffic.
This dataset provides a production-grade analytical environment containing 100,000 granular observations and 180+ features. It bridges the gap between unstructured social media video engagement (hooks, rewatches, drop-offs) and enterprise financial performance (ROAS, CPA, Net Profit Margins, Multi-Touch Attribution).
💡 Why This Dataset Beats Standard Datasets End-to-End Customer Path: Tracks the complete journey from video impression and 3-second hook to product conversion and return risk.
Modern Social Commerce Vectors: Includes influencer tiering, affiliate tracking, and viral video hook dynamics that traditional retail datasets ignore.
Advanced Financial Metrics: Measures true net margins after creator commissions, platform fees, shipping costs, and customer acquisition expenses.
Multi-Touch Attribution Engine: Features First-Touch, Last-Touch, Linear, Position-Based, and Data-Driven revenue attribution models.
🧬 Dataset Schema & Feature Vectors (180+ Features) The data is structured into 6 interconnected analytical vectors:
1. 🎬 Video Performance & Content Hook Dynamics Video runtime, unique impressions, likes, shares, saves, retention curves at 3s and 10s, exact drop-off second, rewatch ratio, viral score, and hook clarity ratings.
2. 👤 Creator & Affiliate Intelligence Creator tier (Micro, Mid, Mega), follower count, engagement velocity, historical conversion authority, affiliate commission tier, and influencer ROI impact.
3. 💰 Financials & Unit Economics Gross revenue, net profit margin, Cost Per Acquisition (CPA), Customer Acquisition Cost (CAC), Return on Ad Spend (ROAS), platform fees, and marketing ROI.
4. 📈 Multi-Touch Attribution Models First-click revenue, last-click revenue, linear attribution, time-decay attribution, position-based model, and machine learning attribution confidence scores.
5. 🛒 Funnel & Traffic Analytics Product page clicks, add-to-cart rate, checkout abandonment rate, profile visits, traffic origin, and scroll depth.
6. 📦 Operations, Logistics & Profitability Shipping duration, delivery failure rate, return reason codes, processing cost, refund impact, and net contribution margin.
🎯 Machine Learning & Predictive Modeling Targets Use this dataset to build models targeting real-world business challenges:
Is_Viral (Binary Classification): Predict whether a video will achieve viral reach based on early 3-second hook metrics.
Net_Profit_Margin (Regression): Predict true profitability after subtracting creator commissions, return rates, and CAC.
Attribution_Model_Winner (Multi-Class Classification): Classify which attribution model best reflects true conversion impact based on customer path complexity.
Customer_Return_Probability (Predictive Modeling): Predict item return risk using product category, video claims, and shipping delays.
🚀 Quick-Start Analysis Snippet (Python)
import pandas as pd import numpy as np
df = pd.read_csv("tiktok_shop_viral_analytics_100k.csv")
df['CAC_Efficiency'] = df['Net_Revenue'] / (df['CAC'] + 1e-5)
creator_perf = df.groupby('Creator_Tier')[['ROAS', 'Net_Profit_Margin', 'CAC']].mean() print("=== Creator Economics Overview ===") print(creator_perf)
df['Attribution_Delta'] = df['Data_Driven_Revenue'] - df['Last_Click_Revenue'] print(" === Revenue Misattributed by Last-Click Model ===") print(df['Attribution_Delta'].describe())
🛠️ Recommended Portfolio Projects Project 1: Executive Creator Economics Dashboard (Power BI / Tableau) Visualize ROAS vs. Creator Commission rates across 100K transactions.
Project 2: Multi-Touch Attribution Engine (Python / SQL) Compare First-Touch vs. Data-Driven Attribution to reallocate marketing spend.
Project 3: Viral Hook & Drop-off Predictor (XGBoost / LightGBM) Predict video conversion failure using early 3-second viewer retention signals.
📌 Community Engagement & License License: CC0: Public Domain (Free for personal, academic, and commercial portfolio use).
Support: If you find this dataset valuable for your research, portfolio, or machine learning experiments, please ⭐ UPVOTE to support future dataset releases!
Showcase:Feel free to share your EDA notebooks and dashboard links in the Code and Discussion tabs!
Facebook
Twitterhttps://aussda.at/en/aussda-scientific-use-licence-for-data-and-cc-by-for-documentationhttps://aussda.at/en/aussda-scientific-use-licence-for-data-and-cc-by-for-documentation
Full edition for scientific use. The present dataset documents the use of the platform TikTok by Austrian politicians and political parties at the federal, state, and European levels. It includes all active TikTok accounts of members of state governments, the National Council, and the European Parliament, as well as those of state and federal political parties. For each account, the dataset provides reach and activity metrics. In addition, the ten most successful videos per account were categorized using a standardized content analysis conducted by a trained coding team. The dataset offers a systematic overview of reach, thematic emphases, and communication patterns of political actors on TikTok and provides an empirical basis for further research on political communication on visually oriented social media platforms.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
📱 About Dataset Overview This Social Media Engagement Dataset contains comprehensive engagement metrics from 5,000 social media posts across six major platforms: Instagram, Twitter, Facebook, LinkedIn, TikTok, and YouTube. The dataset spans over 2 years (2024-2025) and provides valuable insights into content performance, audience engagement patterns, and influencer analytics.
Dataset Contents The dataset includes 20 detailed features covering various aspects of social media engagement:
Post Information Post_ID: Unique identifier for each post Timestamp: Date and time when the post was published Platform: Social media platform (Instagram, Twitter, Facebook, LinkedIn, TikTok, YouTube) Content_Type: Type of content (Photo, Video, Reel, Tweet, Story, etc.) Category: Content category (Technology, Fashion, Food, Travel, Fitness, Education, Entertainment, Business, Lifestyle, Gaming, Health, Sports) Engagement Metrics Likes: Number of likes/reactions received Comments: Number of comments on the post Shares: Number of shares/retweets/reposts Views: Total number of views Saves: Number of bookmarks/saves Engagement_Rate: Calculated engagement rate percentage Account Information Follower_Count: Number of followers of the account Influencer_Tier: Classification (Nano, Micro, Mid-tier, Macro) Is_Verified: Whether the account is verified (True/False) Content Characteristics Hashtag_Count: Number of hashtags used Content_Length: Length in characters (text) or seconds (video) Sentiment: Sentiment analysis (Positive, Neutral, Negative) Has_Media: Whether post contains media (True/False) Temporal Features Hour_of_Day: Hour when the post was published (0-23) Day_of_Week: Day of the week (Monday-Sunday) Use Cases This dataset is perfect for:
📊 Predictive Analytics: Build ML models to predict engagement rates 📈 Data Visualization: Create insightful dashboards and charts 🤖 Machine Learning: Classification, regression, and clustering tasks ⏰ Time Series Analysis: Analyze posting patterns and optimal timing 🎯 Content Strategy: Optimize content strategy based on data insights 🔍 Sentiment Analysis: Study correlation between sentiment and engagement 📱 Platform Comparison: Compare performance across different platforms 💼 Influencer Marketing: Analyze influencer tier performance Technical Details Format: CSV Size: ~651 KB Rows: 5,000 Columns: 20 Time Period: January 2024 - December 2025 Missing Values: None Potential Research Questions What time of day generates the most engagement? Which platform has the highest engagement rates? How does content type affect performance? Does verified status impact engagement? What's the optimal hashtag count? How does sentiment correlate with engagement? Notes Engagement metrics are platform-realistic and proportional All data is synthetically generated for educational and research purposes Suitable for beginners and advanced data scientists
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This is a dataset accompanying the paper “The DSA's Blind Spot: Algorithmic Audit of Advertising and Minor Profiling on TikTok” presented at the FAccT 2026 conference, designed to analyze video interactions, ad classifications, and user engagement patterns. It contains records of video interactions, including metadata about the videos, user demographics, and ad classifications, allowing the full replication of results presented in the paper.
The video excerpts included in this dataset are used solely as units of content for analytical purposes. They do not represent, reflect, or imply the personal views, intentions, or stance of the individuals who created them. Content should be interpreted as data artifacts, not as statements attributable to any person.
To minimize the risk of third-party misuse, the dataset is available only to researchers for non-commercial research purposes upon verification of their email address associated with academic organisation.
Paper: https://dl.acm.org/doi/10.1145/3805689.3812355
Preprint: https://arxiv.org/abs/2603.05653
GitHub repository: https://github.com/kinit-sk/ai-auditology-advertising-and-minor-profiling-tiktok
Acknowledgemet: This work was partially funded by the EU NextGenerationEU through the Recovery and Resilience Plan forSlovakia under the project AI-Auditology, No. 09I03-03-V03-00020.
If you use this dataset in any publication, project, tool or in any other form, please, cite the following paper:
@inproceedings{10.1145/3805689.3812355,
author = {Solarova, Sara and Mosnar, Matej and Tibensky, Matus and Jakubcik, Jan and Bindas, Adrian and Liska, Simon and Hossner, Filip and Mesar\v{c}\'{\i}k, Mat\'{u}\v{s} and Srba, Ivan},
title = {The DSA's Blind Spot: Algorithmic Audit of Advertising and Minor Profiling on TikTok},
year = {2026},
isbn = {9798400725968},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3805689.3812355},
doi = {10.1145/3805689.3812355},
abstract = {Adolescents spend an increasing amount of their time in digital environments where their still-developing cognitive capacities leave them unable to recognize or resist commercial persuasion. Article 28(2) of the Digital Service Act (DSA) responds to this vulnerability by prohibiting profiling-based advertising to minors. However, the regulation's narrow definition of “advertisement” excludes current advertising practices including influencer paid partnerships and brand promotional content that serve functionally equivalent commercial purposes. We provide the first empirical evidence of how this definitional gap operates in practice through an algorithmic audit of TikTok. Our approach deploys sock-puppet accounts simulating a pair of minor and adult users with matching interest profiles. The content recommended to these users is automatically annotated, enabling systematic statistical analysis across four video categories: containing formal, disclosed, undisclosed advertisement and non-advertisement; as well as advertisement topical relevance to user's interest. Our findings reveal a stark regulatory paradox. TikTok demonstrates formal compliance with Article 28(2) by shielding minors from profiled formal advertisements, yet both disclosed and undisclosed ads exhibit significant profiling aligned with user interests (5-8 times stronger than for adult formal advertising). The strongest profiling emerges within undisclosed commercial content, where creators/brands fail to label paid partnership/promotional content and the platform neither corrects this omission nor prevents its personalized delivery to minors. These results demonstrate that minors remain exposed to algorithmically targeted commercial content through the same recommendation mechanisms the DSA seeks to constrain. We argue that protecting minors requires expanding the definition of advertisement in EU law to encompass influencer and brand promotional content, and ensuring that any such expansion is accompanied by a corresponding prohibition on profiling-based targeting of minors, so that commercial content cannot circumvent protections merely by operating outside formal advertising channels.},
booktitle = {Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency},
pages = {4811–4835},
numpages = {25},
keywords = {Digital Services Act, advertisement, algorithmic auditing, minor profiling, TikTok},
location = {},
series = {FAccT '26}
}
The logs of video presented to individual simulated users are provided in the ai-auditology-advertising-and-minor-profiling-tiktok_video_data.csv file. It is structured into 31 columns, capturing details such as session and video identifiers, timestamps, ad classifications, visual indicators, user demographics, and video metadata.
|
Column Name |
Data Type |
Description |
Example Value |
|
session_id |
string |
Session identifier captured during browsing |
1765302414.743265 |
|
video_id |
string |
Platform video identifier |
[anonymized] |
|
timestamp |
datetime |
Timestamp when the record was captured |
2025-12-09T17:47:56.296448 |
|
is_ad |
boolean |
Whether the video was classified as an ad |
false |
|
ad_type |
string (nullable) |
Ad classification type when is_ad is true |
other |
|
ad_topic |
string (nullable) |
Detected topic for ad content |
beauty |
|
visual_indicators |
array[string] |
List of visual indicators used to classify ads |
["hashtag #clearskin"] |
|
reasoning |
string |
Model reasoning for the ad classification |
No disclosure label visible. |
|
interaction_number |
integer |
Sequential interaction count within the session |
1 |
|
search_term |
string |
Search term used to find the content |
clear skin |
|
video_action_skip |
boolean |
Whether the user skipped the video |
False |
|
video_action_watch |
boolean |
Whether the user watched the video |
True |
|
video_action_like |
boolean |
Whether the user liked the video |
True |
|
video_action_bookmark |
boolean |
Whether the user bookmarked the video |
True |
|
video_time_watch_loop_start |
float (nullable) |
Timestamp when watch loop started |
1765302470.8245792 |
|
video_time_watch_loop_end |
float (nullable) |
Timestamp when watch loop ended |
1765302477.842666 |
|
video_time_skip |
float (nullable) |
Timestamp when the video was skipped |
nan |
|
video_time_like |
float (nullable) |
Timestamp when the video was liked |
1765302471.8269806 |
|
video_time_bookmark |
float (nullable) |
Timestamp when the video was bookmarked |
1765302477.3054323 |
|
video_time_predict_interaction |
float (nullable) |
Timestamp for predicted interaction (if any) |
nan |
|
topic |
string |
User interest topic used for personalization |
beauty |
|
gender |
string |
User gender |
female |
|
country_code |
string |
User country code |
DE |
|
date_of_birth |
date |
User date of birth |
2009-11-29 |
|
agent |
string |
Agent identifier added during processing |
Beauty_minor |
|
video_url |
string |
Full URL to the video | |
|
video_author |
string |
Account handle of the video author |
[anonymized] |
|
video_description |
string |
Video description text |
little bonus - your waist? |
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
The dataset provides structured information about the top 100 influencers from various countries globally. Each entry represents an influencer and includes the following attributes:
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Find the top TikTok accounts.
What's inside is more than just rows and columns. Make it easy for others to get started by describing how you acquired the data and what time period it represents, too.
Data source: https://hypeauditor.com/top-tiktok/