
A 460 GB dataset containing metadata for billions of public TikTok videos has been posted to Hugging Face, bypassing official research restrictions.
AI-generated summary
TikTok restricts broad data scraping under Section 3.4 of its U.S. terms of service and limits its official Research Tools to approved academic institutions.
A developer who goes by hashfunction has posted metadata for about 5.6 billion public TikTok videos on the open-source AI repository Hugging Face, free to download. The dataset, published under the datasocial account, runs from July 2014 through October 2026.
It's metadata, not footage. Each row carries the caption, hashtags, sound ID, on-screen text, TikTok Shop product and seller IDs, and counts for views, likes, comments, shares, saves, and downloads. The whole thing is 460 GB of Parquet files, one per month.
Some fields are TikTok's own labels: whether a video is flagged as AI-generated, and whether TikTok keeps it off the For You page. The AI flag is often empty for older videos, which come from an archive.
The dataset isn't gated, so anyone can pull it. Hugging Face's counter shows 1,181 downloads so far. Itâs also available online on https://datasocial.ai/
Why should anyone care? Because AI developers want this kind of data, and TikTok has made it hard to get. Captions, hashtags, sound IDs and engagement counts for billions of posts are raw material for models that predict what goes viral or what sells on TikTok Shop, what words and phrases click and which donât, and for language models learning how people write in short-form video.
TikTok's official route to the data is much narrower. Its Research Tools are open to qualifying researchers at academic institutions in regions including the United States, EEA, UK, Canada, and Switzerland, plus some not-for-profit bodies in the EU, and require an application and approval.
Broad data scraping isn't allowed by the platform. Section 3.4 of TikTok's U.S. terms of service bars extracting data from the platform with automated software unless TikTok approves it in writing.
DataSocial's write-up describes reaching TikTok's private mobile API with generated device identities that pass as Android phones, reverse-engineered request signatures, and a spoofed TLS handshake. It says the system collected 3.23 billion creator profiles, 5.94 billion videos, and 2.8 billion comments in three weeks. No login or account was involved, according to the write-up.
The free tier doubles as a storefront. The license is CC BY-NC 4.0, and the card says it's free for research and personal use. Commercial use, creator profiles, and daily updates are routed to datasocial.ai, and the scraper's source code is sold separately.
Scraping fights are already in court. Reddit sued Perplexity and three data-scraping firms in October 2025, alleging an "industrial-scale" scheme to harvest its content for AI training. In July, a federal judge largely declined to dismiss the case, per Law360.
The surviving claims include DMCA allegations that SerpApi circumvented Google's anti-bot protections. TikTok isn't a party, and the case involves scraping through Google search results, not a private mobile API.
As for the data itself, rows include captions, music used, and different checkboxes. There is no creator handle column, but it shows a lot of data and information including captions, hashtags, people mentioned on the video, etc. That said, the user ID is actually available in Datasocial.
The data trove already has company. Copies of a 4.5-billion-video set, also described as collected from TikTok's mobile API, are on Hugging Face too.
The data is free for non-commercial use. The code that collected it costs $1,699.
AI outlook â possibilities, not facts
TikTok may issue removal requests or take legal measures regarding the scraped data.
Likely · Within weeks

Google launched Nano Banana 2.1, an updated AI model that generates and edits images from text, now available in Gemini app, Search AI Mode, Ads and developer tools. The model offers better visual design, mask-based editing and subject consistency, scores 1,050 ELO in preference tests, and reduces API costs to $0.0336 per 1K image, half the price of its predecessor.

Mistral AI launched Mistral Large 4, a 1-trillion-parameter AI model using a mixture-of-experts design with 49 billion active parameters per query. The model is positioned as a competitive open-weight alternative to Claude Opus 5.5 and GPT-6 Astra, with pricing at $1.36 per million input tokens and $4.18 per million output tokens. Benchmarks show strong performance in coding and automation tasks, trailing only Kimi K3 and Gemini 4 Argon in some tests. Mistral plans to release the model weights by end of October.

Cardano's CIP-113 proposal for programmable tokens adds compliance controls but risks temporarily blocking unrelated assets and ADA that share the same transaction output.

The Solana Foundation announced Solana DvP, an open-source settlement program designed to cut institutional securities settlement times from days to seconds using blockchain infrastructure.

The Ethereum Economic Zone has successfully tested its first atomic cross-chain transaction between the Ethereum mainnet and a layer-2 network, aiming to eliminate the need for traditional bridges.

The Solana Foundation has launched Solana DvP, an open-source escrow program enabling banks to settle trades on the Solana blockchain with atomic delivery-versus-payment finality in seconds. Developed with input from J.P. Morgan and released under the MIT license, the tool supports SPL Token and Token-2022 standards and aims to replace bespoke smart contracts with a reusable infrastructure for institutional finance.