A researcher uploaded 4.5 billion TikTok video records to Hugging Face
The 289GB dataset was scraped through a private Android API. It breaches TikTok's terms, and nothing technical stopped it being published.
An independent researcher has uploaded a dataset of 4.5 billion TikTok video records — 289GB — to Hugging Face, scraped via a private Android API. The upload breaches TikTok's terms of service, and the dataset carries stated restrictions on its use that nothing enforces.
The episode is a clean illustration of how little stands between a platform's data and a public training corpus.
Terms of service are not a control
TikTok's terms prohibit scraping. The private Android API was not intended for third-party use. The dataset's own licence restricts what downloaders may do with it.
All three are statements about what people should not do, and none is a mechanism preventing it. The data was collected, published, and is now distributed by whoever fetched it before any takedown — which is the property that makes dataset publication different from most terms-of-service breaches. Once a 289GB file has been mirrored, removing the original changes nothing.
The restrictions on the dataset are the least enforceable part. A researcher who publishes 4.5 billion records with a note asking people not to use them commercially has expressed a preference.
What the records are worth
Metadata at this scale is more revealing than it sounds. Video records covering billions of items describe what was posted, when, by what kind of account, and how it performed — which is the raw material for studying how content spreads on the largest recommendation system in the world.
That is genuinely valuable for research nobody can otherwise do. Platform recommendation dynamics are studied almost entirely through the platforms' own disclosures or through small scraped samples, and independent measurement of what actually propagates is close to impossible.
The Institute for Strategic Dialogue's finding this week — 150 AI-remixed extremist videos across 71 TikTok accounts with 5.4 million views — came from targeted discovery precisely because platform-wide data is unavailable. The researchers said so explicitly.
So the dataset serves a real need that exists because the legitimate route does not work.
The legitimate route exists on paper
The Digital Services Act requires very large online platforms to give vetted researchers access to data for studying systemic risks. That provision was designed for exactly this: independent measurement without scraping.
It has been slow, contested and narrow in practice. Platforms have been restrictive about who qualifies and what they receive, and the process takes months. A researcher who wants to study recommendation dynamics this year has a choice between a formal process that may not deliver and an Android API that will.
That is the policy failure this dataset represents. Article 40 access working properly would remove most of the justification for scraping at this scale.
The awkward position for the host
Hugging Face is where the dataset sits, and Hugging Face is being acquired by Nvidia for $12.9 billion, with the deal expected to close in the first half of 2027.
A platform hosting a scraped dataset from a competitor's service is an ordinary content-moderation problem. A subsidiary of Nvidia hosting one is a different exposure, with a corporate parent that has substantially more to lose from litigation and considerably more regulatory attention on it.
Nvidia has committed to keeping the hub open to competing models and chips. It has said nothing about content policy, and the dataset question is the first place that will matter.
TikTok has not commented, and the dataset remained available at the time of reporting.
Runs the newsroom. Rename this profile in the studio to your own byline.
Related
Every weekday, the AI stories that moved money or shipped code.
No cross-posting, unsubscribe anytime. See all newsletters