Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets
Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.
187 million files totaling 7.6 petabytes of public AI training data on Hugging Face contained 221,303 live credentials, including 349 GitHub tokens with repo write access that can compromise software supply chains, enabling attackers to push malicious code to millions of users. This exposes a critical vulnerability for teams shipping AI models and agents that rely on this data. It necessitates immediate secret scanning and revocation for anyone using Hugging Face datasets.
A scan of 7.6 PB of public Hugging Face datasets found 221,303 live unique credentials across 6,003 datasets, including write-capable GitHub and Docker Hub tokens with supply-chain impact. If you train, fine-tune, RAG-index, or let agents operate over public datasets, treat the data plane as an active secret-ingestion path: add secret scanning and revocation workflows before ingestion, and assume exposed tokens can become deployable code or infrastructure compromise, not just benign training noise.