Agents & InferenceHacker News

Scanning 7.6 Petabytes of HuggingFace Training Data for Secrets

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

187 million files totaling 7.6 petabytes of public AI training data on Hugging Face contained 221,303 live credentials, including 349 GitHub tokens with repo write access that can compromise software supply chains, enabling attackers to push malicious code to millions of users. This exposes a critical vulnerability for teams shipping AI models and agents that rely on this data. It necessitates immediate secret scanning and revocation for anyone using Hugging Face datasets.