Is My Place in The Stack? - a Hugging Face Space by HuggingFaceCode
huggingface
| Source: Mastodon | Original article
Hugging Face launches The Stack V3, a massive 15.9 TB dataset of source code.
Hugging Face has launched a new space, "Am I in The Stack?", allowing developers to check if their open-source code is included in The Stack v3, a massive 15.9 TB dataset of source code. This dataset, crawled from GitHub in 2025, spans 713 programming languages and 173 million repositories. The move is significant as it addresses concerns about data privacy and ownership in the context of AI model training.
As we reported earlier, OpenAI's breach of Hugging Face's models has fueled calls for stricter AI regulation. The Stack v3's scale and scope have raised questions about the extent of data collection and usage. By providing a way for developers to opt out and check if their code is being used, Hugging Face is taking steps to increase transparency.
What to watch next is how the developer community responds to this initiative and whether it will lead to more stringent regulations on data scraping and AI model training. The fact that Hugging Face has implemented a license detection step in The Stack v3 suggests an effort to respect developers' rights, but the effectiveness of this measure remains to be seen.
Sources
Back to AIPULSEN