Researchers reveal source of AI’s unending pool of knowledge

According to an article published in the Washington Post, the internet and human behaviors on it provide a tremendous reservoir of information for artificial intelligence (AI). Researchers examined a dataset of over 500,000 personal blogs, accounting for 3.8% of the total “tokens” in the dataset.

Google’s C4 dataset, which contains the contents of 15 million webpages, has been used to train high-profile English-language AIs such as T5 and LLaMA from Facebook. Websites from many areas, including journalism, entertainment, software development, medical, and content production, are included in the collection. However, the information also includes at least 27 other sites that the US government has designated as pirate and counterfeit marketplaces.

The websites in Google’s C4 data set that are reportedly responsible for training chatbots include patents.google.com, wikipedia.org, scribd.com, nytimes.com, journals.plos.org, latimes.com, theguardian.com, huffpost.com, patents.com, washingtonpost.com, coursera.org, fool.com, frontiersin.org, instructables.com, ipfs.io, businessinsider.com, chicagotribune.com, booking.com, theatlantic.com, and about 80 others.

Despite the fact that C4 is a huge dataset, large language models are believed to need even larger ones. For example, OpenAI’s GPT-3 training data, which was released in 2020, began with up to 40 times the amount of web scraped data seen in C4. GPT-3’s training data also includes the whole English language Wikipedia, a collection of free books authored by unpublished writers, which are widely used by large technological businesses, and a compilation of text from Reddit users’ favorite links.

Many firms do not document the contents of their training data, either internally or externally, because to worries about uncovering personally identifiable information, copyrighted material, and other data obtained without authorization, according to experts.

The sources for this piece include an article in WashingtonPost.

Researchers reveal source of AI’s unending pool of knowledge

Top Stories

Cisco breach exposes 300+ repos after supply chain attack

OpenAI cuts Sora with $1M-a-day costs to free up resources for core AI models

AI boom faces new constraint as helium disruption slows chip production

New Google update targets sensitive data exposure in search results

GitHub Copilot to train on user data by default

Microsoft pulls Copilot Chat from core Office apps for enterprise customers

Related Articles

Top 10 reflections on information technology developments in 2025

Alphabet to buy data centre and energy firm to boost AI capacity

AI is reshaping how people look for information, Google’s Year in Search 2025 shows

Former Shopify product chief joins OpenAI to lead ChatGPT app platform

TND Newsdesk

TND Newsdesk

Jim Love

Follow Us

Popular categories

Tech News Delivered