tech

Rani Molla4/18/24

Meta’s not telling where it got its AI training data

Today Meta unleashed its ChatGPT competitor, Meta AI, across its apps and as a standalone. The company boasts that it is running on its latest, greatest AI model, Llama 3, which was trained on “data of the highest quality”! A dataset seven times larger than Llama2! And includes 4 times more code!

What is that training data? There the company is less loquacious.

Meta said the 15 trillion tokens on which its trained came from “publicly available sources.” Which sources? Meta told The Verge’s Alex Heath that it didn’t include Meta user data, but didn’t give much more in the way of specifics.

It did mention that it includes AI-generated data, or synthetic data: “we used Llama 2 to generate the training data for the text-quality classifiers that are powering Llama 3.” There are plenty of known issues with synthetic or AI-created data, foremost of which is that it can exacerbate existing issues with AI, because it’s liable to spit out a more concentrated version of any garbage it is ingesting.

AI companies are turning to such data because there’s not enough good, public data on the entire internet to train their increasingly greedy AI models. (Meta had reportedly floated buying a publisher like Simon & Schuster to satisfy its insatiable data needs.)

Meta, of course, isn’t the only company that’s tight-lipped about where its AI data is coming from. In a now infamous interview with WSJ’s Johanna Stern, OpenAI’s chief technology officer Mira Murati was unable to answer questions about what Sora, OpenAI’s video generating app, was trained on. YouTube? Facebook? Instagram — she said she wasn’t sure.

Meta’s battle with ChatGPT begins now

Meta’s battle with ChatGPT begins now

What is that training data? There the company is less loquacious.

Meta said the 15 trillion tokens on which its trained came from “publicly available sources.” Which sources? Meta told The Verge’s Alex Heath that it didn’t include Meta user data, but didn’t give much more in the way of specifics.

It did mention that it includes AI-generated data, or synthetic data: “we used Llama 2 to generate the training data for the text-quality classifiers that are powering Llama 3.” There are plenty of known issues with synthetic or AI-created data, foremost of which is that it can exacerbate existing issues with AI, because it’s liable to spit out a more concentrated version of any garbage it is ingesting.

AI companies are turning to such data because there’s not enough good, public data on the entire internet to train their increasingly greedy AI models. (Meta had reportedly floated buying a publisher like Simon & Schuster to satisfy its insatiable data needs.)

Meta, of course, isn’t the only company that’s tight-lipped about where its AI data is coming from. In a now infamous interview with WSJ’s Johanna Stern, OpenAI’s chief technology officer Mira Murati was unable to answer questions about what Sora, OpenAI’s video generating app, was trained on. YouTube? Facebook? Instagram — she said she wasn’t sure.

More Tech

tech

Anthropic reportedly doubles current fundraising round to $20 billion

Anthropic has doubled its current fundraising round to $20 billion on strong investor demand, according reporting from the Financial Times. The new fundraising round would value the company at a staggering $350 billion. That’s up 91% from September, when it raised at a valuation of $183 billion.

The company reportedly received interest totaling 5x to 6x its original $10 billion fundraising goal, and it’s expected to haul in several billion more than that tally before the current round closes.

Anthropic’s success with enterprise customers and the popularity of its Claude Code product are boosting the company’s momentum as it chases the current valuation leader of the AI startup pack: OpenAI.

Anthropic doubles VC fundraising to $20bn on surging investor demand

Anthropic doubles VC fundraising to $20bn on surging investor demand

The company reportedly received interest totaling 5x to 6x its original $10 billion fundraising goal, and it’s expected to haul in several billion more than that tally before the current round closes.

Anthropic’s success with enterprise customers and the popularity of its Claude Code product are boosting the company’s momentum as it chases the current valuation leader of the AI startup pack: OpenAI.

Produce At Whole Foods Market's Flagship Store

Amazon says it’s doubling down on opening Whole Foods stores. That sounds familiar.

The company says it’ll open 100 Whole Foods locations in the next few years. That sounds similar to plans Whole Foods’ CEO laid out in 2024 for opening 30 stores a year. Since then, it appears to have added 14, total.

Incredulous Man

One year after the DeepSeek freak, the AI industry has adjusted and roared back

A look back at how the Chinese startup shattered conventions, changed the way Big Tech thought about AI, and blew a $1 trillion hole in the stock market that got filled right back up... and then soared to new levels.

tech

Georgia lawmakers introduce data center construction moratorium amid statewide pushback

More and more communities across the US are wrestling with the pros and cons of having a data center come to town. Georgia has become a hotspot of resistance to the data centers planned by Big Tech, according to a new report from The Guardian. The Atlanta metro area led the nation in data center construction in 2024.

Georgia state representatives introduced legislation that would place a one-year moratorium on data center construction in the state. Ten Georgia municipalities have already passed local bans on data centers.

Per the report, at least three other states have seen similar data center moratorium legislation introduced in the last week, including Maryland and Oklahoma.

Georgia leads push to ban datacenters used to power America’s AI boom

Georgia leads push to ban datacenters used to power America’s AI boom

Georgia state representatives introduced legislation that would place a one-year moratorium on data center construction in the state. Ten Georgia municipalities have already passed local bans on data centers.

Per the report, at least three other states have seen similar data center moratorium legislation introduced in the last week, including Maryland and Oklahoma.

Tesla’s Robotaxi is way cheaper than Uber, Lyft, or Waymo — but you’ll have to wait a lot longer for one

New data from ride-share comparison app Obi shows how much cheaper and less available Tesla’s autonomous ride-share service is.

Self Driving Taxi Company Waymo Voluntarily Issues Software Recall Over Cars Not Stopping For School Buses

Latest Stories

Sherwood Media, LLC produces fresh and unique perspectives on topical financial news and is a fully owned subsidiary of Robinhood Markets, Inc., and any views expressed here do not necessarily reflect the views of any other Robinhood affiliate, including Robinhood Markets, Inc., Robinhood Financial LLC, Robinhood Securities, LLC, Robinhood Crypto, LLC, or Robinhood Money, LLC.

©2026 Sherwood Media, LLC