Tech
tech

Meta scrambling to defend its AI after Llama 4 benchmark bungle

This weekend, Meta surprised everyone and released two flavors (“Maverick” medium and “Scout” small) of its highly anticipated Llama 4 AI model. Llama 4’s release is a big deal, as the company has been hyping it up as the key to its AI plans in the coming year.

When a major new model drops, people do two things: check to see how the model scored on major benchmarks, and load up the model and kick the tires.

Llama 4’s benchmark scored some eye-popping results for ChatbotArea, a popular human-powered benchmark that’s a sort of blind taste test for AI models with side-by-side results. But after looking at the fine print, some in the community cried foul, as Meta achieved the higher score using an “experimental chat version” of Llama 4 that was not available to the public.

A footnote to a chart that highlighted Llama 4’s standout score read “LMArena testing was conducted using Llama 4 Maverick optimized for conversationality.”

In response to the controversy, LMArena (which runs the Chatbot Arena benchmark) updated its guidelines for testing:

“Meta’s interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that ‘Llama-4-Maverick-03-26-Experimental’ was a customized model to optimize for human preference. As a result of that we are updating our leaderboard policies to reinforce our commitment to fair, reproducible evaluations so this confusion doesn’t occur in the future.”

This led to some unfounded accusations that Meta had trained its model on test datasets — akin to giving a kid the answers to a quiz before having them take the test.

To quell the firestorm of questions surrounding the model’s release, Meta’s head of generative AI, Ahmad Al-Dahle, refuted the claims in a post on X yesterday.

The release was also unusual for what was missing from the release: the extra-large version of the model named “Behemoth.” Meta said the model was still being trained, but boasted about its performance nonetheless.

“Llama 4 Behemoth outperforms GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. Llama 4 Behemoth is still training, and we’re excited to share more details about it even while it’s still in flight.”

Meta did not immediately respond to a request for comment.

More Tech

See all Tech
tech

Getty Images suffers partial defeat in UK lawsuit against Stability AI

Stability AI, the creator of image generation tool Stable Diffusion, largely defended itself from a copyright violation lawsuit filed by Getty Images, which alleged the company illegally trained its AI models on Getty’s image library.

Lacking strong enough evidence, Getty dropped the part of the case alleging illegal training mid-trial, according to Reuters reporting.

Responding to the decision, Getty said in a press release:

“Today’s ruling confirms that Stable Diffusion’s inclusion of Getty Images’ trademarks in AI‑generated outputs infringed those trademarks. ... The ruling delivered another key finding; that, wherever the training and development did take place, Getty Images’ copyright‑protected works were used to train Stable Diffusion.”

Stability AI still faces a lawsuit from Getty in US courts, which remains ongoing.

A number of high-profile copyright cases are still working their way through the courts, as copyright holders seek to win strong protections for their works that were used to train AI models from a number of Big Tech companies.

Responding to the decision, Getty said in a press release:

“Today’s ruling confirms that Stable Diffusion’s inclusion of Getty Images’ trademarks in AI‑generated outputs infringed those trademarks. ... The ruling delivered another key finding; that, wherever the training and development did take place, Getty Images’ copyright‑protected works were used to train Stable Diffusion.”

Stability AI still faces a lawsuit from Getty in US courts, which remains ongoing.

A number of high-profile copyright cases are still working their way through the courts, as copyright holders seek to win strong protections for their works that were used to train AI models from a number of Big Tech companies.

tech

Norway’s wealth fund, Tesla’s sixth-largest institutional investor, votes against Musk’s pay package

Norway’s Norges Bank Investment Management, the world’s largest sovereign wealth fund, said Tuesday that it voted against Tesla CEO Elon Musk’s $1 trillion pay package, ahead of the EV company’s annual shareholder meeting Thursday. The fund, which has a 1.2% stake in Tesla, is the company’s sixth-largest institutional investor, according to FactSet, and the first major investor to disclose how it voted on the matter.

Tesla is down nearly 3% premarket, amid a wider pullback in equities that’s most pronounced in AI-related stocks.

“While we appreciate the significant value created under Mr. Musk’s visionary role, we are concerned about the total size of the award, dilution, and lack of mitigation of key person risk- consistent with our views on executive compensation,” NBIM said in a statement.

Tesla’s board considers Musk’s mammoth, performance-based pay package necessary to retain Musk. For what it’s worth, prediction markets are quite certain investors will pass the proposition.

Tesla is down nearly 3% premarket, amid a wider pullback in equities that’s most pronounced in AI-related stocks.

“While we appreciate the significant value created under Mr. Musk’s visionary role, we are concerned about the total size of the award, dilution, and lack of mitigation of key person risk- consistent with our views on executive compensation,” NBIM said in a statement.

Tesla’s board considers Musk’s mammoth, performance-based pay package necessary to retain Musk. For what it’s worth, prediction markets are quite certain investors will pass the proposition.

tech

Waymo to expand robotaxi service to Detroit, Las Vegas, and San Diego

Google’s Waymo robotaxi service is expanding to three new cities — Detroit, Las Vegas, and San Diego — where it has previously tested its driverless vehicles. Waymo plans to bring its Jaguar I-Pace and Zeekr RT vehicles to those three markets this week, but they won’t be immediately available to the public.

Currently Waymo is available in five US cities: Atlanta, Austin, Los Angeles, Phoenix, and San Francisco.

Tesla is currently testing in Las Vegas, while Amazon’s Zoox has limited service in the city.

Currently Waymo is available in five US cities: Atlanta, Austin, Los Angeles, Phoenix, and San Francisco.

Tesla is currently testing in Las Vegas, while Amazon’s Zoox has limited service in the city.

Latest Stories

Sherwood Media, LLC produces fresh and unique perspectives on topical financial news and is a fully owned subsidiary of Robinhood Markets, Inc., and any views expressed here do not necessarily reflect the views of any other Robinhood affiliate, including Robinhood Markets, Inc., Robinhood Financial LLC, Robinhood Securities, LLC, Robinhood Crypto, LLC, or Robinhood Money, LLC.