Exclusive: Google Plans New ‘Frozen’ Chip to Run Its AI Models Much More Efficiently Save 25% to access this list

Sign in
Subscribe

    Data Tools

    • About Pro
    • Enterprise Software Startup Takeover List 2026
    • The Next GPs 2026
    • The Executives Leading the Data Center Race
    • The Next GPs 2025
    • The Rising Stars of AI Research
    • Leaders of the AI Shopping Revolution
    • Enterprise Software Startup Takeover List 2025
    • Org Charts
    • The Information 50 2025
    • Generative AI Takeover List
    • Generative AI Database
    • AI Chip Database
    • AI Data Center Database
    • Tech IPO Tracker
    • Tech Sentiment Tracker
    • Gigafactory Database

    Special Projects

    • The Information 50 Database
    • VC Diversity Index
    • Enterprise Tech Powerlist
  • Org Charts
  • Deep Research
  • Tech
  • Finance
  • Weekend
  • Charts
  • Events
  • TITV
    • Directory

      Search, find and engage with others who are serious about tech and business.

    • Forum

      Follow and be a part of discussions about tech, finance and media.

    • Brand Partnerships

      Premium advertising opportunities for brands

    • Group Subscriptions

      Team access to our exclusive tech news

    • Newsletters

      Journalists who break and shape the news, in your inbox

    • Video

      Catch up on conversations with global leaders in tech, media and finance

    • Partner Content

      Explore our recent partner collaborations

      XFacebookLinkedInThreadsInstagram
    • Help & Support
    • RSS Feed
    • Careers
    Sign in
  • About Pro
  • Enterprise Software Startup Takeover List 2026
  • The Next GPs 2026
  • The Executives Leading the Data Center Race
  • The Next GPs 2025
  • The Rising Stars of AI Research
  • Leaders of the AI Shopping Revolution
  • Enterprise Software Startup Takeover List 2025
  • Org Charts
  • The Information 50 2025
  • Generative AI Takeover List
  • Generative AI Database
  • AI Chip Database
  • AI Data Center Database
  • Tech IPO Tracker
  • Tech Sentiment Tracker
  • Gigafactory Database

SPECIAL PROJECTS

  • The Information 50 Database
  • VC Diversity Index
  • Enterprise Tech Powerlist
Deep Research
TITV
Tech
Finance
Weekend
Charts
Events
Newsletters
  • Directory

    Search, find and engage with others who are serious about tech and business.

  • Forum

    Follow and be a part of discussions about tech, finance and media.

  • Brand Partnerships

    Premium advertising opportunities for brands

  • Group Subscriptions

    Team access to our exclusive tech news

  • Newsletters

    Journalists who break and shape the news, in your inbox

  • Video

    Catch up on conversations with global leaders in tech, media and finance

  • Partner Content

    Explore our recent partner collaborations

Subscribe
  • Sign in
  • Search
  • Opinion
  • Venture Capital
  • Artificial Intelligence
  • Startups
  • Market Research
    XFacebookLinkedInThreadsInstagram
  • Help & Support
  • RSS Feed
  • Careers

In-depth insights in seconds. Ask Deep Research.

AI Agenda

AI Evaluators Struggle with Models That Know When They’re Being Tested

Art via Getty Images
By
Rocket Drew
[email protected]Profile and archive

AI researchers are starting to make progress on a confounding problem: AI models are getting better at telling when they are in an evaluation.

That could become a problem for AI companies that use evaluations to gauge the capabilities and behaviors of their models before releasing them. If models act differently during testing, that could mean they get released with undesirable tendencies. It could also undermine their creators’ ability to show off test scores to potential clients. 

Evaluations are important for “convincing customers that our products are better at their use case than other products,” said Silas Alberti, who works on evaluations at Cognition, the AI coding startup.

And as models get smarter, they are gaining even more eval awareness, as researchers call it. For example, in testing of its non-public Mythos model, Anthropic found that Mythos more often mentioned that it was being tested than its predecessors Claude Opus 4.6 and Sonnet 4.6.

Recommended