Center for AI Standards and Innovation: Model Testing

The Center for AI Standards and Innovation, or CAISI, evaluates commercial and foreign models, coordinates security research, and reviews frontier systems before release.

kindGovernment
founded2023
Event history

From AI Safety Institute to CAISI

The body now called CAISI was created inside NIST in November 2023, the day after President Biden signed an executive order on AI, and announced at the UK's Bletchley Park AI Safety Summit as the US AI Safety Institute. Elizabeth Kelly, a former White House economic policy adviser, was appointed to lead it in February 2024, and it received a $10 million budget allocation the following month. [2]

In June 2025, under the Trump administration, Commerce Secretary Howard Lutnick announced the institute would be renamed the Center for AI Standards and Innovation. The stated shift was away from the original safety-evaluation framing and toward, in Commerce's own description, positioning CAISI as "industry's primary point of contact within the U.S. government to facilitate testing and collaborative research" on commercial AI systems. NIST's current description of CAISI's role spans voluntary evaluation agreements with AI developers, assessment of security vulnerabilities including possible backdoors, coordination with defense and intelligence agencies, and representing US interests in international AI standards work. [3][4]

What CAISI actually publishes

CAISI's own evaluations are its clearest public output. A May 2026 assessment of DeepSeek V4 Pro, an open-weight Chinese model, tested it across cybersecurity, software engineering, natural sciences, abstract reasoning and mathematics using nine benchmarks, including some CAISI keeps non-public. It found DeepSeek V4 the most capable Chinese model CAISI had evaluated to date but roughly eight months behind the frontier in aggregate, a weaker showing than DeepSeek's own reported comparisons to frontier US models, particularly on reasoning and agent tasks, while beating GPT-5.4 mini on cost in five of seven benchmarks tested. In March 2026, CAISI signed a cooperative research agreement with OpenMined, a nonprofit building privacy-preserving computation tools, to develop methods for evaluating AI systems on sensitive data without exposing the underlying models, datasets or benchmarks. [5][6]

On May 5, 2026, CAISI signed pre-deployment evaluation agreements with Google DeepMind, Microsoft and xAI, extending an arrangement OpenAI and Anthropic had already made in 2024: the labs let CAISI test models before public release rather than only after. Five weeks later, the same government suspended foreign-national access to Anthropic's Fable 5 and Mythos 5 and held back OpenAI's GPT-5.6 rollout on national-security grounds, a separate action taken by Commerce's export-control authority rather than by CAISI itself. [7][8]

A mandate that changed shape, not just name

Atlas interpretation: The 2025 rename tracks a real change in emphasis, from an institute framed around AI safety risk to one framed around voluntary evaluation partnerships with industry and competitive assessment of foreign models, alongside a continuing national-security testing role. Both framings still produce the same basic function this page cares about: an arm of the US government that tests models other than its own claims about them, whether the public rationale offered for that testing is safety or competitiveness. This page does not take a position on whether the narrower, industry-facing mandate serves either goal better than the original one did. [2][4][5]

Sources

  1. NIST Seeks Collaborators for Consortium Supporting Artificial Intelligence Safety

    NIST · Nov 2, 2023

  2. Center for AI Standards and Innovation

    Wikipedia · Sep 9, 2026

  3. Statement from U.S. Secretary of Commerce Howard Lutnick on Transforming the U.S. AI Safety Institute into the Center for AI Standards and Innovation (CAISI)

    U.S. Department of Commerce · Sep 9, 2026

  4. Center for AI Standards and Innovation (CAISI)

    NIST · Sep 9, 2026

  5. CAISI Evaluation of DeepSeek V4 Pro

    NIST · May 1, 2026

  6. Announcement: CAISI Signs CRADA with OpenMined to Enable Secure AI Evaluations

    NIST · Mar 27, 2026

  7. Trump admin moves further into AI oversight, will test Google, Microsoft and xAI models

    CNBC · May 5, 2026

  8. Statement on the US government directive to suspend access to Fable 5 and Mythos 5

    Anthropic · Jun 12, 2026