"The infrastructure provided by NIST will provide common data, metrics and scoring to help developers understand the performance of their models," the press release said.
AITE: an isolated testbed for AI model safety
The National Institute of Standards and Technology on Monday launched the AI Technology Evaluation (AITE), a voluntary testing vehicle that gives researchers access to an isolated testbed environment to evaluate artificial intelligence models against a set of predefined commands. NIST bills the program as focused on safety analysis: the testbed supplies blind data for models to process so evaluators can gain objective insights into capabilities without exposing production systems or live datasets.
Image analysis tasks: quantum science, genomics and public safety
At rollout, AITE will concentrate on image analysis work performed by large vision language models and will limit its initial scope to three specific domains: quantum science, genomics and public safety. The program’s designers say more tasks will be added over time, but the first tranche of evaluations is narrow by design—targeted image-analysis challenges in these three technical areas.
Blind data, not training material — rules for participants
AITE provides "blind" datasets—that is, material intended for models to process during testing rather than material that should be used for training. The evaluation data is explicitly not intended to serve as training data for the models being assessed. Both categories of participants—data providers and model providers—must submit materials related to testing. Data providers are asked to supply original datasets that are inaccessible publicly and to pair those datasets with a "meaningful" task suited to the data. Model developers will submit their AI models to be run against those datasets within the isolated environment.
Connection to the Commerce Department and a voluntary model-evaluation strategy
NIST’s formation of AITE is described in the source as the latest step in the Trump administration’s strategy to work with major AI developers through voluntary model submissions to advance model safety. Related activity at the Commerce Department includes a renegotiated deal announced in May between the agency and three companies—Google Deepmind, Microsoft and xAI—to evaluate their models through the Center for AI Standards and Innovation. NIST’s program adds a specialized testbed and a rubric intended to standardize measurement across submissions.
How model developers, data providers, and the Commerce Department will act
- Model developers: will submit their AI models to AITE for evaluation in the isolated environment; the first set of evaluations is scheduled to start in August 2026.
- Data providers: will provide original, nonpublic datasets plus a meaningful task description that AITE can use to assess model performance without making those datasets public or using them for model training.
- The Commerce Department: through its prior renegotiated agreement and the Center for AI Standards and Innovation, will continue evaluating models under the administration’s voluntary-submission strategy while AITE supplies a common infrastructure and rubric.
NIST frames AITE as an effort to produce a universal rubric and to determine the state of the art for model performance. By supplying common data, standardized metrics and scoring, the agency aims to give developers an objective baseline for comparison—while keeping the evaluation data sequestered from model training. The first practical test of that theory begins in August 2026, when the initial evaluations commence.
What remains explicit in the record is straightforward: NIST has created an isolated, voluntary testing mechanism focused initially on image analysis in three technical domains; data and model providers must submit controlled materials; and the work sits alongside a Commerce Department arrangement with Google Deepmind, Microsoft and xAI conducted through the Center for AI Standards and Innovation. The next concrete milestone—AITE’s opening evaluations in August 2026—will show whether a common rubric and blind datasets can produce the objective comparisons the agency intends.




