Skip to main content
AI & Machine Learning

Agencies Misattribute AI Hallucinations, Overlook Data Context Challenges

Person stands amidst rows of computer servers and data storage units in a quiet, softly lit data center.

"Quite often, when people see an inaccurate response from an AI system, the reaction is that the models aren’t good," said Kevin Bohan, Director of Product Marketing at Denodo.

That observation cuts to the center of a growing disconnect inside federal modernization efforts: agencies are pouring energy into procuring larger, more capable generative models while the errors that attract headlines often originate in the information fed to those models. The distinction matters because, as agencies move from pilots to production, the reliability of AI systems depends less on model size and more on the provenance, meaning, and runtime controls around the data the models receive.

Kevin Bohan on where hallucinations really begin

Bohan frames the problem in operational terms: "What’s the trusted source? What’s the golden record? And what additional context helps the AI understand how that information should be used?" He argues that AI systems "only work with the information they receive," and when that information lacks clear definitions, business context, or authoritative sources, even advanced models can produce inaccurate or misleading responses. As agencies explore agentic AI and autonomous decision-support capabilities, he says, ensuring systems can access trusted and contextualized information becomes increasingly important.

Distributed government data isn't solved by bigger models

The source material underlines that government data often sits scattered across agencies, programs, cloud environments, on-premises systems, and external partners. Centralizing every dataset into a single repository is, the piece notes, "difficult to execute across large and complex government landscapes." Instead of moving everything, the practical strategy proposed is to "unify access, governance, and semantics across distributed environments" so users, applications, and AI systems can access trusted information regardless of location. The author stresses that "data will remain distributed due to operational, regulatory, security, or mission requirements."

Active Context and the AI Data Layer: what agencies must build

To address the root cause of many hallucinations, Bohan and the article introduce the notion of an AI Data Layer that provides "Active Context" — described concretely as live, governed, and semantically trusted data from across the enterprise. A semantic layer supplies a shared business and mission vocabulary, but the article emphasizes that semantic alignment alone is not sufficient. Agencies also need governed access, lineage, runtime policy enforcement, and "live awareness of the current state of the data" so models receive metadata that gives the raw facts meaning.

The argument is explicit: "The improvements aren’t necessarily going to come from improving the models," Bohan says. "The improvement is going to come from how well you’re able to provide those models with the appropriate data and the metadata that gives it meaning." For builders, that translates into pointing applications and agents at a trusted layer rather than forcing each project team to recreate governance, integration, and validation work.

What this means for AI builders, policymakers, and agency IT leaders

  • AI builders: Stop treating data as an afterthought. Bohan urges that "AI builders shouldn't have to worry about data" and should be able to "point their applications and agents to a trusted AI Data Layer" where context, governance, and controls are already applied.
  • Policymakers and regulators: Emphasize runtime controls and semantic standards. The article specifies that governance and security controls must be enforced at runtime to prevent sensitive or unauthorized information from reaching models, applications, or agents.
  • Agency IT leaders: Plan for federation, not total consolidation. The recommended modernization posture is to unify access and semantics across distributed environments without requiring every source to be copied into one platform, acknowledging operational, regulatory, and security constraints.

Scaling beyond pilots: enforce controls before models see data

The article warns that what works for a pilot often unravels at enterprise scale: individual project teams cannot sustainably solve the full stack of data integration, governance, security, and context problems on their own. It prescribes a trusted data foundation that provides unified access to distributed information, consistent business context, and governance and security controls "enforced at runtime, before sensitive or unauthorized information reaches an AI model, application, or agent." That foundation is meant to let mission teams and developers focus on outcomes rather than searching for, interpreting, and validating information.

Hallucinations, the piece concludes, are frequently a data and context problem rather than a model flaw. For agencies seeking to modernize responsibly, "one of the most important investments they can make is establishing active context: live, governed, and semantically trusted access to distributed data, so AI systems can operate with greater accuracy, security, and confidence."

Original story