The Starting Point of AI Sovereignty: Data and Governance

This column was written by Dtonic CEO Yong Joo Jun. It was originally published in Korean by ETNews on July 15, 2026, and has been translated into English and reposted here with permission. Read the original article in Korean [here]

——

No matter how far artificial intelligence technology advances, one fact never changes: more than 80% of the time spent on an AI project goes not into modeling, but into "preprocessing" — cleaning and preparing data. This is the so-called "80/20 rule of AI," proposed by the AI consulting firm Cognilytica, and it points to a truth that remains constant even amid rapid technological advancement: the real Achilles' heel of AI is data.

Yet the market's attention has largely been fixed on the dazzling performance of large language models (LLMs), and this fact has been overlooked. The "Anthropic Mythos incident" showed just how dangerous that sweet dependency had become — a ticking time bomb. This is also the backdrop against which countries have elevated "sovereign AI" to a national strategic agenda.

Here we need to soberly identify the essential point. The Mythos incident was a crisis triggered by a model being cut off, but paradoxically, the survival card that lets us break through such a crisis lies not in the model, but in the data. The power to immediately switch to an open-source or alternative model when a particular model is blocked comes from data. What makes that possible is data that has been standardized and refined so that it can plug into any model instantly, like a module. Rather than competing with global Big Tech on model size, the true starting point of AI sovereignty is data governance — the set of rules by which an organization controls and operates its own data.

AI today is rapidly evolving beyond simply executing instructions, into "agentic AI" that understands goals, formulates plans, and executes them on its own. Gartner forecasts that by 2028, 15% of everyday enterprise decisions will be made autonomously by agentic AI, and that 33% of enterprise software will include this capability. AI is evolving into an agent of business decision-making.

However, poor governance translates directly into critical risk. Gartner warns that low-quality data and the absence of governance cost enterprises an average of $12.9 million annually in financial losses, and that as a result, more than 40% of agentic AI projects will be discontinued by 2027.

If the era of digital transformation (DX) was defined by a "quantitative competition" to amass as much data as possible, the era of AI transformation (AX) is defined by a "qualitative competition" over data. The reason so many companies' AI projects stall at the initial proof-of-concept (PoC) stage is not a lack of technology, but the absence of a data governance framework capable of aligning the formats of scattered data and guaranteeing its quality. Only when the entire lifecycle of data — from creation, to use, to disposal — can be systematically controlled and its history traced, can reliability be guaranteed in an actual operational environment.

The defense sector demonstrates the importance of data most starkly of all. On the future battlefield, centered on manned-unmanned teaming (MUM-T) and multi-domain operations (MDO), vast volumes of heterogeneous data pour in real time from countless systems.

Only when data from disparate sources can be integrated into a single situational context, and AI can understand the meaning and relationships within that data on its own and apply it appropriately to the situation, does the value of that hard-won data truly shine.

On top of that foundation, situational-awareness AI is completed — AI capable of explaining the very basis for its judgments. Because this is a domain where lives and national security are directly at stake, an explainable and trustworthy data operating framework is itself combat power.

This shift is not confined to any single industry. The more a field generates data in real time — manufacturing, distribution, energy, smart cities, defense — the greater the importance of data governance becomes.

Dtonic, too, has spent more than a decade developing platform technology that uses AI to integrate and operate spatiotemporal, heterogeneous data across a wide range of real-world settings.

We have focused not only on streamlining data collection, refinement, and connection — the biggest bottleneck in AI projects — but on building a data-centric platform that allows AI to accurately understand and make use of the meaning and context of data. In the defense sector as well, we are partnering with LIG D&A to develop "L-NODE," an AI platform specialized for the defense industry.

Ultimately, AI performance is not determined by the model alone, but by how correctly AI can understand and make use of trustworthy data.

The recent restrictions on AI model access pose an important question to all of us: "If we could no longer access the AI model we use, starting tomorrow, what would remain?" AI models can change at any time. But a framework for understanding and using data cannot be built overnight.

Now that we have entered the era of agentic AI, the first thing we must invest in is data governance — the discipline that makes data understandable and usable. That, too, is where the starting point of AI sovereignty lies.

Next
Next

Five industries that will be redefined by the Data OS & what it means for everyone else