Former clients have asked,
“If we produce our documentation in DITA XML, aren’t AI answer engines less likely to hallucinate when people ask questions about our products?”
They were caught off guard when I answered, “Not necessarily.” I followed that declaration with this statement:
Using DITA XML to create documentation doesn’t, on its own, mean an AI answer engine will provide more accurate answers.
There was quiet in the meeting room as the team considered my words. One writer spoke up, breaking the silence to ask,
“I thought DITA’s labels and structured content help AI systems understand our content.”
”You’re right,” I replied, “but only if (and when) the system can retrieve that information, and subsequently, recall it and act upon it. Tech writers must produce accurate facts and clearly explain the relationships customers (humans and machines) need to understand the instructions. AI systems must find the right information and apply it correctly.”
Producing valid DITA topics doesn’t guarantee complete content. Consider, for example, a document that follows DITA rules, but fails to state an important prerequisite or another dependency. When that happens, an AI answer engine may be forced to guess (infer) the missing information and perhaps get it wrong.
My first experience with XML authoring came in 1999 at a pharmaceutical company trying to improve how it produced new drug applications. These submissions could approach 100,000 pages and contained information the U.S. Food and Drug Administration needed to evaluate drugs for approval. The company applied information science to producing and managing that content as reusable XML components.
I’ve advocated for semantic structured content ever since. Much of my subsequent consulting work involved helping tech docs teams adopt DITA, an XML specification, and a component content management system (CCMS).
What DITA Identifies
Tech writers use DITA markup to create content in accordance with the DITA specification. A <task> structures procedural content. Within it, <prereq> identifies prerequisites and each <step> may contain a <cmd> describing an action. The task’s <result> describes its expected outcome. An application can use the prerequisite label to identify requirements the author says apply before the procedure begins.
DITA’s <concept>, <reference>, and <glossentry> identify other types of information. Maps organize topics and their relationships, while metadata can identify the product or audience to which content applies.
Can Valid DITA XML Leave An Answer Unstated?
Yes, it can. And, perhaps more often than some might realize. The docs we produce can conform to the DITA specification while at the same time omitting facts or relationships needed to answer a customer’s question. A correctly structured restoration task, for example, might never state how long trial account users have to restore deleted projects.
Validity also doesn’t establish how fully a team uses DITA’s capabilities. Authors may use less specific markup than the specification allows or omit metadata that would clarify when instructions apply. If the required facts or relationships aren’t explicit elsewhere in the information the engine receives, it may infer them incorrectly. More detailed markup can supply useful cues, but the appropriate level of detail depends on the content and the questions it needs to support.
Suppose an answer engine receives two DITA topics, including their relevant markup and metadata.
The topics contain these statements:
Deleted projects are retained for 30 days. Trial account users can restore deleted projects.
The customer’s question is:
How long does a trial account user have to restore a deleted project?
A product could retain deleted project data for 30 days while giving users 14 days to restore it. If no other available information establishes the restoration period for trial accounts, an answer of “30 days” depends on assuming that retention and restoration have the same duration. An engine could make that assumption incorrectly or recognize that the source doesn’t specify the duration.
Stating the Missing Rule
If the product team verifies a 30-day restoration period for trial accounts, the documentation could say:
Trial account users can restore deleted projects for up to 30 days after deletion.
The author could state this rule within the same DITA elements. The added content gives the restoration period for trial accounts and identifies deletion as its starting event. The engine no longer needs to infer restoration access from a retention statement.
For a customer asking whether they can restore a particular project, the engine also needs evidence that the rule applies to their particular account and project type. Checking the elapsed time requires the deletion date and current date, with precise timestamps if the cutoff requires them. If that information is unavailable, the engine may infer timing or applicability incorrectly.
Encoding More Than Publishing Decisions
In my experience, many tech docs teams limit their use metadata to publishing capabilities, which is helpful, certainly, but not the only use for documentation metadata.
For example, maybe a writer tags a topic so it will display in your developer portal but not in your customer help center. That’s useful for ensuring the appropriate topics are displayed in the relevant output channel, but they don’t ensure the content surfaced in each channel is accurate and trustworthy.
Publishing metadata directs where content may appear but that type of metadata doesn’t encode information like restoration duration (see the previous documentation examples above). The documentation would need to represent that rule elsewhere in the content or metadata. A link between the retention and restoration topics doesn’t establish that the two policies have the same duration.
DITA specialization lets teams introduce more specific semantics, and subject schemes can define controlled values and relationships among subjects. When a team models account types, each label needs a defined scope so the application can tell which policy applies to it.
Related reading: DITA Specialization Overview
Writers could state the restoration rule in a DITA topic or encode it in metadata that the application understands. An external vocabulary, such as Schema.org, can help when its terms accurately describe the relationship you need to express. Schema.org is a collaborative, open-community project that provides standard tags for adding structured data to web content.
What Validation Can Establish
DITA grammar validation checks structural rules, such as which elements are permitted and where they can occur. Text inside permitted elements can be incomplete or incorrect. The content of a <prereq> element can, for example, omit something the customer needs to know before starting the task.
If trial users have only 14 days to restore their projects, the revised 30-day statement gives them the wrong duration. Detecting that error requires evidence of the product’s behavior. An additional validation rule could flag the discrepancy by comparing the documented duration with an authoritative value.
Checking What the Engine Receives
The publishing and retrieval processes tech docs teams use can change which cues reach answer engines. Converting docs to HTML might preserve a prerequisite through a heading, and a retrieval system might carry a product restriction in metadata. Original XML markup can be removed while these representations preserve its meaning.
Tech docs team must inspect the content and metadata supplied to the answering system to determine what survives. If a product restriction is omitted, the engine may apply the content to the wrong product (one the source content excludes).
How to review it
Choose a real customer question.
Trace the proposed answer back to the information the answer engine receives.
Verify that the source docs support every part of the answer.
For example, if “30 days” comes from a retention policy, it doesn’t confirm how long trial users can restore a project. Find an approved source, such as a product specification or test result, that states the restoration period. Confirm the documented behavior before changing the source content.
Then make sure the answer engine receives that rule. To determine whether the change reduces wrong answers, test the revised content with the answer engine. 🤠







I have been wondering about the "Checking What the Engine Receives" issue for a while now, and am still unclear. I see tech writing articles tersely state how DITA/semantic tags help AI engines—it's easy to see the logic there. However, my understanding is that DITA tags are not present on the websites (aka CDPs) where published content lives (it all transforms to HTML). And, wouldn't most AI answer engines just be looking at the website content (without the DITA tags)?
Perhaps there is an easy answer that I'm missing? Are folks pushing raw DITA to their publishing site (similar to other behind-the-scenes metadata) and somehow wiring up the AI engine to look at that instead of the HTML content? If so, isn't this a bit of a major paradigm shift that CMS companies and tooling ppl should be making noise about?
I would like someone to break down in nuts and bolts terms the practical ways that you would have an AI engine on your docs website keying off DITA tags that live in your authoring tool.