Data Quality and Knowledge Readiness for Agentic AI Workflows
An agentic workflow is ready only when its data, identities, knowledge sources, definitions and provenance support the decisions and actions the agent is expected to make.
Define the evidence contract
For every task, state the required fields, permitted sources, time boundary, validation rule and meaning of missing evidence. The contract prevents the agent from substituting plausible text for information the decision actually requires.
Resolve identity and entity quality
Test duplicate customers, inconsistent asset identifiers, supplier names, account hierarchies and document versions. An agent that joins the wrong records can produce a coherent but materially false case.
Measure completeness, validity and timeliness
Quality should be assessed for the intended workflow, not as one enterprise score. A field can be valid for reporting and too stale for operational action. Record how missing or conflicting values change the route.
Prepare governed knowledge sources
Assign owners and effective dates to policies, manuals, contracts and approved guidance. Retrieval should respect jurisdiction, product, language and version. Remove obsolete material or make its status explicit.
Preserve provenance
Log which source, record and version supported each material conclusion. The reviewer should be able to distinguish retrieved evidence, model inference, deterministic calculation and human judgement.
Design for source failure
Specify what happens when a system is unavailable, a source conflicts with another or retrieval returns no applicable evidence. Safe responses include requesting information, using a verified fallback, pausing or escalating.
Monitor data after deployment
Track schema changes, missingness, identity collision, source freshness, retrieval coverage and changes in workflow populations. Data monitoring should connect each defect to affected decisions and cases.
