Data governance: structure your data before integrating AI

Without clean, structured and governed data, your AI projects will fail. Here's how to lay the foundations of effective data governance. The problem: scattered data: 85% of AI projects fail (Gartner). The primary cause isn't technological but organisational: data scattered across silos (ERP, CRM, Excel, emails), undocumented, inconsistent and often poor quality. The 4 pillars of data governance: 1) Quality: is data accurate, complete, up-to-date? 2) Catalogue: do you know what data exists, where, and what it means? 3) Lineage: can you trace each data point's origin and transformations? 4) Access and security: who has access to what, with what justification? Data mesh vs data warehouse: Data warehouse centralises all data in a single repository. Data mesh distributes data responsibility to the business teams that produce it. For SMEs, a pragmatic data warehouse (PostgreSQL + dbt) is often the best starting point. Preparing data for AI: AI needs labelled data (for supervised ML), clean data (no duplicates), representative data (avoid selection bias), and data accessible via structured APIs. Data prep typically takes 60-80% of total AI project time. Our method at Powehi: We start every AI project with a data audit: source inventory, quality assessment, gap identification. Then we structure a lightweight data warehouse (PostgreSQL + dbt), set up a data catalogue, and automate quality controls.

Key takeaways

  • 85% of AI projects fail, often due to data
  • 4 pillars: quality, catalogue, lineage, access
  • 60-80% of AI time = data preparation
  • PostgreSQL + dbt = pragmatic SME data warehouse
  • Data audit = first step before any AI project