Who Needs A Data Model Anyway?

Will AI eliminate the need for data models?
With data lakes offering to store raw data and promising schema-on-read access, data warehouses moving in-memory for vastly enhanced query performance, and even BI tools improving ease-of-use with artificial intelligence (AI), many in the IT industry are proclaiming the imminent death of the data model.
“A model is so limiting for innovative users,” they say, “and building one is so terribly time-consuming. The idea is so last century. Let it fade away like the images on some old celluloid movie.”
I beg to disagree. I don’t dispute that data models are changing dramatically or that new ways of building and using them must emerge. However, the term data model includes three different concepts (application-scope modeling, information structuring and simplification, and enterprise data models) and we must examine each to see how the ongoing evolution in information and process will re-create the data model landscape.
Furthermore, as discussed in a previous Upside article, data modeling occurs at three levels — conceptual, logical, and physical. Of these, only the physical level is directly affected by technology change, although the implications of AI for the other two levels require deeper consideration than we have heretofore seen.
Keep in mind that each of these flavors of data model serves a different purpose and audience, so the usefulness and prospects of each in today’s information and process technology environment must be considered separately. In addition, the sources of the data under consideration play a role in this evaluation. Internally sourced data — process mediated data, as I have described it in Business unIntelligence — is typically well-defined and well-structured but often semantically complex in comparison to externally sourced data.
All three concepts do indeed have their roots in the last millennium. Application-scope modeling was the first incarnation, introduced in 1976 by Chen with the original entity-relationship diagram. The aim was to allow “important semantic information about the real world” to be incorporated into the then-emerging databases in the design of operational applications. The important characteristic of such modeling is that it is local in scope, driven by the specific needs of a particular business function. It applies to operational systems as well as to limited-scope informational systems, such as data marts.
Modern improvements in database performance and data storage methods certainly reduce — and, in some cases, eliminate — the need for physical data modeling.

