Here’s the formatted transcript with punctuation, capitalization, and paragraph breaks:
---
Hi and welcome to the Data Strategy Gurus podcast. Today, we're diving into the invisible yet foundational systems that power every AI analytics effort: data governance. And helping us unpack today is Lauren Mafo. She's a senior AI/ML program manager for the state of Maryland. She's also an author and a thought leader and the voice behind designing data governance from the ground up. Lauren, welcome.
>> Thanks very much for having me. I'm excited to be here.
>> Lauren, we were discussing already what brought you a bit into the space of data analytics, AI, machine learning.
>> It's not a field that I thought I was going to be in. I was a liberal arts student in university, and then for my graduate education, I really saw myself going into the media space, particularly as a reporter. So I oriented my university, my academic majors for both my bachelor's and master's degrees, around studying media and communications with a particular focus on rhetoric and propaganda.
I also interned at various news organizations as an intern, ranging from radio to television. And so I really wanted to work in the digital media space at a time when you had the triple whammy of the recession happening around 2008 that also coincided with both a decline in ad revenue for the news industry at a time when many tech organizations were rapidly growing and becoming public.
And so that confluence of factors meant that the news organization as we knew it was not going to return. And so what I did after graduate school was work as a freelance tech reporter, covering the tech sector from Europe because I did my graduate degree in London. So I was based in London and did freelance tech reporting on European startups for news outlets like The Guardian and The Next Web.
So that was my first introduction to the tech space at a time when it was rapidly growing. Then a few years later, I did pivot to working in the tech sector for a variety of reasons. And what brought me into the data AI space was working as an analyst at Gartner. I spent four years as a research analyst for them, covering trends in AI and data for their small and midsize business clientele.
I started out covering project management and accounting, then quickly pivoted to AI and data when I realized I thought that was more interesting with a lot more upside. And so I really became intrigued with it by this question of whether AI would grow to a degree that it could replace humans at work and if it could replace jobs en masse. That was something I was interested in.
And so that was the basis of the aptitude that I ended up gaining at Gartner for uncovering different trends in AI and machine learning technologies. And that really was the catalyst about seven years ago for my pivot into the space. So I started as an analyst covering those more philosophical questions about job loss.
And it's pretty interesting, seven years later, to see how that has actually played out because what I've realized is that many of the predictions Gartner was making at the time about what AI would do have actually turned out to be true.
>> Yeah. So a mouthful. I understand you come from journalism. So I see some things which are very important skills for a lot of people in the data space but very often not used, like being critical, being investigative, and trying to understand why things happen.
You already answered one of the questions I had on my mind: how do you start with your data governance? Is it the why, or do you start with the governance? How do you come into that? So I'm kind of intrigued by how you made the switch from being an analyst, being into journalism, then being exposed more to analytics and the tech scene of Europe. But what triggered you especially to go deeper into working with that data?
---
(Note: The transcript cuts off mid-question at the end, so the formatting stops there.)
Here is the formatted transcript with punctuation, capitalization, and paragraph breaks:
---
That's a great question. And after the answer is, after four years at Gartner, I was interested still in analyzing the AI/ML space, but I also did want more hands-on experience working with cross-functional teams to build and—specifically—to design these systems, because I do think that the design component is very often lost in this broader conversation about data and AI. But that really is where a lot of these critical decisions are made—in those foundational early weeks when you're designing the system itself.
And so what I did was, I pivoted to working as a systems and service designer for a company called Steampunk, which works on human-centered design problems for US federal government clients. And I very quickly started working with clients that were building and shipping and updating AI products—and not so much AI products, I misspoke—but data products of various types.
So, for instance, my first project with Steampunk was working on a database—a public database that was housed at a federal agency. It had data going back to the 19th century, including, you know, I think up to 50 million now—individual data points. So it had a wide array of publicly available data for people to use, and it had diverse user groups ranging from economists to people in ad production. So various different types of users needed this data.
And when I, as a service designer, looked behind the hood to see how these monthly—you know—and regular reports were being produced, there were very often no automation involved of any kind. There were no standards for data quality. There were no opportunities to do, you know, QA or any of the best practices around data and pipelines that many tech teams know today.
And that was when the idea really clicked in my head that this was a governance problem, because if you don't know what methodologies you're using, if you don't have a good grasp of the people, processes, and tools that you're using to manage data at scale, you don't have a data strategy—and therefore, you don't have data that is ready to be used for AI.
So that's really where the idea for my book came into play, and that was the impetus for drafting it—was seeing this really wide chasm between the modern tech stack alongside user needs and how those were not being addressed, coupled with the lack of a data program strategy, quality guidelines. And I have continued to—this is not unique to that organization. This is actually the norm in most organizations.
And that's also why it's very difficult to have dialogues on the record with folks—is the fact that this work is not seen as cool or sexy, especially when there's so much pressure on companies to deliver AI products. And so this lack of data readiness comes with a lot of hesitation, I would say.
So a lot of the human aspect—what you bring in here—you say, looking as well at the process. So that's where I feel trying to understand what the human is doing and then bringing the systems to help to automate it as such—but as well, on the other hand, helping people understand what is possible to automate and by repetitive way of doing reports as such—and playing the guard as well, if you talk about data quality and understanding where things go wrong.
So in such a place in your book, you say you built the data governance from the ground up. Why do you explicitly say "from the ground up"? Because you can do that top level. You can start with just first defining all the governance in place and then start to build. But most of the time, it's an afterthought. We need to grant access to a few of the business—
---
(Note: The transcript cuts off abruptly at the end, so the final sentence remains incomplete.)
People in that department need access as well. We then have an issue or data breach, and we start digging into how the flow is and the observability from the data products and the data systems as well. But you explicitly say from the ground up—what do you try to trigger and bring to the people if you say that verb?
>> Yeah, it's a great question. When I talk about data governance being designed from the ground up, there are two points made there. The first is that data governance is fundamentally a design problem to be solved. I wrote the book encouraging readers to view the work of data governance through that lens of design, particularly human-centered design, because the book focuses very closely on that process of identifying who your subject matter experts are regarding each data domain that you identify.
Then it gives a framework for bringing those people—those subject matter experts—into a way of improving data quality for respective domains and having a say in how data quality looks for their respective domains as part of the wider organization. That is a unique way of working. It's something that we practice at the State of Maryland, which I'm very proud of, but it's certainly not the norm.
I think this gets tricky when the audience is highly technical and used to working in a very specific way with very specific people in an organization. If you're from an inherently siloed organization, this way of working is disruptive and is not always well-received. It's also messy to implement because humans are messy.
But the flip side of it is that I said "from the ground up" because I wanted to emphasize that really, if you are a chief data officer, a CIO, CTO, or CDO, your job is to create a holistic data strategy that manages the people, processes, and tools overseeing your data at scale. And then the question becomes, how do I do that?
I wrote the book for people who know that this is on their plate professionally. They believe that their data quality needs to change and evolve, but they're not sure how to get there. "From the ground up" is also interesting framing because businesses are in a catch-22, I think, where they have so much data already—and data is ultimately information.
If you think about your Google Drive and how many documents are on there, it's very easy for it to turn into a landslide—to not be structured the right way, to not have the right naming conventions. Then you're constantly going into the Google Drive looking for one document or a sentence in a document, and you can't find what you need. Data is very similar.
I think the issue then becomes you have this onslaught of data on a day-to-day basis that is very likely not governed or organized in a way to be meaningful. So then the data governance "from the ground up" references that process of stepping back to look at what you can be doing at the highest level to organize it and bring in the right people throughout your organization to help you organize it.
>> Yeah. Yeah, you define a six-step program as well. I've heard you say people know what's on their plate, but they kind of don't know what to tackle first. If we look at the DMB, it's 12 domains, which you can tackle, and everything is in the catch-22, as you say.
So what are the six steps that help them to focus and win their battles very fast and get some engagement from the people as well? Because most of the time, we have to do this and this and this extra. So it feels like a burden starting to do data governance—not having access, filling out papers, whatever.
Checklists—what they do most of the time, it's not practical. It's just, "Okay, fill it out, put in the check marks, and you're compliant." For me, it's kind of, yeah, that's the paperwork, but I still can hack your systems. So, it doesn't guarantee anything. It's just, "Well, I sign off. My responsibility goes out of the way." So, I'm quite interested in what are the six simple steps that you pronounce to help these people get started.
Well, I want to start by acknowledging that what you say certainly feels true, and that's a real risk when you do any type of governance work, whether it's related to data, privacy, or AI. There's very often a confluence of those three areas in an organization, and we worry about this—the concept of being too punitive with people. What is that line between giving guardrails and guidelines versus being the enforcer? I know that on my current team, that is the last thing we want. Our job exists to help people across our state and across our agencies use AI productively and safely. So, we are not a legal body or entity as it relates to compliance.
But I think the first thing to start with is to really have a handle and do an honest assessment of what your data landscape looks like today in your organization. This is a painful process because many organizations, for various reasons, are not where they would like to be or need to be with their data. So then the first question is: What is our organization's mission, and how does our data usage, as is, enhance that mission?
If you speak to any C-suite leader of an organization, they are going to be able to tell you why that organization exists and the value it brings. What they very rarely can do is talk about how their current data structure and organization or governance enhances that mission. And that is a really fundamental gap to close, especially with senior leaders, because again, they are tasked with the strategy to enhance and grow the business in various ways, regardless of what that business does.
So, once you talk about the chasm between the data that they are meant to use and what they exist to do, then I think you can have a conversation about how the business is structured. Because you really want your data to be organized in such a way that it mirrors and enhances the way your organization exists today. This could also be an opportunity to take an honest look at how your organization is structured because maybe there are some higher-level transformation opportunities that you need to take advantage of.
But you want to ultimately look at your people, processes, and tools, which are being used to manage data at a high scale, and think about how to structure your data strategy in a way that is going to mirror and enhance the organization. I'm a big fan of data domains and subdomains because this is really where you get into the organization of your data. It's where you classify different types of data based on the category it falls into, what domain it falls into, and then you can tag it with the right subsequent metadata.
And again, like if, like me, you've worked in digital publishing and information architecture, I see a lot of parallels between information architecture as a discipline and data architecture and design. So, that goes back to why I really do think data governance is a design challenge to solve because there's a lot of overlap.
Once you have a sense, as the data leader, of the processes, tools, and design of your data plan, the next thing that you need to do is look across the organization at the subject matter experts who are closest to that data in each respective domain. These are very often senior leaders of each respective domain—and so your most senior marketing person, your most...
Senior salesperson, these are colleagues who should already exist in the organization.
I sometimes see job descriptions for data stewards, and I think that fundamentally misses the mark on what this role does well, which is the fact that this person does not need to be a data engineer. They just need to be a subject matter expert in the domain so that they can speak to the quality of it.
And then the book does go through the steps that you can take to do the basics of getting a data council together, crafting a road map for your first data product in terms of tasks and milestones that you'll use to measure success. It talks about the ins and outs of shipping that product and some best practices that other tech organizations use.
And then finally, it does talk about the beginnings of the compliance pieces about data destruction policies and how you can accommodate or troubleshoot machine learning drift in your models. So it goes through all of that, and it does assume again that you need a governance strategy for your data but haven't started one yet. The book is really meant to be that 100-page guide to getting off the ground and then creating something that you can build upon.
So yeah, the 100 pages—it sometimes feels like there's nothing much to say about data governance, but you kept it very practical. If I hear you talking on getting started and then finding your way through all the complexities of what you have to put in place, but the practical things like you say—putting in place the data council, knowing how to find your data champions, data stewards—because these are the people already in your organization, but you help them identify in a certain way and avoid being siloed in the organizations to get your data governance and data systems or data strategy very well aligned to the business strategy.
What struck me was that the senior leadership understands what the business needs to do and what is expected from them, but they don't know what type of data they need. For me, it's more like—you know the questions you want to ask and which kind of answers you're looking for. So that's for me a type of direction where you can say, yes, in a certain way, you know what type of data you're looking for.
I understand that you don't have any understanding of the CRM systems or ERP systems that are running out there. So that's the responsibility of the more technical people that can help you understand—is it available? Can we answer these questions? Is there a kind of mechanism? What you have in place to help that align and kind of align the strategic level together with the technical level? Because here, it's fun to build systems and look at it, but we're not aligned to the business strategy, and the business strategy doesn't really understand what is already available in such a way.
Right? I liken this with data stewards who aren't in technical roles to saying that—you know, most of us do not need to know the fundamentals of being a mechanic in order to know how to drive a car. You need to know which steps to take to drive the car, but only a few small people need to know how that car works and is powered under the hood.
Likewise, I think there is an analogy with data. You do need the most senior leaders in your organization to be bought in as participants in the data governance strategy, which I would argue involves tying the success of quality data in their respective domains to their own roles. I think historically, this is an area where the structure of most organizations does not support data governance because data has very historically been seen as one person or one team's job. And really, there is—
Too much data in any organization for it to not affect us, and it affects every job now. Every role, no matter how qualitative, is now tied to metrics of some sort. So I would argue it's on all of us to be talking about how we need data in our organizations and in our roles to succeed.
In terms of bringing people along, this is an area where you might need to work with HR and your peers in senior-level leadership to weave metrics about data governance and participation into each person's roles, job descriptions, tying bonuses or something to participation in the data governance council. Now, this can be a big ask. There's also a lot of nuance here depending on each organization, so I don't want to be too prescriptive about what that looks like. But I do think this is an area where organizations need to be thoughtful about how we're motivating and rewarding our peers for enhancing the quality of our organization's data.
That is something very few organizations do, but I think it's increasingly necessary because if you're not sure why people are behaving a certain way, all you have to do is look at how they're being rewarded or not. That typically gives a clue as to why they're behaving that way. It also gives you information to consider how you might motivate them to behave differently if that's what you need.
Yeah, fully understand. If you say "reward them," it feels always financial, but it sometimes can mean making your life easier, doing your job in a faster, easier way with fewer errors. That's what you need to find out—what motivates who—and then work in that way to set up your program. The human part of complete data governance is very important.
So we're now in the world where everybody is talking about AI, large language models, ChatGPT all over the place, Gemini, CL, and everything like that. How does data governance change the outcomes with the rush of all organizations wanting to do AI and thinking, "Let's do AI, and that will solve all our problems"?
I see two things happening here with the popularity of LLMs like ChatGPT but also Claude, Gemini, Copilot. I think there's becoming a wider distinction between AI for productivity versus building AI models because that gap was very distinct even a year or two ago. Now, if you have access to a closed LLM through your enterprise, you can actually do quite a bit on your own with AI to create shortcuts and tools without any necessity to code.
One example I really like to give and show people is Notebook LM, which is a tool within Gemini. Notebook LM allows you to upload documents to a notebook—you choose which documents you upload. The interface is as simple as uploading an attachment to an email. That's how simple it is to upload documents to your Notebook LM.
The goal is to create a notebook as specific to a group of questions as possible. This goes back to the domains I talked about, where you want to keep your data organized to align with your business domains. It makes sense because then it's easier to find within your lakehouse. With the notebook, you can build and train your own notebooks to query against.
For example, let's say you upload 20 documents about your...
Here’s the formatted transcript with proper punctuation, capitalization, and paragraph breaks:
---
Data governance policies in your organization, and then somebody on your team wants to know how many data domains you have with subsequent metadata, or subcategories, subdomains underneath. You should be able to then query that notebook to get the answer that you need because the answer presumably is within those documents. So rather than having to go through them all one by one, line by line, you should be able to query against the notebook and get what you need in seconds.
For those of us who are in organizations with a lot of documents, this is a highly productive use case. But it's also something that you have to socialize with a lot of folks because most people will not be familiar with using AI in general, and so you have to make it real for them in that particular way.
The flip side of it is that then you have those more advanced teams. These are the more rare teams that have those skills in forward deployment, deployed engineering, and they can build ETL pipelines. They can write Python scripts. Those are the people who are more valuable than ever but also exist to build something new and/or build on top of what already exists.
I think there's an opportunity to leverage this group to actually automate quality checks for the data. Now, the flip side is that you have to have standards for what quality data looks like, which many organizations do not have. But if you do have those highly technical, highly skilled workers available to you either in-house or through a contractor, I think using them to assess the quality of your data at scale across domains and then automate checks for how it can scan your data at scale to get it to adhere to your standards—that's a really untapped opportunity, I would say.
So I hear you saying, let's have that large language model internally deployed and available to a lot of people. Use notebook LLMs. The notebook LLM kind of excels in what it was in the past, but this has more opportunities, and use the LLM as well to help us with AI scan the notebooks to understand what is available and in a faster way, get a view of all the data sets and data measures you have built over time instead of having to go through your Google Drive or all the Excel files on the network in a certain way.
This is where I see as well that LLMs and generative AI will help us discover in a faster way. We're still facing a lot of memory issues with the large language models because it doesn't know anymore what was yesterday or the day before, and it forgets what you just did. So these are the challenges we're facing for the time being. But it helps us in a certain way.
I think we need to train a lot of people to understand what the capabilities are, and even in a certain way, currently, we still need to do a lot of prompt engineering to get the right answers back from what we have. But over time, these systems—you see it's insanely quadrupling the performance and the quality of what it's getting back every month—it goes at an amazing speed.
So I think in a year's time, it will be really what we have in mind, capable of doing the large language models, generative AI, and helping us bring more out of our data in a certain way.
So what are, in fact, your personal values that drive you in your work to do data governance? And just being so fascinated by all the possibilities of data—if you can call it possibilities—I look at it, I can analyze a lot of things, and now with AI, I can do it at an amazing speed.
---
(Note: The transcript retains all original spoken words, including repetitions and filler words like "uh," while improving readability with punctuation and structure.)
Here is the formatted transcript with proper punctuation, capitalization, and paragraph breaks:
---
I can't write these queries that fast. I see the large language models still make mistakes a lot of times, not finding the right fields. They go on and they do the reasoning to understand how the structure of your database works. So it's kind of mimicking our behavior of doing exploration of the data systems in a certain way where I think, "Yes, but I know this is the field." So there is still a gap in what generative AI is capable of doing as what I'm expecting it to do already. That's my kind of frustration by working so much with the large language models for the time being. But I try to understand what drives you so much and what are your personal values in working with data and governance.
That's a great question, and the main value I'd say is trying to create equitable access to AI in terms of socializing it as a tool to help as many people improve their quality of life as possible, especially at work. But also, what drives me most is going back to that initial question of why I got into data governance and AI to begin with. It really was that fundamental question of whether AI was going to cause mass job loss, which, in some ways, that is happening. It's been happening for some time, but it's been scaling to other industries and sectors on a wider scale in recent years.
And so there's that aspect of it. But then when I started researching even more the possible effects of AI and technology on taking human employment, especially in a country without many social safety nets, it really depends on the state that you live in. The other side of that was that in about 2018, that's when the conversation in computer science really started gaining steam around the results of AI and how there was a total lack of governance regarding these AI models. And so even in cases where the technology was deployed, it was having really adverse effects on a lot of users.
As a result of that, we are trying to then deploy these products to production for people to use, and very often these tools are used for people or used on people's behalf without their consent, without their knowledge, especially for surveillance purposes, especially for surveillance of groups which are often of minority status and often already more persecuted.
So the longer answer to a shorter question is thinking about how this incredibly powerful technology really has the opportunity to do substantial harm while also being highly effective and productive. And so I really am motivated by this concept of trying to democratize AI and teach as many people to use it in as diverse ways as possible while also underpinning the AI—the technology that lies below it, which, and really, it's not even technology, just the data underneath.
Yeah. So you still believe in the very good of the technology and are not too afraid of things going wrong. Yes, we're afraid of it in a certain way, but we make sure that the systems get in place with governing it in the right way, helping people understand what is possible.
Lauren, as we come to an end of our conversation, I always like to ask my guests, what are your favorite band or the types of music that helps you think or inspires you?
It's tough to say only because I love music, and I've realized I can't really listen to music while I work. I get too distracted too easily, and I spent years listening to music while I worked with not the best results. I actually find that I need quiet in order to do the work. But I am a big fan of alt pop and alt rock music. And so sometimes when I'm stuck on a problem, I will put a band like the...
1975 on to listen to.
>> Wow, that's, uh, that's nice. Uh, Lauren, as a, as a final thing, I like to fire a few questions to, um, just tease your mind a little bit. Um, data mesh or a mesh of network of humans—what would you choose?
>> I think done well, data mesh is a network of humans because you have—you should have—uh, by defining those data domains, included the leaders of those respective domains in the governance process and in the creation process. I—it's interesting that word "governance." It carries a lot of baggage in a lot of different sectors. And so even though I wrote a book on the concept, I try to steer people to think about it as more of a co-creation process. But I do think the data mesh should be at its heart a mesh of humans.
>> Yeah. Top-down mandates or bottom-up stewardship?
>> It needs to be bottom-up stewardship with top-level sponsorship. In my experience, you can do very little as it relates to data AI ethics without the buy-in of your most senior leadership. So if you have to spend your precious time convincing them why you need to design ethical AI, I think your time, frankly, is better spent elsewhere with another leader in another organization. So, it does need to be bottom-up stewardship with top-level sponsorship.
>> You're giving me a hard time just making the choices. What do you prefer? A conference keynote or a deep-dive workshop?
>> I would say a conference keynote, but with the caveat that I am writing a keynote speech now, and it is not—it's not as easy as it appears, especially if you are going to talk about anything more personal, at least in my experience. So I do enjoy a deep dive, and I will be doing more of those in my role as program manager for the state of Maryland. There is a lot of professional satisfaction that comes out of leading those, but for now, I will choose the keynote.
>> Policy-driven AI or product-driven AI?
>> I think it needs to be product-driven AI with policies as a backbone. Uh, because you do need guardrails and policies to know what's permissible and what's not—what you are working towards. That's the most important thing in product: what are we actually doing? Why are we deploying this? And there is a culture in AI of building for fun, which of course is great, but that doesn't necessarily mean you need to deploy every product to production. You can still tinker all you want and experiment all you want, but that doesn't mean that you have to deploy something for users. But if you're really going for product-market fit, it needs to be product-driven. To be product-driven, you need to have an answer to the why.
>> Indeed. Ethics by design or compliance by default?
>> I think it needs to be ethics by design. You need to really prioritize ethics at the outset of any AI product and weave it into every aspect of the product, including at that earliest stage of designing the project or the pilot itself. We also run a lot of AI pilot projects with the intent that it is grounds for experimentation. But before we can begin that experimentation, we need to make sure that we've designed the project not only to succeed but to meet our internal guidelines.
>> Lauren, as you work for the city—for the state of Maryland—citizen trust or tech efficiency?
>> It needs to start with citizen trust because if citizens do not trust what we're doing, we really don't have any leg to stand on, and most of them are not going to care about the tech efficiency. We're at a critical juncture in the United States where trust in government is at all-time lows. And so we fundamentally have to work first and foremost on getting to a level of—
Trust with our citizens gives them the faith that they are going to get the services they deliver faster in a more highly quality state than they have in the past. Because if they don't trust us, it doesn't matter how efficient our tech is.
And then the last one, we briefly already touched on it. Start with the why or start with the data.
You have to start with the why because that determines which data you need and what your data strategy is going to be in the first place.
Okay, Lauren, thanks very much for your time. Great to have you on the show, and I've learned a lot again. So, thanks very much for being here.
Oh, thank you so much for having me. I would love for folks to check out the book if they feel inclined. It's called *Designing Data Governance from the Ground Up*, and you can get it wherever books are sold.
Okay, great. Thanks. Any other places where people can find you, Lauren? Online?
I'd be happy to connect with folks on LinkedIn. You can find me there under my name, Lauren Mafo. That's the social platform that I'm on for business most. And then I am also available on GitHub if people want to connect there. We actually just put our AI enablement team's work into a GitHub repo. So if people are interested in what we're doing in the state of Maryland with AI as well as where we're going, they can feel free to follow our repo.
Thanks very much. Thanks for having me.