Skip to content
7wData Data and AI tools, companies, events, podcast
  • Tools
  • Companies
  • Podcast
  • Articles
  • Events
  • Newsletter
  • Research
  • Sponsor

Table of Contents

General Articles 2021 • By Yves Mulkers

Operationalizing AI — Managing the End-to-End Lifecycle of AI

Operationalizing AI — Managing the End to End Lifecycle of AI
4 min read
CI/CD, Data lineage, Data Science
Curated from medium.com →

We know that the Build-Run-Manage paradigm needs to integrate with existing CI/CD mechanisms to shepherd data science assets from Dev to QA to Production. It also needs to work with popular platform technologies like Docker and Kubernetes / OpenShift. We know that the Manage paradigm needs to establish correlations between model metrics and business KPIs.

Let us take a closer look at each of the five areas.

A data science team is a precious resource for any organization, but they’re often swamped with use case requests from across the company. How can the team scope and prioritize the requests to ensure optimum use of the team?

They start by exploring business ideas in detail, insisting on clarity around the business KPIs. Without that clarity, a team is far too likely to judge the success of a project by the performance of the model — rather than by its impact on the business. Consider an example. A company’s HR department wants to use AI to predict which managers will be high performers. In this case, the business metric might be something outside the feature set and model output — say, a business KPI like employee attrition rate. How will the data science team capture this metric if it’s neither a model input nor a model output? How will they hope to correlate model performance metrics like precision/recall to this external business KPI?

Too often, a business captures its KPIs in emails, slides, or meeting notes, where they’re impossible to track, especially as they change. Insist on capturing and understanding business KPIs up front (even if they change later). Doing so allows the data science team to prioritize and create swim lanes for each use case they undertake. Later in the lifecycle, these KPIs will need to be evaluated and correlated to model performance.

We all know the adages, “There is no Data Science without Data” and “Garbage In, Garbage Out”. It’s clear that good data stands at the heart of any successful AI project. But it’s still worth asking whether the data science team fully understands the data sets they’re dealing with. Do they understand metadata and how it maps to a business glossary? Do they know the lineage of the data, who owns the data, who has touched and transformed the data, and what data they’re actually allowed to work with?

On one occasion, we were working with a client data science team on a fraud detection use case. Our models performed well on hold-out data but performed poorly in production. We considered whether we needed to re-train the model to catch new fraud patterns, but no: something else was fundamentally wrong. Much to the surprise of all, we learned that we’d been working with data generated using rules, not ground truth, despite assurances from the data provider. It was a failure to understand and insist on data lineage — and a demonstration of how important it can be to work with a data steward.

Empower your data science team to “shop for data” in a central catalog — using either metadata or business terms to search, as if they were shopping for items online. Once they get a set of results, give them the ability to explore and understand the data — including its owner, its lineage, its relationship to other datasets, and so on. Based on that exploration, they can request a data feed. Once approved (if approval is needed), make the datasets available via a remote data connection in your data science development environment. As much as possible, respect data gravity and avoid moving data. If the team needs to work with PII or PHI data, establish the appropriate rules and policies to govern access. For example, you can anonymize/tokenize sensitive data. This stage is really about understanding the data you need for your AI initiative within a context of smooth data governance.

Most data scientists relish the build phase of the AI lifecycle — where they can explore the data to understand patterns, select and engineer features, and build and train their models.

In the 7wData directory

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Compare the tools & companies behind this topic

Browse the directory →
  • QlikCompany
  • LakebaseTool
  • SurrealDBCompany
  • PProject NessieCompany
  • WeaviateCompany
  • ClickHouseTool
  • Astra DBTool
  • OpenMetadataCompany

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at medium.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.

Want the structural read on any AI or data company?
INS7GHTS

Want a sharper read on this topic?

Ask ins7ghts how the players compare, what people are actually shipping with, and where the trade-offs land.

Tweet LinkedIn Bluesky Threads Email

Related Articles

Why venture capital isn’t the only way to build a ‘successful’ business
General Articles

Why venture capital isn’t the only way to build a ‘successful’ business

3 min read • May 2016
The 4 Mistakes Most Managers Make with Analytics
General Articles

The 4 Mistakes Most Managers Make with Analytics

4 min read • Jul 2016
Best Practices for Cloud Incident Response
General Articles

Best Practices for Cloud Incident Response

3 min read • 2021
7wData

Independent reporting on AI and data: daily newsletter, podcast, deep dives.

Read

  • Ins7ghts newsletter
  • AI Beat newsletter
  • Latest articles
  • Podcast
  • Research guides

Use

  • Tools directory
  • Company directory
  • Research
  • Events
  • ins7ghts

Company

  • About
  • Contact
  • Sponsor a slot
  • Media kit
  • RSS feed

Follow

  • LinkedIn
  • X
  • YouTube
  • Instagram

© 2026 7wData. Independent. Belgium-based.

Privacy Cookies Terms Imprint Cookie settings
INS7GHTS
New · ins7ghts Drops

The AI governance conversation already moved. Most 2026 plans missed it.

Drop #1 · 60 pages · Launch week €99 (then €149) · ends Thu 9 July

Read Drop #1 →
Cookies on 7wData

We use strictly necessary cookies for the site to work, and optional analytics cookies to understand how readers use 7wData. We never share your data with advertisers. See our Cookie Policy.

Get the AI & data signal, daily. 335k+ already do.
Thanks. Check your inbox to confirm.
Get the AI & data signal

One curated email a day. 335k+ data & AI professionals already read it.

No spam. Unsubscribe anytime.

Check your inbox.

We just sent a confirmation. Click the link to start receiving the daily signal.