Marquez

Marquez is an open-source project hosted on GitHub under the MarquezProject organization that provides a metadata service for data lineage, data discovery, and data governance.

Reviewed by 7wData

On this page

Profile

Marquez is an open-source metadata service that collects, stores, and visualizes data lineage and data asset metadata across data pipelines.

Marquez is an open-source project hosted on GitHub under the MarquezProject organization that provides a metadata service for data lineage, data discovery, and data governance. The project, which has accumulated 2,200 stars and 399 forks as of June 2026, is designed to collect and visualize the lifecycle of datasets as they move through data pipelines. It is primarily aimed at data engineering and data platform teams who need to track the provenance of data assets across complex, distributed systems.

The repository shows 2,861 commits on the main branch, with 66 branches and 94 tags, indicating active development. Marquez integrates with Apache Airflow and other orchestration tools, offering a web UI and a REST API. The project does not appear to have a corporate parent, dedicated funding, or disclosed revenue; it is maintained as a community-driven open-source initiative.

No recent funding rounds, headcount figures, or commercial product offerings were found in the research dossier. The project's GitHub presence and documentation suggest it is used by organizations that operate data pipelines at scale, but no named clients or commercial adoption metrics were available in the provided sources. The project's long-term viability depends on continued community contributions and sponsorship, as it lacks the financial backing of a venture-backed startup or a large vendor.

Track Marquez and 240+ vendors.

335k+ subscribers read the daily AI & data note. One email, both newsletters. Unsubscribe anytime.

Who buys this

  • Data engineering teams that need to trace data provenance across multiple systems
  • Data platform teams building internal data catalogs and governance tools
  • Organizations running Apache Airflow or other orchestration frameworks that require lineage tracking
  • Enterprises subject to data compliance regulations that demand audit trails for data transformations

Strengths and what to watch

Strengths

  • Active open-source community with 2,200 GitHub stars, 399 forks, and 2,861 commits on the main branch as of June 2026
  • Provides a standardized API and web UI for data lineage, which is a common pain point for organizations with complex data pipelines
  • Integrates with Apache Airflow, a widely adopted workflow orchestrator, lowering the barrier to adoption for existing Airflow users

Watch for

  • No disclosed funding, revenue, or corporate parent, making the project's long-term maintenance dependent on volunteer contributors and uncertain sponsorship
  • No named commercial customers or adoption metrics in the research dossier, raising questions about real-world deployment scale
  • The project competes with well-funded commercial alternatives (e.g., Atlan, Alation, Collibra) that offer broader data governance features and dedicated support

Key Information

Industry
Data Management
Founded
1986

Frequently Asked Questions

What is Marquez and what does it do?

Marquez is an open-source metadata service that collects, stores, and visualizes data lineage and data asset metadata across data pipelines. It helps data engineering teams track the lifecycle of datasets as they move through complex, distributed systems.

How does Marquez track data lineage?

Marquez collects and visualizes the lifecycle of datasets as they move through data pipelines. It provides a standardized API and a web UI for data lineage, allowing teams to trace data provenance across multiple systems and understand data transformations.

Does Marquez integrate with Apache Airflow?

Yes, Marquez integrates with Apache Airflow and other orchestration tools. This integration lowers the barrier to adoption for existing Airflow users, enabling them to capture data lineage automatically as workflows execute.

Is Marquez a commercial product or open source?

Marquez is an open-source project hosted on GitHub under the MarquezProject organization. It does not have a corporate parent, dedicated funding, or disclosed revenue. The project is maintained as a community-driven initiative with active development.

Who typically uses Marquez for data governance?

Marquez is used by data engineering teams that need to trace data provenance, data platform teams building internal catalogs, and organizations running Apache Airflow. It also helps enterprises subject to data compliance regulations requiring audit trails for transformations.

How active is the Marquez open source community?

As of June 2026, the Marquez GitHub repository has 2,200 stars, 399 forks, and 2,861 commits on the main branch. It also has 66 branches and 94 tags, indicating active development and community contributions.

Sources

  1. github.com — Project repository details: 2,200 stars, 399 forks, 2,861 commits, 66 branches, 94 tags, integration with Airflow, web UI, REST API
  2. investors.mheducation.com — McGraw Hill Q3 2026 financial results (unrelated to Marquez, but included in dossier)
  3. www.msd.com — Merck Q1 2026 financial results (unrelated to Marquez, but included in dossier)
  4. techcrunch.com — Benchmark raises $2B in capital (unrelated to Marquez, but included in dossier)