Diffbot Knowledge Graph

Diffbot Knowledge Graph is a machine learning-curated database of 10+ billion interconnected entities—people, companies, products, articles, and discussions—containing over 1 trillion facts extracted from the public web.

Reviewed by 7wData

On this page

Publisher review

Diffbot Knowledge Graph is a machine learning-curated database of 10+ billion interconnected entities—people, companies, products, articles, and discussions—containing over 1 trillion facts extracted from the public web. Unlike search engines that return websites, Diffbot returns structured business intelligence: a company record includes 50+ fields covering financials, employees, competitors, and relationships. The system rebuilds itself every 4-5 days with fresh data sourced across all languages and geographies.

Users query via Diffbot Query Language (DQL), supporting both visual builders and code-based filters to find, rank, and analyze billions of linked entities. The Knowledge Graph powers three primary use cases: lead research (finding companies and decision-makers with specific traits), competitive intelligence (tracking organizational changes and market moves), and data enrichment (adding provenance-backed facts to existing CRM or analytics systems). Diffbot's core differentiator is autonomy—the graph updates itself via machine learning without human editors, unlike Wikipedia or Google's knowledge graphs.

This cuts latency but also means data reflects whatever patterns exist in web content at scale; bias in source material can propagate. Enterprise teams use it embedded in Excel, Tableau, Power BI, and Zapier. The product sits in a middle ground: more structured than raw web scraping, cheaper per query than traditional B2B data vendors, but less finalized than hand-curated databases. Pricing is consumption-based, making costs predictable for steady-state research but unpredictable during exploratory phases.

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

How it works

  1. 10+ Billion Entities with 50+ Fields

    Each entity (person, company, product, article) contains 50+ data fields including employment history, financials, location data, relationships, and metadata sourced from web-wide extraction.

  2. Diffbot Query Language (DQL)

    Native query interface supporting both visual builders and SQL-like syntax to filter, sort, facet, and analyze billions of interlinked entities without switching tools.

  3. 4-5 Day Refresh Cycle

    Fully autonomous machine learning system rebuilds the entire knowledge graph every 4-5 days, capturing new entities and facts without manual intervention.

  4. Multilingual Web Crawl

    Extracts and indexes structured data across French, Chinese, Cyrillic, and English-language web content with equal accuracy and coverage.

  5. Direct Platform Integration

    Native connectors to Excel, Google Sheets, Tableau, Power BI, Zapier, and REST API enable dataset creation and embedding without manual export/import.

  6. Entity Linking and Relationship Mapping

    Automatically identifies and links related entities across the graph, surfacing connections between companies, people, and technologies that separate databases miss.

  7. Web Data Extraction APIs

    Companion extraction tools allow users to convert unstructured web pages into structured JSON, complementing the pre-built knowledge graph for custom data needs.

Strengths and trade-offs

Strengths

  • Autonomous machine learning curation removes human bias and editorial lag; the system updates on a 4-5 day cycle across 10+ billion entities.
  • Significantly cheaper per-query than traditional B2B data vendors (Dun & Bradstreet, ZoomInfo) for basic research; startup tier at $299/month includes 250K API calls.
  • Multilingual coverage and real-world entity focus (local suppliers, colleagues, SMBs) go beyond Google or Wikipedia's public-figure emphasis.

Trade-offs

  • Credit-based pricing creates uncertainty at scale—overage costs ($0.0009/call) can balloon during exploratory analysis or bulk enrichment runs, unlike flat-rate competitors.
  • Algorithmic biases in source data propagate directly into results; no human verification layer means inaccurate web pages can pollute the graph until the next rebuild cycle.
  • UX and DQL learning curve deter non-technical users; competitors like Bright Data or import.io offer visual builders that require less SQL fluency.

Pricing context

Diffbot offers four plans: Free ($0/month, 10K credits, 5 calls/min), Startup ($299/month, 250K credits), Plus ($899/month, 1M credits), and Enterprise (custom). All plans charge $0.0009 per overage credit. A credit represents one API call or one Knowledge Graph query; data extraction, crawling, and enrichment consume credits at different rates.

No per-seat licensing; pricing is pure consumption-based. Free tier and Startup tier target small teams and projects; Plus and Enterprise are for organizations running continuous research or bulk enrichment. The credit system rewards steady-state usage but penalizes exploratory spikes—a month of light research may cost $50, but a bulk company enrichment project may cost thousands if not budgeted carefully.

Alternatives

User reviews

No user reviews yet. Be the first to write one.

Sources

Reporting on this tool draws on these publicly available sources.

  1. www.diffbot.com — Core product description, 10+ billion entities, 50+ fields, DQL query interface, entity types (people, companies, products, articles, discussions)
  2. www.diffbot.com — All pricing tiers (Free, Startup $299/month, Plus $899/month, Enterprise), credit rates ($0.0009 per overage), API call limits (5/min, 5/sec, 25/sec), monthly credit allowances
  3. blog.diffbot.com — Product announcement and overview, founding date (2018), multilingual support, autonomous curation approach, 1+ trillion facts, use cases (people, companies, locations, articles, products, discussions, images)
  4. blog.diffbot.com — Technical approach, machine learning curation method, 4-5 day rebuild cycle, web-scale crawling, autonomous system design, comparison to human-curated alternatives (Google, Wikipedia)
  5. www.historytools.org — Competitive landscape, alternatives including Bright Data, ScraperAPI, Scrapeless, KnodeGraph; credit-based pricing criticism
  6. craft.co — Company founding (2011), founder/CEO (Mike Tung), headquarters (Menlo Park, CA), employee count (35), funding ($12.5M from Felicis Ventures, Tencent, Bloomberg Beta)