Scaling Graph Analytics Without ETL: Inside PuppyGraph’s Architecture

Scaling Graph Analytics Without ETL: Inside PuppyGraph’s Architecture

May 31, 2026 · 54 min · Episode 510

About this episode

Weimo Liu discusses the architecture of PuppyGraph's zero-copy graph querying engine and its applications in various data sources.

Summary In this episode Weimo Liu, co‑founder of PuppyGraph, talks about the engineering behind their “zero-copy” graph querying engine for lakehouse and database sources. He explores how PuppyGraph lets you run Cypher and Gremlin traversals and graph algorithms directly on data in Iceberg, Delta, Hudi, Hive, and even MongoDB—without loading into a separate graph store. Weimo explains their edge-sharded, vectorized, MPP architecture that tackles hub nodes, multi-hop traversals, and shuffle at scale, targeting sub-second to single-digit-second workloads. He digs into practical graph data modeling on top of normalized and denormalized tables, logical views, and flexible mappings; strategies for caching, adaptive reads, and leveraging Iceberg metadata; and how PuppyGraph’s operator-based engine unifies query and algorithms. He also covers real-world applications—from cybersecurity log analysis to entity resolution and agentic workflows—when to choose embedded or transactional graph databases instead, and what’s next for enterprise features and broader warehouse integrations. Announcements Hello and welcome to the Data Engineering Podcast, the show about modern data management This…

People in this episode

Host: Tobias Macey

Guest: Weimo Liu

Topics covered

Keywords

Sponsors

DataDriven.io

Mentioned in this episode

Organizations: PuppyGraph, Iceberg, Delta, Hudi, Hive, MongoDB

Explore listener stats, chart rankings, contacts and more on the Data Engineering Podcast podcast page.