Podcast profile
Data Engineering Podcast
By Tobias Macey
This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.
Latest episode
What Context Really Means in Data Engineering and AI
53 min
- Cadence
- Weekly
- Typical length
- 54 min
- Latest release
- Sep 15
- Language
- EN
- Activity
- Current — released within 30 days
Podcast Classification
Topics
Building and operating data pipelines, transformations, orchestration, and data platforms.
Distributed data systems, warehouses, lakes, streaming, and large-scale data platforms.
Database systems, storage engines, query processing, transactions, and database operations.
Formats
Host-led conversation where a guest supplies most of the subject matter.
Detailed technical, architectural, or research-oriented examination.
Audience level
Regularly assumes specialist knowledge and discusses implementation, architecture, protocols, code, or research in depth.
People behind the show
Hosts
Stay connected
From the feed
Latest episodes
Episodes shown — 10 of 25
What Context Really Means in Data Engineering and AI
Summary In this episode Soham Azumdar, co-founder and CEO of Wisdom AI, talks about what “context” really means in data engineering and AI systems.
Specialized AI for Data Engineers: Inside Astronomer’s Otto
Summary In this episode Yetunde Dada discusses Otto, Astronomer’s AI agent for Airflow, and the broader challenge of making agentic tooling actually useful for data engineers.
Why Multi-Agent Systems Need Shared State, Graph Semantics, and Governance
Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems.
Building the Context Flywheel for AI Data Agents
Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations.
Holding Kafka Right: Product-Friendly Streaming with TypeStream
Summary In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka.
Text to Data Products: Kaarvi’s End-to-End AI for Ingestion, Quality, and Dashboards
Summary In this episode Shravan Gunda, founder and CEO of Kaarvi AI, talks about building an AI-native, agent-driven data platform designed to eliminate the janitorial work that consumes most data teams.
Scaling Graph Analytics Without ETL: Inside PuppyGraph’s Architecture
Summary In this episode Weimo Liu, co‑founder of PuppyGraph, talks about the engineering behind their “zero-copy” graph querying engine for lakehouse and database sources.
Maximizing GPU Utilization: Heterogeneous Pipelines with Ray and Kubernetes
Summary In this episode Robert Nishihara, co-founder of Anyscale and co-creator of Ray, talks about maximizing hardware utilization for AI and data-intensive workloads.
The AI-First Data Engineer: 10–50x Productivity and What Changes Next
Summary In this episode, I sit down with Gleb Mezhanskiy, CEO and co-founder of Datafold, to explore how agentic AI is reshaping data engineering.
Treat Metering Like Finance: Building Data Platforms for Consumption Economics
Summary In this episode Himant Goyal, Senior Product Manager at Salesforce, talks about how data platform investments enable reliable, accurate metering for consumption-based business models.