Skip to content
Core Digital Technology

Top 10

Data, Databases & Analytics podcasts

Data engineering, analytics, data platforms, databases, governance, and data architecture.

Sector list

Top 10 Data, Databases & Analytics Podcasts of 2026

Showing 10 of 10

Listen now

Latest in Data, Databases & Analytics

All latest episodes
The Data Stack Show49 min

Re-Air: AI, Abstractions, and the Future of Data Engineering with Pete Hunt of Dagster

This week on The Data Stack Show , Brooks and John welcome Pete Hunt, CEO of Dagster Labs. During this conversation, Pete takes listeners on a fascinating journey through the evolution of data platforms, sharing insights from his experiences at Facebook and X (Twitter). Pete discusses the critical challenges of managing data complexity, emphasizing the importance of software engineering best practices and strong abstractions in data orchestration. He explores how emerging technologies like AI are transforming data workflows, breaking down traditional team boundaries, and enabling more collaborative, self-service data engineering. The conversation also highlights Dagster's innovative approach to building a data control plane, with a focus on empowering different stakeholders through flexible, composable tools. Listeners will gain valuable perspectives on the future of data engineering, the role of AI in development, the importance of creating adaptable, user-friendly data infrastructure, and so much more. Highlights from this week’s conversation include: Pete's Background and Journey in Data (1:36) Evolution of Data Practices (3:02) Integration Challenges with Acquired Companies (5:13) Trust and Safety as a Service (8:12) Transition to Dagster (11:26) Value Creation in Networking (14:42) Observability in Data Pipelines (18:44) The Era of Big Complexity (21:38) Abstraction as a Tool for Complexity (24:41) Composability and Workflow Engines (28:08) The Need for Guardrails (33:13) AI in Development Tools (36:24) Internal Components Marketplace (40:14) Reimagining Data Integration (43:03) Importance of Abstraction in Data Tools (46:17) Parting Advice for Listeners and Closing Thoughts (48:01) The Data Stack Show is a weekly podcast powered by RudderStack, customer data infrastructure that enables you to deliver real-time customer event data everywhere it’s needed to power smarter decisions and better customer experiences. Each week, we’ll talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data. RudderStack helps businesses make the most out of their customer data while ensuring data privacy and security.

It’s About Data18 min

Context Engineering and Data Products (Justin Borgman, CEO of Starburst)

Matt's book Fundamentals of Data Engineering is available through the O'Reilly Learning platform. Use the code MATT20 for a 20% discount on a one year subscription. We welcome Justin Borgman, cofounder and CEO of Starburst, back to the show to discuss data and context. Justin argues that context engineering is a human in the loop process.

DataFramed51 min

#378 The Data Engine for AI with Ledion Bitincka, CTO at Cribl & Nikhil Mungel, Head of AI R&D at Cribl

As agents take over more of the actual coding, the nature of technical work is shifting from producing software to judging it. Engineers increasingly spend their time specifying what should be built and then checking whether an agent's output actually solved the problem, rather than writing every line themselves. That changes what skills matter for a career in data and software: problem-solving and evaluation start to outweigh knowing today's specific tools or syntax. It also raises a harder question for teams — if agents can generate work this fast, how do you know which of it is actually worth shipping? Ledion Bitincka is Co-Founder and CTO of Cribl, where he leads the engineering organization with a first-principles approach to product delivery. Before Cribl, he was an Advanced Development Architect at Splunk, where he worked on Search-Time Schema, Hunk, and SmartStore, and before that founded Triangulus Communications. Nikhil Mungel is Head of AI R&D at Cribl, based in San Francisco, with over 15 years building distributed systems and AI teams at companies including Substack, Splunk, and ThoughtWorks — he now leads teams building LLM-powered systems for IT and security data. In the episode, Richie, Ledion, and Nikhil explore AI agent disasters and cost shocks, software telemetry fundamentals, using agents to analyze telemetry, AI-powered software factories, the shift from knowledge work to judgment work, skills for the agentic era, avoiding runaway AI spend, and measuring product value through growth metrics, and much more. Links Mentioned in the Show: • Jocko Willink's book, Leadership Strategy and Tactics: A Field Manual • Cribl • Cribl AI (Copilot) • Claude Code • AWS Graviton • NVIDIA • Connect with Ledion: LinkedIn • Connect with Nikhil: LinkedIn • AI-Native Course: Intro to AI for Work • Related Episode: AI Agents at Work: What Actually Breaks (and How to Fix It) New to DataCamp? Learn on the go using the DataCamp mobile app Empower your business with world-class data and AI skills with DataCamp for business

DataTalks.Club1 hr 3 min

Decoding the Open Lakehouse - Sahil Walia

In this talk, Sahil Walia, Senior Technical Architect at Snowflake, shares his extensive data engineering expertise from writing simple SQL queries early in his career to architecting enterprise-scale Apache Iceberg lakehouses today. We explore the evolution of the modern data platform, the mechanics of open lakehouses, and how decoupling storage from compute is transforming data architecture. TIMECODES: 00:00 Defining the Modern Data Platform 11:34 What Exactly is a Lakehouse? 17:20 Comparing Hive Metastore and Apache Iceberg 22:34 Handling GDPR and Data Deletions in Iceberg 27:33 Apache Iceberg vs. Delta Lake 34:10 Do You Need a Lakehouse for Small Data? 40:11 Understanding Apache OSI (Open Semantic Interchange) 49:00 Managing Schema Changes in Data Warehouses 54:20 How to Contribute to Open Source Projects 01:00:40 The Role of AI in Open Source Contributions This talk is designed for data engineers, data architects, and analytics professionals looking to modernize their data infrastructure. It provides essential architectural insights for technical teams evaluating Apache Iceberg, dealing with large-scale data transformation, or aiming to scale their open-source contributions.

The Data Flowcast20 min

Orchestrating data across 30 companies at itti

Orchestrating data across more than 30 companies means most Airflow users on the platform aren't data engineers. In this episode, Kenten Danas talks with Lucas Trubiano , Data Engineer at itti , the technology company within Grupo Vázquez in Paraguay. Lucas walks through the custom YAML framework his Center of Excellence built on top of Airflow, how they baked data quality and custom operators into it, and how a spec-driven AI workflow now lets product and business users contribute to templates without knowing Python. Key Takeaways: 00:00 Introduction. 01:47 What itti and Grupo Vázquez do, and the Data Engineering Center of Excellence's mandate to build a 360-degree view of the customer across more than 30 companies. 02:56 How Airflow fits in as a central task orchestrator (not a processing engine) across around 300 production DAGs. 04:25 Managing enterprise-scale Airflow: preferring Airflow-as-a-service, plus enabling self-service for non-technical users through YAML. 05:44 Why itti built a second, more opinionated YAML framework after DAG Factory-style customization created a code review bottleneck. 07:14 More than 80% of new DAGs are now created with the new framework because it's simply faster. 07:54 How the framework works end to end: Python DAGs, Jinja templates, YAML configs, and CI/CD compilation. 09:10 A Google Sheets ingestion example that shows how prevalidation, download, and processing tasks are hidden behind a simple YAML config to preserve reliability. 10:14 Building an in-house data quality tool that tests per partition instead of full-scanning tables, triggered via custom Airflow operators. 11:17 Custom operators for dbt, the in-house data quality tool, and AWS services like QuickSight dashboard refreshes, and how OSS Airflow makes them portable across instances. 15:35 Spec-driven development with a fork of GitHub Spec Kit so business users can describe what they want and let agents generate DAGs against certified templates. 17:24 Slack-native error routing: every DAG has an owner team, common errors ship with explanations, and only deep issues escalate to the central team. 19:29 Where they're heading next: agent-triggered pull requests for self-healing pipelines. 20:22 Airflow 3 wishlist: backfill improvements, event-driven orchestration for streaming pipelines, and Human in the Loop for generative AI DAGs. Resources Mentioned: Apache Airflow Apache Spark DAG Factory GitHub Spec Kit Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations. #AI #Automation #Airflow

The Data & AI Chief37 min

AI Literacy Starts in the C-Suite

Most AI playbooks are built for a workforce tethered to a desk, but many workers are not near a keyboard. In this episode, Beth Miles, Chief AI Officer at Advantage Solutions, breaks down how she's building an AI-ready culture across a distributed workforce of more than 60,000 teammates. She shares why AI literacy has to start with leadership, why redesigning a broken process matters more than simply automating it, and how she measures AI success well beyond adoption metrics. Key Moments: From Chief of Staff to Chief AI Officer (05:00) : Beth explains how her chief of staff role gave her the cross-functional view needed to build Advantage's AI office. Hands-On Keyboards: AI Training in the C-Suite (11:00) : Beth breaks down 10+ hours of executive AI training, including a CEO who sat down for a coding workshop. Inside the AI Champions Build Summit (22:00) : Over 40 teammates spent five hours building AI use cases together, breaking down silos across functions. Why You Shouldn't Optimize a Bad Process (26:00) : Redesigning broken processes before automating them matters more than speed. Adoption Is Step One, Not the Finish Line (31:00) : Usage is only step one. Beth points to service excellence and time to complete as the metrics that matter. Key Quotes: “AI can be an accelerant, but you have to have all the mechanics and the process and the data to actually make it run really well.” - Beth Miles “It may require moving slower now to move faster later, to actually redesign and rethink work so you’re not optimizing a bad process.” - Beth Miles “ To lead a change, we need to understand it.” - Beth Miles Mentions: AI Transformation Requires Redesigning Work, Not Cutting Roles East of Eden by John Steinbeck Shoe Dog by Phil Knight Guest Bio: Beth Miles leads Advantage Solutions’ AI strategy, enablement, and governance to drive value for the company’s clients and teammates. She drives AI deployment across the enterprise, alongside the Data Office and Growth & Strategy Office. Beth joined Advantage in early 2024 and has served the last year as Chief of Staff, leading the company’s AI workstreams from diagnostic through scaled execution, driving enterprise transformation initiatives, and overseeing CEO Office operations. Before Advantage, Beth spent more than a decade leading technology and corporate transformation at global software firms, including senior roles at Sage and Cerner (now Oracle Health). At Cerner, she drove the redesign of Cerner’s portfolio management process for their $400m R&D investment. At Sage, she led GTM execution across new revenue streams and managed a large cloud partnership, leading cross-functional teams across governance and delivery. Hear more from Cindi Howson here .

Data Engineering Podcast53 min

What Context Really Means in Data Engineering and AI

Summary In this episode Soham Azumdar, co-founder and CEO of Wisdom AI, talks about what “context” really means in data engineering and AI systems. He explores why context has become such an overloaded term, spanning everything from semantic layers and data catalogs to tribal knowledge, query logs, dashboards, and even agent memory. Soham explained that the big shift is that context is no longer being prepared primarily for human analysts, but for LLMs and agents that can’t reliably fill in missing gaps on their own. That change raises the bar for how context is represented, validated, benchmarked, and maintained so that AI systems can produce trustworthy outcomes. Announcements Hello and welcome to the Data Engineering Podcast, the show about modern data management Your host is Tobias Macey and today I'm interviewing Soham Mazumdar about what "context" actually means in data engineering Interview Introduction How did you get involved in the area of data management? One of the perennial challenges of engineering in all forms is building a shared understanding of what a given word means. "Context" is one that is being used for an increasing number of purposes with the introduction of AI agents. Can you start by sharing some of the ways that this terminology overload has caused problems in your own experience? Data engineering has arguably always been about context engineering, but at the scale of human consumers. What are the substantive changes that AI/agentic consumers bring to the discipline? While we all understand the notion of "context", turning it into a useful and re-usable component is a different matter entirely. What are some of the ways that "business context" or "technical context" manifests as a tangible artifact? This also brings up the question of data modeling. What are some of the key attributes that are necessary when storing, enriching, evolving, and joining into that context? One could argue that the entire history of data warehousing is about building organizational context. What are the real differences in approach for today's work of capturing and activating that context? How does your work at Wisdom AI address the technical and operational burdens of capturing, modeling, and exposing context at the speed necessary to keep up with organizational demands? What are the most interesting, innovative, or unexpected ways that you have seen Wisdom AI used? What are the most interesting, unexpected, or challenging lessons that you have learned while working on Wisdom AI/context engineering? When is Wisdom AI the wrong choice? What do you have planned for the future of Wisdom AI? Contact Info LinkedIn Parting Question From your perspective, what is the biggest gap in the tooling or technology for data management today? Closing Announcements Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems. If you've learned something or tried out a project from the show then tell us about it! Links Wisdom AI Context Engineering Knowledge Graph Ontology Snowflake Open Semantic Interchange (OSI) Semantic Layer Palantir Foundry The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

You don't have ICs anymore, you have managers (Emilie Schario)

Emilie Schario is co-founder and head of product and engineering at Kilo Code, the open-source, model-agnostic coding agent platform Anaconda acquired in July 2026. She joins Tristan Handy to talk why nobody on her 20-person engineering team is really an individual contributor anymore, and why she thinks data people are better prepared than software engineers for a world where everyone manages a portfolio of agents. The Analytics Engineering Podcast is sponsored by dbt Labs.

Data Skeptic23 min

Recommender Systems Today and Tomorrow

In the final episode of our Recommender Systems season, we explore the growing questions of trust, manipulation, privacy, fairness, sustainability, and user control. From fake reviews and shilling attacks to explainable recommendations and user-selected algorithms, we look at what happens when recommender systems must answer not only for what they recommend, but for the consequences of those choices.