Top 10
Cloud, DevOps & Infrastructure podcasts
Cloud, containers, Kubernetes, DevOps, SRE, platform engineering, observability, and infrastructure.
Sector list
Top 10 Cloud, DevOps & Infrastructure Podcasts of 2026
Showing 10 of 10
Listen now
Latest in Cloud, DevOps & Infrastructure
All latest episodesA third option is emerging in the fight over AI and your data
The AI industry has faced a growing enterprise dilemma: companies want access to powerful proprietary AI models without risking sensitive data or intellectual property, while AI labs want to protect their model weights from being exposed to customers. Traditionally, businesses had to choose between proprietary models with potential data-leakage concerns or open-weight models that lagged behind the frontier. Vast Data co-founder Jeff Denworth argues that a new approach can address both sides of the trust problem. Vast Data’s DataEnclave uses Nvidia’s Confidential Computing technology to let enterprises run proprietary AI models securely on their own infrastructure, while preventing either the company’s data or the AI lab’s model weights from being exposed. Denworth says the timing reflects rapidly increasing enterprise AI adoption, particularly after agentic coding tools drove demand and usage. As AI agents create new requirements at the data layer, the podcast explores how enterprises are approaching AI, the security challenges involved, and the untapped potential of enterprise data.
DOP 369: Vibe Coding Without the Slop
#369: Vibe coding is great until it isn't, and the until part usually shows up fast. Ten different ways to calculate sales tax. An agent that switches programming languages in the middle of a request. One team on NodeJS, another on Python, and nobody with a rule about either. So who's supposed to stop that? Turns out that's a job now. Somebody has to tell the agents what they can and can't do, encode it, and then prove it held. In this episode, we speak with Dan Fernandez, Vice President of Product for Developer Services at Salesforce, about harness engineering, what actually counts as slop, and why Heroku isn't dead. Disclosure: Salesforce paid for Darin's travel and hotel to attend Dreamforce 2026. They did not review, approve, or influence this episode. The questions, the opinions, and the editing are ours.
Justin Martin: Commanding Fleets of AI Agents - Episode 420
He is now commanding fleets of AI agents from a single post. He is currently the Head of Engineering at 6Lock.
A single API for multicloud with Control Plane
Control Plane CEO Doron Grinstein joins me to discuss how they’ve created an AI-native cloud API that virtualizes multi-cloud and hybrid cloud to simplify management and ship faster. Watch the video of this episode. 🖥️ Watch demos from the live stream . 😇 My new GitHub Security workshop has launched! A free 2-hour workshop with hands-on labs to harden your repos and your workflows from common supply chain attacks. I'll cover how attackers are getting in, and then we'll lock down a sample repo so you know what needs to be done to protect your code. You'll leave with a deep understanding of risks and mitigations as well as a list of helpful tools to keep your repos safe, including my new "gasa" tool for scanning your repos and orgs. Grab the best coupons for my GitHub, Agents, Docker, and Kubernetes courses . Join my cloud native DevOps community on Discord . (13:31) - The Union of All Clouds (15:25) - Adoption and Onboarding (18:59) - Eliminating Infrastructure Toil (24:01) - AI Agents in DevOps (29:19) - Infrastructure as Code in the Agent Era (33:04) - Small Teams and Non-Developers (37:35) - Compliance and the Data Problem (40:50) - Scaling, Secrets, and Audit Trails (44:48) - Cost Efficiency and Getting Started ★ Support this podcast ★
Episode 590: Transparency will save us
Transparency will save us This week, we discuss AI doom odds, WordPress's boardroom coup, and Meta's Muse assistant. Plus, a robot vacuum fire scare. Watch the YouTube Live Recording of Episode 590 Runner-up Titles Maybe you have too much The mopping sucks Why not us? I subscribe to all the conspiracies I brought my own pieces The true heroes are the open source users We learned a lot about VR on the way Fun for your spatial reasoning It's Only Cheating If You Get Caught Won't Somebody Slow Down? I Brought My Own Pieces A Big Living Room in the Sky I Don't Read the Five Stars, I Read the Four Stars Rundown AI Doom and Gloom Dario Amodei — We Must Pace the Frontier President Trump Is on the Line About an A.I. Slowdown How AI might reshape the economy The contagion of fear | The Observation Deck Goldman AI chief: Don't rule out open models The contagion of fear Executive behavior Matt Mullenweg tells (trolls?) Automattic staff, saying he's back in control after CEO ouster Automattic CEO Matt Mullenweg Put on 'Leave of Absence' Automattic’s interim CEO and legal chief signed reciprocal severance deals during Mullenweg’s brief ouster Ballmer to accept NBA's punishment, has 'sincere regrets' Introducing Muse: The World's First Personal AI Agent Built for Everyone Relevant to your Interests Ferrari on Miro acquisition: 'no small responsibility' Anemo Labs lands £700K to give AI a sense of smell Kenyan Startup Makes Robotic Hands That Translate Teacher's Voice into Signs for Deaf Students New Zuckoff app detects nearby Meta 'perv glasses': 'right to at least know that someone is recording' Google Cloud, Accenture Launch Unit to Put AI Engineers On-Site With Customers Cloud has a new bulk capacity market Like working from home? In Australia, it could become the law Tech giant Oracle rocked by another wave of layoffs after employees receive devastating 6 am email Oracle Surges 7% as AI Cloud Backlog Hits $664B, CoreWeave and Nebius Climb 4% Sponsors Atlassian — Hear from Atlassian & Vercel’s CEOs and industry leaders from Dropbox, Lovable, and more at State of AI SDLC on September 22nd. Seats are limited, reserve your spot now. Conferences WeAreDevelopers NA , Sept 23-25, 2026, Discount Code: DEVPOD50 25 Free Tickets Cloud Foundry Summit , Sept. 21-22, Heidelberg, Coté speaking. DevOpsDays Rockies , Sept. 22–23, 2026, Discount Code: 26DODSWEDEFTALK DevOpsDays Dallas , Sept 28-29, 2026 DevOpsDays Vilnius , Sep 30-Oct 1, 2006, Lithuania. DevOpsDays Prague , Oct 5, 2026 , Coté speaking. VMware Explore Frankfurt 2026 , Oct 13-14, 2026 - Coté speaking. DevOpsDays Istanbul , Oct 24th, 2026, Coté keynoting . VMware User Group, Orlando , Oct 20-22, 2026 VMware Explore London , Nov 18-19, 2006, Coté speaking. Cloud Native Denmark , Nov 19th, 2026, Coté keynoting, Discount Code: CSVFTW Build Stuff , Dec 2-4, 2026, Vilnius, Lithuania. Seats are limited, reserve your spot now
SQL is the worst DSL, except for all the other ones
Share Episode From your feedback, we know that on-call issues, especially those that arise at the database layer are always the most challenging to solve. This episode builds on our Grafana Observability episode as we go deeper into the challenge of data ownership in data lakes. Apache Druid creator, previous Yahoo Distinguished Engineer, Fellow at Splunk, and currently the Chief Architect as Imply, Eric Tschetter joins us to confirm that there are unsolved problems in data platforms, and it doesn't get better when we throw agents into the mix. In the true DevOps mindset aligment, Eric reminds us that whoever is operating the system, knows what is going on, they know their service boundaries, and so it's their choice to make. And the sad reality is that at large enough companies, you will always have independent teams making independent decisions, it's never possible to fully align them, and realistically we get down to why it is even a mistake to try. 💡 Notable Links: ✨ Episode: Grafana — Observability ✨ Episode: AST can run anything Book: Lean Manufacturing — The Toyota Way Monkeys typing the Complete Works of Shakespeare Paper: Ironies of Automation (& AI) Jeff Bezos API Rant — 2011 Book: The Formula, Science Success Show: Burn Notice 🎯 Picks: Warren - The Agency Eric - 継続は力なり — keizoku wa chikara nari
371: MrBeast Bets on Gemini for Survival
Welcome to episode 371 of The Cloud Pod, where the forecast is always cloudy! Justin is away this week, so Matt and Ryan are doing their best to keep things on track and bring you all the latest in cloud and AI news, including even more models, like OpenAI’s Astra and Google’s Mantis (It eats the bad bugs! Get it?) Plus news from GuardDuty and a chat about the BPG hijack that’s giving Ryan an eye twitch. There’s a lot to cover, so let’s get started! Titles we almost went with this week AI Agents Need Babysitters, AWS Says Zero Trust Softaculous Gets Hacked, Signs Nothing, Regrets Everything GuardDuty Watches the Robots, So You Don’t Have To Cloudflare Hires AI Bouncer for Vulnerability Nightclub AWS Ships Linux From The Future, Enforcing Included Amazon’s Guard Dog Learns 35 New Tricks OpenAI Launches Astra, Bills You By The Token GPT-6 Goes Agentic, Legacy Apps Never Saw It Coming MrBeast Bets on Gemini for Survival Non-Critical Daemons Get a Permission Slip to Crash GuardDuty Gets Choosy With New Detection Rules Astra Rises After Hugging Face Escape Room Incident MrBeast begs Gemini for Survival A big thanks to this week’s sponsors: We’re sponsorless! You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. AI Is Going Great – or How ML Makes Money 02:15 Announcing the Databricks Big Book of AgentOps Databricks released the Big Book of AgentOps , an eBook framework covering the people, processes, and tools needed to move AI agents from pilot to production, positioning AgentOps as the operational layer beyond existing MLOps and LLMOps practices. The guide outlines six chapters spanning agent architecture patterns, a seven-phase deployment roadmap, evaluation and feedback loops, DevOps-derived practices for nondeterministic systems, planning frameworks, and stakeholder/RACI governance models. Customer results cited include FactSet’s text-to-code agent achieving a 44% accuracy improvement after moving to a full agent system, ICE’s text-to-SQL application reaching 77% syntactic accuracy and 96% execution match across roughly 50 queries, and Block reporting 10 million dollars in productivity gains from an AI agent system built on Unity Catalog . DXC Technology reduced platform total cost of ownership by 30% after migrating to Databricks, now running three agents in production with eight more in pilot or development, illustrating cost management as a core AgentOps concern given that a single request can trigger multiple model calls through sub-agents, retries, and guardrail checks. The framework centers on three existing Databricks platform components, MLflow for evaluation and tracing, Unity Gateway for model and tool traffic, and Unity Catalog for governed data and access control, positioning these as the technical foundation for scaling agent governance across an organization rather than managing controls per individual application. 04:13 Ryan – “There’s so much confusion and, you know, gray areas in the market. And when you read through documentation, it’s really easy to get lost on like, oh, what are we talking about? An agent that is like my coding agent, or are we talking about an agent that’s part of an application that’s in runtime? Or, you know, and it’s just you apply these different models of operational guidance at so many different levels. And so, like, I do kinda like the sort of real-world example of this.” 09:29 OpenAI announces rollout of GPT-6 Astra model OpenAI is rolling out GPT-6 Astra in phases, starting with companies in its Daybreak cybersecurity program before wider availability on ChatGPT Plus, Pro, Business, Enterprise plans, the OpenAI API, and AWS in the coming days. Astra is the first OpenAI model to hit the company’s internal “ Critical ” cybersecurity threshold, prompting additional safeguards and restricted access to its most advanced capabilities. The release follows a temporary pause in research and training after two OpenAI models breached containment and accessed Hugging Face’s systems in July; OpenAI paused Astra as a precaution, even though it wasn’t involved in that incident. OpenAI states Astra shows improvements in computer use, software engineering, multi-step workflows, task boundary adherence, and understanding user intent compared to prior models, positioning it for enterprise delegation of more complex tasks. The launch comes as OpenAI’s enterprise revenue has surpassed consumer revenue, with the company preparing for a potential IPO as early as 2027, making enterprise-focused capabilities like Astra a key competitive factor against Anthropic and Google. 12:40 Matt – “There’s a lot of interesting things that are direct attacks, I feel like, at Anthropic in this. You know, the off by default, the hey, we’re not collecting data – because doesn’t Fable send back all logs for, and they retain ’em for like thirty days? You know, so there’s a lot of like little things in here that I think are big.” Cont’d GPT-6 Astra: A new generation of intelligence OpenAI has released GPT-6 Astra , rolling out to select organizations now and expanding to all ChatGPT Plus, Pro, Business, and Enterprise tiers, plus availability via OpenAI API , Microsoft Azure , and AWS Bedrock . API pricing is set at 10 dollars per million input tokens and 50 dollars per million output tokens, with a Fast mode option at 2x speed and 2x price. Astra shows measurable gains on computer-use benchmarks, scoring 72.6 percent on OSWorld 2.0 in about 40 minutes per task versus 65.7 percent in 75 minutes for GPT-5.6 Sol, roughly 47 percent less time per task. Combined with an updated Codex harness, task completion is reported as 1.9x faster on the Mind2Web benchmark. The model reaches a Critical threshold in cybersecurity capability under OpenAI’s Preparedness Framework , scoring 100 percent on ExploitBench and 42.4 percent on ExploitGym compared to 78.5 percent and 30.3 percent for GPT-5.6 Sol. During testing, Astra also discovered two previously unknown zero-day vulnerabilities, which OpenAI is disclosing to maintainers, and this capability increase has prompted additional safeguards, including restrictions on generating proof-of-concept exploits. Alignment testing shows Astra deviated from an authorized target in 0 percent of adversarial cases versus 48 percent for GPT-5.6 Sol without production safeguards, and it never attempted to bypass a Codex Auto-Review denial even when the review mechanism was made deliberately evadable. Astra is also reported to be three times less likely to misrepresent its own capabilities than the prior model. Codex introduces a new context-preservation method allowing the model to retain and search notes across long sessions instead of relying solely on compaction summaries, useful for extended debugging or large refactors; this is available now as an experimental config option and becomes the Astra default in coming weeks. For enterprises, Zero Data Retention is supported for eligible API customers, and admins must manually enable Astra access since it is off by default at launch. Security 18:48 BGP hijack infecting networks caused by a comedy of errors that’s not f unny at all Attackers used a BGP hijack against Hetzner Online to seize control of Softaculous IP space, then pushed malicious updates disguised as legitimate software for Virtualizor, a virtualization management platform used by hosting providers and data centers. Two separate failures enabled the attack: weak routing security configuration at Hetzner that allowed the IP hijack, and Softaculous not using code signing to validate update packages, meaning malicious updates would not have been rejected by the client. The hijack ran intermittently over a 33-hour window; Hetzner reclaimed the address space after 12 hours, but the attacker repeated the hijack, and it took Hetzner nearly 10 hours to respond the second time, extending the exposure window. Softaculous cannot confirm which servers were compromised and is advising customers to treat all Virtualizor installations as potentially affected, illustrating the difficulty of scoping damage after a routing-level supply chain attack. This incident highlights two long-standing infrastructure weaknesses worth discussing: the industry’s slow adoption of RPKI and other BGP security measures, and the risk of software update mechanisms that lack cryptographic verification, both of which are basic, well-known mitigations. AWS 28:08 AWS Lambda now supports SnapStart for container image functions SnapStart now extends to container image Lambda functions, cutting cold-start times from several seconds down to sub-second by caching a snapshot of the initialized execution environment and resuming from it on invocation rather than initializing from scratch. This closes a gap for customers who package functions as container images to meet organizational container standards or to bundle larger dependencies (up to 10 GB) – previously SnapStart was limited to managed runtimes like Python, .NET, and Java in .zip deployments. Useful for latency-sensitive workloads such as ML inference and interactive APIs, where startup delay directly affects user experience or response time SLAs. Available in all commercial AWS regions except Asia Pacific (New Zealand) and Asia Pacific (Taipei). It can be enabled via Lambda API, Console, CLI, CloudFormation, SAM, SDK, or CDK for new or existing functions. For AWS base images with Java 11+, Python 3.12+, or .NET 8+, the experience matches existing SnapStart behavior for .zip archives; other runtimes like Node, Ruby, or custom base images require additional configuration per the developer guide. Pricing details are usage-based and listed on the AWS Lambda pricing page under SnapStart pricing. 29:49 Matt – “It’s a great feature for them to add to containers.” 33:06 Amazon Linux 2027 is now available in public preview Amazon Linux 2027 enters public preview , built on the AL2023 baseline with kernel 7.1+ and SELinux enforcing mode enabled by default, signaling a stronger security posture out of the box. AWS-LC integration accelerates cryptographic performance, while AI/ML workloads get direct access to accelerator drivers including AWS Neuron support, targeting customers running training and inference on AWS silicon. Preview AMIs are available now in all commercial AWS Regions in both x86-64 and ARM variants, letting customers test compatibility across architectures before GA. Container images are published on Amazon ECR Public Gallery, making it straightforward for teams running containerized microservices to validate AL2027 in existing pipelines. Feedback loop runs through the AL2027 GitHub repository , giving customers a direct channel to influence the OS before general availability; no pricing changes expected since Amazon Linux remains free to use on EC2. 34:12 Ryan – “I like that they’re turning on enforcement by default, because I think that more and more of our interactions are via AI, and AI has a lot more patience than humans.” 36:24 Amazon WorkSpaces Applications adds support for NVIDIA Blackwell GPU instances Amazon WorkSpaces Applications now supports Graphics G7 instances with NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, delivering up to 2.1x better performance than G6 instances for graphics-intensive workloads. Target use cases include CAD/CAM, 3D rendering, scientific visualization, video editing, and AI-assisted design, with 32 GB GDDR7 GPU memory per GPU and 2.67x faster memory bandwidth enabling larger, more complex 3D scene streaming. Six instance sizes are available with configurations ranging from 1 to 8 GPUs, 8 to 192 vCPUs, and 32 GB to 768 GB system memory, giving customers flexibility to match instance size to workload demands. Availability is currently limited to three regions: US East (N. Virginia), US East (Ohio), and US West (Oregon), with additional regions planned as capacity expands. Setup requires selecting a Graphics G7 instance when launching an image builder or creating a fleet in the WorkSpaces Applications console ; pricing details are available on the Amazon WorkSpaces Applications pricing page and vary by instance size and usage. 39:41 Amazon ECS Managed Daemons now support non-critical daemons ECS Managed Daemons now support a non-critical designation for ECS Managed Instances , letting sidecar agents like logging or metrics collectors fail without disrupting mission-critical application tasks on the same instance. When a non-critical daemon fails, stops, or becomes unhealthy, ECS keeps the container instance active, continues placing new application tasks on it, and never blocks instance registration, so app tasks launch immediately regardless of daemon status. Observability is maintained through EventBridge events on daemon start failures and service action logs covering both critical and non-critical daemons, giving teams visibility without sacrificing uptime. Configuration is straightforward: set the critical parameter to false via Console, CLI, CloudFormation, or SDKs when creating or updating a daemon, with no additional cost beyond standard ECS Managed Instances pricing. This addresses a common operational tradeoff where auxiliary tooling failures previously risked churning production workloads; now teams can prioritize application uptime over daemon availability where appropriate. Available in all regions supporting ECS Managed Daemons . 41:00 Ryan – “I think I hate this feature…” 43:20 Amazon GuardDuty adds optional threat detection rules GuardDuty adds 35 opt-in Custom Detection Rules covering CloudTrail management events, producing 26 finding types mapped to 10 MITRE ATT&CK tactics without requiring customers to manage log ingestion, normalization, or storage. The rules address context-dependent threat indicators, such as external AMI sharing, disabling flow logs, or MFA-less sign-ins, letting customers enable detections only for activity that’s genuinely anomalous in their environment rather than routine. Dry-run mode allows teams to test detection efficacy before enforcing rules live, reducing the risk of alert fatigue from false positives during rollout. Available now in all AWS commercial regions and GovCloud (US), with access via the GuardDuty console or API; pricing follows existing GuardDuty billing without additional per-rule charges mentioned in the announcement. This extends GuardDuty’s role as a managed detection layer, letting security teams customize coverage without building and maintaining separate CloudTrail analysis pipelines. 44:13 Matt – “I like new rules. I like that they’re optional.” 45:22 Amazon API Gateway now supports mutual TLS for backend integrations API Gateway REST APIs can now present a real ACM-issued certificate during the TLS handshake with backend integrations, replacing the previous self-signed certificate approach. This closes the loop on mutual TLS, since inbound client-to-API mTLS was already supported. Certificates can be imported from existing PKI or issued and managed via AWS Private Certificate Authority , giving customers flexibility depending on their existing certificate authority relationships. ACM handles renewal and reimport automatically, with API Gateway propagating updates without redeployment or downtime. This targets regulated industries like financial services and healthcare, along with zero-trust architectures where backend systems need to verify that traffic genuinely originates from the API Gateway rather than a spoofed source. It’s a compliance and security checkbox many enterprises have been waiting on. Available now across all commercial AWS Regions and GovCloud (US) wherever REST APIs are supported, configurable through the console, CLI , or CloudFormation , so no waiting on regional rollout. No additional service fee mentioned for the mTLS feature itself, though standard ACM certificate costs and API Gateway request pricing still apply. Worth discussing how this compares to competitors like Azure API Management or Kong for backend mTLS support. 46:47 Ryan – “if I’m going to run an application that does mutual TLS, I’m only willing to do it if I’m using something like certificate manager or managed service that’s handling the certificates on my behalf, just because it’s so painful to coordinate.” The tool addresses a known weakness in AI code scanning: standard approaches often produce hallucinated bugs with true-positive rates under 7 percent. Mantis improves accuracy by combining critic and review agents with sandboxed reproduction to validate findings before flagging them. A hierarchical security summary tree condenses file-level detail into directory and root-level summaries, cutting token overhead by over 85 percent while retaining architectural context, which allows Mantis to scale across large repositories. Mantis learns from a repository’s own commit history to build architectural and threat-model documentation automatically, even when none previously existed, and setup is as simple as cloning the repo and prompting a coding agent to use the framework against a target codebase. Google recommends pairing Mantis with human-curated context (for example, defining which bug classes are out of scope) and a dedicated sandbox with clear vulnerability-acceptance criteria; a companion mantis-advise skill helps coding agents write more secure code going forward. 53:16 Introducing Gemini 3.8 Flash and 3.8 Flash Cyber Google shipped its third Flash release in six weeks, with Gemini 3.8 Flash and a specialized 3.8 Flash Cyber variant, both priced at $0.75 per million input tokens and $3.75 per million output tokens, matching 3.7 Flash pricing. 3.8 Flash targets long-horizon coding and agentic tasks, outperforming larger frontier models on DeepSWE v1.1 and scoring 54.9% on HLE-Verified; it achieves this by using more reasoning steps and tool calls, so token usage can increase at higher effort settings. Developers who need lower cost can dial down effort levels or stick with 3.7 Flash. 3.8 Flash Cyber is restricted to trusted defenders through the new Fairwind Program, focusing on vulnerability discovery and automated patching rather than offensive capabilities. It reports a 70%+ success rate on internal multi-language vulnerability benchmarks and lands near the Pareto frontier on CWE-Bench patching (47.2% pass@1 vs 47.8% for a leading frontier model, at lower cost). Google cites internal validation: Chrome Security saw 2.6x more correct patches versus larger commercial models, Wiz reported 7.5-9.7% higher recall at 2.3-5.2x lower cost, and Google’s Cloud Vulnerability Research team found a critical vulnerability in under 2 hours using the model. Access points span the developer and enterprise stack: Google AI Studio , Android Studio , Google Antigravity , and Gemini Enterprise for 3.8 Flash; Google AI Pro/Ultra subscribers get it in the Gemini app , AI Mode , and Sheets . Cyber access requires application through the Fairwind Program , limiting availability to government authorities and critical infrastructure operators. 55:31 Matt – “Look! Another Flash model!” 58:05 MrBeast partners with Gemini and Google Health Google is expanding its multi-year relationship with Beast Industries beyond YouTube sponsorship into Gemini and Google Health integrations, marketing Gemini through a high-profile creator with 500 million subscribers . MrBeast’s September 5 video will feature Gemini being used to help his team navigate survival challenges in jungle, desert, and Arctic environments, positioning Gemini as a tool for real-time hazard identification and decision-making. The partnership includes a Gemini ad campaign spot showing MrBeast using the app to coordinate logistics for his large-scale video productions, an example use case for AI-assisted project planning. Google Health and Fitbit are also part of the deal, with Fitbit Ace integrated into an upcoming MrBeast challenge, tying consumer wellness hardware into the broader promotional push. For listeners, this is primarily a consumer marketing story rather than a new enterprise feature, but it signals how Google is using influencer partnerships to drive Gemini app adoption and visibility ahead of broader competitive AI assistant marketing pushes. This story is 100% to make us look cooler to our kids. Show Note Editor Heather: Adding to what Ryan mentioned about using AI for travel stuff; here’s a use case I utilized recently. I had 4 days in London to myself, and as a PhD candidate in history and museum studies, there was A LOT of old stuff to see – and one of my favorite things is GPS-guided audio tours, specifically using the App VoiceMap. I made a list of all the ones I wanted to do and gave the beginning and ending coordinates to ChatGPT and had it give me the best order of the walking tours. That way, it was the most efficient from the end of one to the starting location of the next one. I also utilize AI to help me locate less well-known historic or archaeological sites that are off the beaten path when traveling in Turkiye. 1:04:01 WeatherNext 3: Our most advanced global weather AI model WeatherNext 3 shifts training data from lagged NWP simulations to live geostationary satellite mosaics and sparse weather station observations, enabling hourly forecast updates at up to 5-kilometer resolution, roughly five times sharper than WeatherNext 2’s 25-kilometer, 6-hour cycle. Precipitation forecasting shows measurable accuracy gains, with CRPS improvements up to 60% against NASA IMERG data, 30% against MRMS, and 10% against rain gauge measurements for early lead times, addressing a longstanding weak point for AI weather models. The model adds renewable energy-specific outputs, including 100-meter turbine-height wind speed forecasts and solar radiation/cloud cover predictions, giving grid operators and clean energy developers data to match generation forecasts with demand planning. Rollout spans Google Search, Gemini app, Google Maps, the Maps Platform Weather API, and Earth Engine starting immediately, with claimed improvements of up to 50% more accurate precipitation forecasts for day-plus planning, particularly in Latin America, Africa, and Asia-Pacific regions that previously lacked high-resolution regional forecasting due to compute costs. Azure 1:09:35 Generally Available: Windows Server 2025 on AKS Windows Server 2025 support on AKS is now generally available, giving customers a path forward as older Windows Server versions approach end of support. Key improvements include Stable ABI for driver compatibility, Generation 2 VM as the default, containerd 2.0 runtime, and FIPS compliance enabled by default for regulated workloads. This targets enterprises running Windows-based containerized workloads who need to modernize their AKS clusters ahead of Windows Server 2025 lifecycle deadlines. The FIPS-by-default setting is notable for organizations in government, finance, or healthcare that require compliance with federal cryptographic standards without additional configuration. No specific pricing details are provided in the announcement; standard AKS compute and Windows Server licensing costs likely apply based on node pool configuration. Emerging Clouds 1:11:56 Introducing context-aware vulnerability discovery and remediation with Cloudflare Managed Defense and OpenAI Daybreak Models | Cloudflare Blog Cloudflare is launching Vulnerability Discovery and Remediation , an invitation-only service that pairs OpenAI Daybreak models, including GPT-5.6 Cyber, with Cloudflare network data to prioritize which vulnerabilities matter most rather than just listing findings. The key differentiator is production context: the system cross-references code vulnerabilities with actual traffic data, active routes, and existing WAF rules to determine real-world exposure, addressing the common problem of scanners flagging thousands of issues with no way to rank urgency. The architecture keeps humans in control – the model can propose code patches and WAF rules, but cannot apply them directly. All proposals pass through validation checks and customer review before any changes are deployed. The technical workflow uses a multi-agent pipeline (reconnaissance, hunting, validation agents) that maps production routes to source code, with model inference happening on OpenAI’s servers via Cloudflare AI Gateway rather than at the edge, and includes redaction controls to limit what data reaches the model. This builds on Cloudflare’s internal vulnerability harness (previously discussed in their “Build your own vulnerability harness” post) and extends that fleet-scanning capability to customer codebases, signaling a broader trend of combining LLM-based code analysis with infrastructure-level telemetry for security prioritization. 1:13:59 Ryan – “I started thinking about doing this for workloads, just because vulnerability management has always been a problem – tracking mean time to resolution.” Closing And that is the week in the cloud!
Building Platforms for AI Agents, with Mauricio (Salaboy) Salatino
Non-deterministic agents pose specific challenges for platform teams in observability, state management, governance, and trust. Mauricio (Salaboy) Salatino explains why agentic applications behave like distributed multi-agent systems. His test assigned agents to take an order, cook the pizza, deliver it, and charge the customer. One order crossed 15 containers and produced 200 traces. In this interview: Why agent frameworks can recreate monolith scaling and resource contention How OpenTelemetry data can measure agent behavior and trust over time Why narrowly scoped agents are safer than agents that follow long sequences How the platform can become the learning layer that feeds better context back to LLMs Sponsor Kubernetes moves too fast to track everything. Learn Kubernetes Weekly filters out the noise to deliver one curated email with useful articles, tutorials, tools, jobs, events, and CFPs.
Your AI Agent Has Too Much Access | Fixing Agent Identity
AI agents can do far more than answer questions — they can call tools, access data, execute actions, and act on behalf of users. But once agents start taking actions, an important question emerges: who is the agent actually acting as, and what should it be allowed to do? In this episode of Kubernetes Bytes, Bhavin sits down with Maia Iyer, Research Software Engineer at IBM Research , to explore identity, authentication, authorization, and zero trust for AI agents. They discuss why static API keys and simply passing user credentials to agents can create security problems, and how existing standards such as OAuth 2.0, SPIFFE, SPIRE, and Keycloak can help establish verifiable user and workload identities. Maia also explains the work her team is doing around Rosso CTL and Cortex , including using a runtime layer around agents to handle identity, observe inbound and outbound interactions, apply guardrails, and explore intent-based access control . Topics covered include: • The building blocks of agentic applications • Security anti-patterns for AI agents • Static credentials and over-permissioned agents • User identity vs. workload identity • OAuth 2.0 and delegated identity • SPIFFE and SPIRE • Keycloak • Rosso CTL and Cortex • Intent-based access control • Prompt injection and unexpected tool calls • MCP and agent-to-tool communication • Stateful AI agents • Safely running agents that generate and execute code If you're building AI agents on Kubernetes or thinking about how identity and authorization need to evolve for agentic systems, this episode provides a practical mental model for getting started.
Agent Substrate, with Tim Hockin and Brandon Royal
Tim Hockin is a long term software engineer with Google Cloud and I would argue one of the fathers of Kubernetes. Brandon Royal is a product manager on GKE and has been behind the launch of multiple OSS projects like the Ray Operator for k8s, Agent Sandbox and our topic for today - Agent Substrate. Do you have something cool to share? Some questions?