As customers modernize to lakehouse architectures, they are standardizing on open formats such as Apache Iceberg to create a shared data estate across compatible engines. This enables you to build AI-native lakehouses that turn data to semantic knowledge, enable proactive action, and operate at an agentic scale.
One of the key components of a lakehouse is the catalog, and in the Apache Iceberg environment, that usually means the Iceberg REST Catalog. An Apache Iceberg catalog is responsible for maintaining table pointers, handling atomic commits, and serving as the single source of truth for table locations. But as large enterprise organizations modernize to lakehouses, they have begun to realize that they need a highly scalable and available managed catalog as a part of their lakehouse architecture. This becomes even more important as querying scales with agents. To support a large amount of repeated small queries from agents, you will need to build on top of a managed catalog that provides atomicity, consistency, availability and concurrency at massive scale.
In this blog, we explore the challenges a managed catalog faces in modern cloud environments at agent scale, and show you how Google Cloud’s serverless Lakehouse runtime catalog can help address them. Powered by Spanner, Google Cloud’s always-on database with virtually unlimited scale, and built to meet the open Apache Iceberg REST catalog specification, the Lakehouse runtime catalog is the highly scalable and available foundation you need for the agentic era.
Challenges a managed catalog needs to solve in the Lakehouse
When speaking with data engineers and infrastructure leads running production analytics at scale in a Lakehouse, the following core pain points consistently emerge and need to be solved by a managed catalog:
Atomic commits and concurrency control: Iceberg guarantees ACID transactions via optimistic concurrency control (OCC). A managed catalog must implement a bulletproof atomic compare-and-swap (CAS) operation to swap the current metadata pointer.
High availability and operational maintenance: Because queries fail immediately if the catalog is down, a managed catalog becomes a critical Tier-1 service.
Scaling the database backing the catalog: Catalog architects typically face a difficult trade-off when choosing a backing database for table metadata and state. Traditional scale-up relational databases provide SQL and ACID transactions, but hit vertical CPU, memory, storage and connection limits under heavy concurrent read/write loads unless manually sharded which incurs a huge operational overhead; while scale-out database systems are either eventually consistent, hard to manage, not enterprise-ready, or all of the above.
Table maintenance coordination: A catalog alone does not optimize data; you must build and operate ancillary pipelines for compaction, snapshot expiration, manifest rewriting, and orphan file cleanup.
Governance and security: A catalog acts as the security gatekeeper. The catalog must implement and maintain:
-
Authentication protocols (e.g., OAuth2 token exchange, IAM federation)
-
Access control down to namespace and table levels
-
Vended storage credentials (e.g., generating short-lived tokens so query engines don’t need broad, direct storage credentials)
Lakehouse runtime catalog
To solve these challenges, we built the Lakehouse runtime catalog (GA) with support for Iceberg Rest Catalog. The Lakehouse runtime catalog is a fully serverless, highly available, and unified metadata registry designed from the ground up to support modern open table formats like Apache Iceberg. By natively implementing the Apache Iceberg REST Catalog specification, the Lakehouse runtime catalog decouples metadata discovery from compute engines, helping ensure multiple Iceberg-compatible engines can access a shared data estate and enabling you to take your workloads to production sooner. We’ve helped many customers streamline the migration of their managed catalogs. For example, Etsy migrated its catalog to Lakehouse runtime catalog, joining data in place to accelerate pipeline queries by 60%.






