Amazon Web Services has introduced a major capability for Amazon Aurora PostgreSQL, allowing organizations to directly query operational data alongside large-scale data lake formats—including Apache Iceberg and Apache Parquet—using standard PostgreSQL applications and developer tools. This integration eliminates the traditional requirement for complex Extract, Transform, and Load (ETL) pipelines, fundamentally altering how enterprise applications bridge transactional databases with analytical data repositories.
The feature leverages embedded technology from DuckLabs, the core development team behind the open-source analytical database engine DuckDB, which was recently acquired by Amazon. By embedding DuckDB directly within the Aurora PostgreSQL architecture, query processing occurs natively within the database instance without requiring additional network hops or data duplication. The capability is immediately available across all commercial AWS regions and AWS GovCloud (US) environments for Aurora PostgreSQL versions 17 (starting with 17.11) and 18 (starting with 18.6) at no additional software licensing charge.
Background Context and Operational Challenges
For years, enterprise architecture faced a persistent divide between transactional systems of record and analytical data lakes. Applications requiring real-time transactional data alongside historical archives typically relied on reverse ETL pipelines to copy, convert, and synchronize data between Amazon S3 and relational databases.
These pipelines introduced significant engineering overhead, delayed data freshness, escalated infrastructure costs, and created synchronization vulnerabilities. Furthermore, as organizations increasingly deploy autonomous AI agents capable of reasoning over vast amounts of information, pre-replicating and predicting every dataset an agent might require has become practically impossible.

The integration of DuckDB directly into Aurora PostgreSQL addresses this architectural friction. By treating data lakes as first-class query targets via foreign tables, developers can construct unified queries spanning live operational writes and multi-year historical files residing in Amazon S3 or S3 Tables.
Technical Architecture and Implementation
Implementing the new direct-querying capability requires minimal configuration within the AWS Management Console or via standard database clients such as psql. Administrators must first associate an AWS Identity and Access Management (IAM) role possessing the AuroraAnalytics permissions policy with their Aurora PostgreSQL cluster. This role grants the database secure access to read objects stored in Amazon S3 and resolve catalog metadata through the AWS Glue Data Catalog.
Once the foundational permissions are established, database administrators enable the feature by executing a simple command:
CREATE EXTENSION aurora_analytics;
To bridge the relational database with external data lakes, users create foreign tables. Notably, the architecture supports automatic schema inference. By designating a foreign table pointing to a Parquet file or an Apache Iceberg table, Aurora reads the underlying file metadata directly, removing the need to manually define column definitions. For larger data lake environments containing numerous datasets, administrators can utilize bulk provisioning statements such as IMPORT FOREIGN SCHEMA to map entire AWS Glue Data Catalog databases instantaneously.
The system also supports Iceberg REST Catalog (IRC)-compatible registries through AWS Glue Data Catalog federation. Enterprises can register external catalogs with Glue once, allowing unified queries to seamlessly join operational tables in Aurora with distributed Iceberg tables residing across multiple catalogs without altering existing metadata investments.

Performance Optimization and Query Execution
To maintain high performance as underlying datasets scale into petabytes, Aurora PostgreSQL incorporates sophisticated query optimization techniques. The engine applies predicate pushdown and column pruning, ensuring that only relevant rows and columns are retrieved from Amazon S3.
Frequently accessed data lake segments are cached locally within the Aurora instance, accelerating subsequent query execution. Database operators can inspect these performance metrics at the individual query level by utilizing the aurora_analytics_stat_statements() utility, which reports granular statistics such as scanned rows, bytes read from S3, and cache-hit ratios.
For use cases requiring ultra-low, single-digit-millisecond latency, data engineers can materialize data lake segments directly into native Aurora PostgreSQL tables using standard DDL and DML operations like CREATE TABLE AS SELECT, INSERT INTO ... SELECT, or MERGE INTO. These materialized tables reside fully within the relational database, providing a high-performance path for hot analytical workloads. Because these read queries can be executed across any read replica within the cluster, organizations can effectively offload heavy analytical scans from their primary writer instances.
Chronology of the Integration
The path toward native data lake querying within relational databases reflects broader industry trends favoring decoupled storage and compute.
- Late 2023 to 2024: Industry adoption of open table formats like Apache Iceberg accelerates, as enterprises seek vendor-neutral alternatives to proprietary data warehouses. Organizations increasingly struggle with siloed data architectures spanning transactional operational stores and analytical data lakes.
- Mid-2025: DuckLabs, the engineering organization responsible for the high-performance analytical database engine DuckDB, joins Amazon. The acquisition signals an intentional effort to infuse vector-based analytical efficiency into core AWS transactional services.
- Late 2026: AWS formally announces the integration of DuckDB technology into Amazon Aurora PostgreSQL versions 17 and 18, eliminating ETL requirements for Apache Iceberg and Parquet formats in Amazon S3 and AWS Glue.
Industry Implications and Analysis
The introduction of native data lake querying within Aurora PostgreSQL carries notable implications for database administrators, data engineers, and enterprise software architects. By collapsing the traditional barrier between operational stores and analytical repositories, organizations can streamline system design and reduce infrastructure maintenance expenditures.

From an economic perspective, the feature is offered without additional software markups. Customers incur costs solely for the incremental Aurora compute consumed during query execution and standard Amazon S3 API request fees for reading underlying data files. This pricing structure lowers the barrier to entry for small and medium-sized enterprises seeking advanced data lake analytics capabilities without investing in separate, standalone data warehousing infrastructure.
Furthermore, the capability directly supports the proliferation of generative AI applications and autonomous agents. As enterprises build applications powered by large language models that require immediate context from both live operational tables and historical archives, unified query access ensures agents can retrieve comprehensive information streams in real time. By avoiding data duplication and synchronization delays, organizations can maintain stronger data governance while simplifying application logic across complex hybrid cloud environments.
